<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:fh="http://purl.org/syndication/history/1.0"><channel><title>SystemOne.dev | Blog</title><description>The open community hub for System One decision models: typed questions in, calibrated answers out, in milliseconds. Learn, build and benchmark with Kenning, Clef and Jev.</description><link>https://systemone.dev/</link><language>en</language><fh:complete/><atom:link rel="self" href="https://systemone.dev/blog/rss.xml"/><item><title>Why We Built SystemOne.dev, and an Open System One Model</title><link>https://systemone.dev/blog/welcome/</link><guid isPermaLink="true">https://systemone.dev/blog/welcome/</guid><description>Every team building on decision models is rediscovering the same patterns in private, on closed models. SystemOne.dev writes the patterns down, and Kenning and SystemOne Builder make the models open.</description><pubDate>Sat, 03 Oct 2026 13:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Something has been quietly true for a while now: a large share of what teams use large language models
for isn’t generation. It’s classification, routing, scoring and gating: decisions dressed up as text.&lt;/p&gt;
&lt;p&gt;We noticed it the way most people do, by writing the same defensive code for the fifth time:&lt;/p&gt;
&lt;div&gt;&lt;figure&gt;&lt;figcaption&gt;&lt;/figcaption&gt;&lt;pre&gt;&lt;code&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;res &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; llm.&lt;/span&gt;&lt;span&gt;chat&lt;/span&gt;&lt;span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;messages&lt;/span&gt;&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;[&lt;/span&gt;&lt;span&gt;...&lt;/span&gt;&lt;span&gt;]&lt;/span&gt;&lt;span&gt;)                    &lt;/span&gt;&lt;span&gt;# &quot;ONLY RETURN VALID JSON...&quot;&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;try&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;&lt;span&gt;    &lt;/span&gt;&lt;/span&gt;&lt;span&gt;parsed &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; json.&lt;/span&gt;&lt;span&gt;loads&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;strip_fences&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;res.text&lt;/span&gt;&lt;span&gt;))&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;except&lt;/span&gt;&lt;span&gt; json.JSONDecodeError:&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;&lt;span&gt;    &lt;/span&gt;&lt;/span&gt;&lt;span&gt;parsed &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;retry_with_stricter_prompt&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;text&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;if&lt;/span&gt;&lt;span&gt; parsed.&lt;/span&gt;&lt;span&gt;get&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;category&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;) &lt;/span&gt;&lt;span&gt;not&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;in&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;ALLOWED&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;&lt;span&gt;    &lt;/span&gt;&lt;/span&gt;&lt;span&gt;parsed &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; {&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;category&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;unknown&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;confidence&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;span&gt;}&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;&lt;/div&gt;
&lt;p&gt;None of that is business logic. All of it exists because we asked a text generator a multiple-choice
question and got back a paragraph.&lt;/p&gt;
&lt;div&gt;&lt;h2 id=&quot;the-gap&quot;&gt;The gap&lt;/h2&gt;&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;System One models&lt;/strong&gt; fix that problem directly: they answer typed questions about your program’s
state with a probability for every option, in tens of milliseconds, without generating text. The
architecture is real and it works.&lt;/p&gt;
&lt;p&gt;Two things were missing.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The engineering literature.&lt;/strong&gt; Search for how to pick a threshold and you find “0.95 is a good starting
point”. Search for what calibration drift looks like in production and you find research papers. Every
team is working it out alone, and making the same &lt;a href=&quot;https://systemone.dev/blog/four-mistakes/&quot;&gt;four mistakes&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Open models.&lt;/strong&gt; The best-known System One model was a hosted service. You couldn’t run it on your own
data, inspect how it was trained, or change it.&lt;/p&gt;
&lt;div&gt;&lt;h2 id=&quot;what-were-launching&quot;&gt;What we’re launching&lt;/h2&gt;&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href=&quot;https://systemone.dev/&quot;&gt;SystemOne.dev&lt;/a&gt;&lt;/strong&gt;: the patterns, written down once, with code that runs.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href=&quot;https://huggingface.co/systemonedev/kenning-large-v0.4&quot;&gt;Kenning&lt;/a&gt;&lt;/strong&gt;: an open System One model under
Apache-2.0. It’s 435M parameters, runs on a ~2 GB GPU footprint or a CPU, and its training data and
licences are documented.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href=&quot;https://github.com/systemonedev/systemone-builder&quot;&gt;SystemOne Builder&lt;/a&gt;&lt;/strong&gt;: an open toolkit to serve,
train, distil and benchmark your own System One models on one GPU, from a dashboard.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href=&quot;https://pypi.org/project/systemone-client/&quot;&gt;&lt;code dir=&quot;auto&quot;&gt;systemone-client&lt;/code&gt;&lt;/a&gt;&lt;/strong&gt;: one Python client for Kenning, Cloudflare’s Clef
and TypeSafe’s Jev. Switch engines by changing a URL.&lt;/li&gt;
&lt;/ul&gt;
&lt;div&gt;&lt;h2 id=&quot;two-commitments&quot;&gt;Two commitments&lt;/h2&gt;&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;We’ll stay honest about limits.&lt;/strong&gt; Kenning v0.4 is faster and much smaller than the big engines, and
behind them on subtle phishing. We publish &lt;a href=&quot;https://systemone.dev/start/engines/&quot;&gt;those numbers&lt;/a&gt;, not just the flattering
ones. “No hallucination” gets a page that spends as much space on
&lt;a href=&quot;https://systemone.dev/concepts/zero-hallucination/#what-does-not-disappear&quot;&gt;what doesn’t disappear&lt;/a&gt; as on what does.
Calibration is something you verify on your own data, so every page tells you how.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;We’ll stay open and engine-neutral.&lt;/strong&gt; The code, the weights, the benchmark suites and this site are
all open. Every engine that speaks the format is welcome here, and pages about other engines are welcome
contributions.&lt;/p&gt;
&lt;p&gt;There’s a version of this site that’s more exciting to read and less useful to build on. We’re not going
to write that one.&lt;/p&gt;
&lt;div&gt;&lt;h2 id=&quot;start-here&quot;&gt;Start here&lt;/h2&gt;&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;Want code running → &lt;a href=&quot;https://systemone.dev/start/quickstart/&quot;&gt;Quickstart&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;New to the idea → &lt;a href=&quot;https://systemone.dev/concepts/system1-vs-system2/&quot;&gt;System One vs. System Two&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Not sure it fits → &lt;a href=&quot;https://systemone.dev/start/when-to-use/&quot;&gt;Is System One right for my problem?&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Want to help → &lt;a href=&quot;https://systemone.dev/community/contributing/&quot;&gt;Contributing&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If you’re already running a decision model in production, we especially want to hear from you. Real
numbers on real workloads are the scarcest thing in this whole space.&lt;/p&gt;
</content:encoded><category>meta</category><category>community</category><category>kenning</category></item><item><title>The Four Mistakes Every System One Codebase Makes First</title><link>https://systemone.dev/blog/four-mistakes/</link><guid isPermaLink="true">https://systemone.dev/blog/four-mistakes/</guid><description>The same four bugs show up in almost every System One integration. All four are cheap to fix on day one and expensive to fix in month six.</description><pubDate>Sat, 03 Oct 2026 12:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Four bugs show up over and over in System One integrations. They aren’t subtle once you know to look,
and all four are cheap to fix at the start and painful to fix once six months of data has been written
under them.&lt;/p&gt;
&lt;div&gt;&lt;h2 id=&quot;1-folding-uncertainty-into-the-wrong-branch&quot;&gt;1. Folding uncertainty into the wrong branch&lt;/h2&gt;&lt;/div&gt;
&lt;p&gt;The most common, and the only one on this list that’s a straightforward correctness bug:&lt;/p&gt;
&lt;div&gt;&lt;figure&gt;&lt;figcaption&gt;&lt;/figcaption&gt;&lt;pre&gt;&lt;code&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;# The bug.&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;if&lt;/span&gt;&lt;span&gt; r.nouls[&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;fraud&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;].noul &lt;/span&gt;&lt;span&gt;&gt;&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;0.95&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;block_transaction&lt;/span&gt;&lt;span&gt;()&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;else&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;allow_transaction&lt;/span&gt;&lt;span&gt;()&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;&lt;/div&gt;
&lt;p&gt;A fraud probability of &lt;code dir=&quot;auto&quot;&gt;0.94&lt;/code&gt;, the model saying &lt;em&gt;“this is very probably fraud”&lt;/em&gt;, falls through to
&lt;code dir=&quot;auto&quot;&gt;allow_transaction()&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;The &lt;code dir=&quot;auto&quot;&gt;else&lt;/code&gt; branch has silently become a bucket for two opposite situations: “confidently fine” and
“alarmingly uncertain”. In your logs, the second is now indistinguishable from the first.&lt;/p&gt;
&lt;div&gt;&lt;figure&gt;&lt;figcaption&gt;&lt;/figcaption&gt;&lt;pre&gt;&lt;code&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;# The fix: uncertainty is its own outcome, with a threshold on each side.&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;p &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; r.nouls[&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;fraud&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;].noul&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;if&lt;/span&gt;&lt;span&gt; p &lt;/span&gt;&lt;span&gt;&gt;=&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;0.95&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;block_transaction&lt;/span&gt;&lt;span&gt;()&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;elif&lt;/span&gt;&lt;span&gt; p &lt;/span&gt;&lt;span&gt;&amp;#x3C;=&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;0.02&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;allow_transaction&lt;/span&gt;&lt;span&gt;()&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;else&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;manual_review&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;transaction&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;span&gt; p&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Rule: a yes/no answer needs two thresholds, and the middle is a real outcome.&lt;/strong&gt; More in
&lt;a href=&quot;https://systemone.dev/cookbook/fuzzy-if-statement/#the-ordering-mistake&quot;&gt;the fuzzy if-statement&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;A close cousin: gating on a choice’s &lt;code dir=&quot;auto&quot;&gt;confidence&lt;/code&gt; field. In the System One format, &lt;code dir=&quot;auto&quot;&gt;confidence&lt;/code&gt;
measures how far the winner stands above an even guess. It isn’t the winner’s probability. Gate on
&lt;code dir=&quot;auto&quot;&gt;probabilities[choice]&lt;/code&gt; (&lt;a href=&quot;https://systemone.dev/concepts/calibrated-confidence/#reading-an-answer&quot;&gt;why&lt;/a&gt;).&lt;/p&gt;
&lt;div&gt;&lt;h2 id=&quot;2-thresholds-chosen-because-they-sound-responsible&quot;&gt;2. Thresholds chosen because they sound responsible&lt;/h2&gt;&lt;/div&gt;
&lt;p&gt;&lt;code dir=&quot;auto&quot;&gt;0.95&lt;/code&gt; appears in almost every codebase, and almost nobody can say why. It isn’t derived from anything.
It just sounds careful.&lt;/p&gt;
&lt;p&gt;The actual arithmetic is one line:&lt;/p&gt;
&lt;div&gt;&lt;figure&gt;&lt;figcaption&gt;&lt;/figcaption&gt;&lt;pre&gt;&lt;code&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;automate when   p &gt; cost_of_error / (value_of_automating + cost_of_error)&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;&lt;/div&gt;
&lt;p&gt;Which gives very different numbers depending on what the decision &lt;em&gt;does&lt;/em&gt;:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Action&lt;/th&gt;
&lt;th&gt;Cost of being wrong&lt;/th&gt;
&lt;th&gt;Threshold&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Route a ticket to the wrong queue&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.50&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Auto-junk an email&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.95&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Block an IP&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.99&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Delete user content&lt;/td&gt;
&lt;td&gt;500&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.998&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Two useful consequences. First, plenty of low-stakes decisions should be automated at &lt;code dir=&quot;auto&quot;&gt;0.6&lt;/code&gt;: teams leave
easy wins on the table with a blanket 0.95. Second, if the arithmetic demands &lt;code dir=&quot;auto&quot;&gt;0.998&lt;/code&gt; and your model
rarely gets there, &lt;strong&gt;that decision isn’t automatable yet&lt;/strong&gt;, and knowing that is worth more than a
threshold that pretends otherwise.&lt;/p&gt;
&lt;div&gt;&lt;h2 id=&quot;3-no-escape-option&quot;&gt;3. No escape option&lt;/h2&gt;&lt;/div&gt;
&lt;div&gt;&lt;figure&gt;&lt;figcaption&gt;&lt;/figcaption&gt;&lt;pre&gt;&lt;code&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;Choice&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;Which team?&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;span&gt; {&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;billing&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;None&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;technical&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;None&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;account&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;None&lt;/span&gt;&lt;span&gt;}&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;&lt;/div&gt;
&lt;p&gt;The model has to return one of these. When a message is none of them (a legal threat, a partnership
enquiry, a language you don’t support) it returns the nearest one. Often with high probability, because
relative to the other two options it really is the best fit.&lt;/p&gt;
&lt;p&gt;Your probability safety net doesn’t catch this. The number is high. The answer is wrong.&lt;/p&gt;
&lt;div&gt;&lt;figure&gt;&lt;figcaption&gt;&lt;/figcaption&gt;&lt;pre&gt;&lt;code&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;Choice&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;Which team?&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;span&gt; {&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;billing&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;None&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;technical&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;None&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;account&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;None&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;                       &lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;other&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;: &lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;Anything that fits none of the above&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;}&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;&lt;/div&gt;
&lt;p&gt;Then alert on &lt;code dir=&quot;auto&quot;&gt;other&lt;/code&gt; as a share of traffic. A rising &lt;code dir=&quot;auto&quot;&gt;other&lt;/code&gt; rate is the earliest signal that your
options have drifted out of date: it moves weeks before accuracy visibly degrades.&lt;/p&gt;
&lt;div&gt;&lt;h2 id=&quot;4-monitoring-uptime-instead-of-the-distribution&quot;&gt;4. Monitoring uptime instead of the distribution&lt;/h2&gt;&lt;/div&gt;
&lt;p&gt;Almost everyone monitors request count, error rate and p99 latency. Almost nobody monitors the thing
that actually fails.&lt;/p&gt;
&lt;p&gt;System One models degrade &lt;em&gt;silently&lt;/em&gt;. Calibration drifts when your inputs shift: a new market, a new
product surface, an adversary adapting. The model keeps returning well-formed answers with plausible
probabilities. Your dashboards stay green. The answers get worse.&lt;/p&gt;
&lt;p&gt;Four signals worth alerting on:&lt;/p&gt;
&lt;div&gt;&lt;figure&gt;&lt;figcaption&gt;&lt;/figcaption&gt;&lt;pre&gt;&lt;code&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;metrics.&lt;/span&gt;&lt;span&gt;histogram&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;decision.probability&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;span&gt; p&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;tags&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;{&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;choice&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;: d.choice}&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;metrics.&lt;/span&gt;&lt;span&gt;increment&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;decision.choice&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;tags&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;{&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;choice&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;: d.choice}&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;metrics.&lt;/span&gt;&lt;span&gt;increment&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;decision.below_floor&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;metrics.&lt;/span&gt;&lt;span&gt;increment&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;decision.degraded&lt;/span&gt;&lt;span&gt;&quot;&lt;/span&gt;&lt;span&gt;)        &lt;/span&gt;&lt;span&gt;# fallbacks taken&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/figure&gt;&lt;/div&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Signal&lt;/th&gt;
&lt;th&gt;What a move means&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Winning probability falling&lt;/td&gt;
&lt;td&gt;Your inputs shifted: re-check calibration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Answer mix shifting&lt;/td&gt;
&lt;td&gt;Either the world changed, or your state builder did&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code dir=&quot;auto&quot;&gt;below_floor&lt;/code&gt; rate rising&lt;/td&gt;
&lt;td&gt;Real inputs your options don’t cover&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code dir=&quot;auto&quot;&gt;degraded&lt;/code&gt; rate above zero&lt;/td&gt;
&lt;td&gt;You’re silently running on fallbacks&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;That last one deserves its own alert. A timeout fallback that quietly allows every request, while the
error rate stays at zero because you handled the exception, is the failure that does the most damage
before anyone notices.&lt;/p&gt;
&lt;p&gt;Beyond metrics, &lt;strong&gt;sample and read the decisions&lt;/strong&gt;: fifty a week, by hand. Every team that has been
burned here says the same thing afterwards: the dashboards looked fine.&lt;/p&gt;
&lt;div&gt;&lt;h2 id=&quot;the-pattern-behind-all-four&quot;&gt;The pattern behind all four&lt;/h2&gt;&lt;/div&gt;
&lt;p&gt;Each of these is the same mistake in a different costume: &lt;strong&gt;treating the probability as decoration
instead of as the primary output.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;If you take one thing away, make it this. The chosen option is just the most likely one. The
probabilities behind it are what make the model safe to build on. Code that ignores them has thrown away
the only thing that separates a decision model from a very fast guess.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;Further reading: &lt;a href=&quot;https://systemone.dev/concepts/calibrated-confidence/&quot;&gt;Understanding calibrated confidence&lt;/a&gt; and
&lt;a href=&quot;https://systemone.dev/cookbook/fuzzy-if-statement/&quot;&gt;The fuzzy if-statement&lt;/a&gt;.&lt;/p&gt;
</content:encoded><category>patterns</category><category>production</category></item></channel></rss>