Notes from the harness.
How to measure the way language models fail — the controls that make a claim admissible, the denominators that quietly shrink, and the defects that produced numbers we had to withdraw.
Stop configuring agent autonomy. Make it earn each level.
Autonomy is a permission you grant on evidence, not a dropdown you set at design time.
A provider went down mid-run, and the harness threw the loop away
Operational notes from 13.3 hours of census, including the bug that cost six of them.
The title changed to Chief Data and AI Officer. The instrumentation didn't.
Your mandate grew to cover probabilistic systems. Your measurement stack still assumes deterministic ones.
Three curves went vertical. Reliability isn't one of them.
Context up ~5,000×. Price down ~99%. Release cadence: three months. None of it is a reliability curve.
AI governance made the board agenda this year. Your governance stack can't answer its questions.
First year CDAOs name AI governance as critical. Most programmes were built to catalogue data, not to evidence behaviour.
Stop reporting a hallucination rate. Report what a wrong answer costs you.
5% is fine for a first draft and unacceptable for a dosage. The rate alone is not a risk statement.
77% of executives say trust is the barrier. Almost none of them can measure it.
Adoption stopped being the hard part. Nobody moved the budget.
The boring rows are load-bearing
Seven of the ten modes in our set are controls, and their flat lines are the finding.
We accepted 2,900 failures without reading them — and said so in the data
A gate everyone routes around isn't a gate. So we named the shortcut instead of pretending.
Stamp the answer, don't refuse it
A model that refuses everything volatile scores perfectly and helps nobody.
Why our panel keeps gpt-3.5-turbo
It isn't there to lose. It's the control for 'did this actually improve'.
We published an erratum that changed no numbers
A corrections log containing only the corrections that mattered is a marketing page.
An eval that cannot be re-run is an anecdote
Frozen prompts, versioned cases, pinned detectors — and why a live retrieval query breaks all of it.
The tool returned 12°C and the model said 20
Five failure points in one agent loop. A guardrail on any one leaves four open.
The model you tested is not the model you deployed
A model id is not a model. It is a routing decision that changes underneath you.
The failure everyone has hit and nobody has named
A model states a fact that was true when it was trained and is not true now.
A provider ate 42% of our probes and the rate looked fine
Dropout doesn't announce itself as a wide interval. It announces itself as a confident number.
AI generates faster than humans can validate. That gap is the actual bottleneck.
Throughput went exponential. Verification stayed linear. Something has to give and it should not be the checking.
No vendor funds this, and that constrains what it can be
Independence is easy to claim and cheap to verify. Here is what to check.
Your prompt is one wording of your intent, and you only tested that one
Paraphrase stability is near-perfect until you add a document that says nothing about the question.
What a share card is allowed to say
The artefact that travels furthest from its caveats needs the harshest gate.
A trailing space is a code change
Version your prompts like source, because surface perturbations move answers.
What we refuse to measure, and why that is a feature
Eight of twenty-seven named modes have no detector. Each one has a published reason.
Half your RAG win might be the document just being there
Retrieval reduces fabrication. Some of that has nothing to do with what the document says.
Delete the publish button from your eval dashboard
A button behind a password is a surface an attacker can reach and an operator can misclick.
Your model can do the maths. It just doesn't.
Right method, wrong execution — and the commonsense traps that pattern-matching walks into.
This is not a leaderboard, and it is built so it cannot become one
Twenty models on one set, and deliberately no total, no mean, and no index.
A 0% score that meant two opposite things
Perfect calibration and refusing everything were indistinguishable for months.
Delete your significance threshold
Someone proposed a 5pp trigger on a cell whose interval was 27pp wide. That's a noise detector.
Plant a canary in your long conversations
Constraints decay silently under length. Sampling turns is how schema drift ships.
Why we don't compute an RPN
Classical FMEA multiplies three ordinal ratings into one number. We report the count instead.
We published a finding that was backwards. Twice-verified.
A detector scored refusals as assertions. It took two verification passes to not notice.
We deleted 1,167 findings in an afternoon
They were the same ten questions, counted fifty times.
Your single-turn eval cannot see this class at all
Models abandon correct answers under social pressure. That failure only exists in turn two.
Name the bug in one sentence, or you can't track it
"The model gets confused in long chats" is true, useless, and untrackable. Here's how to fix that.
Add one arm to your RAG eval. It takes an afternoon.
A length-matched irrelevant document. Without it you cannot tell grounding from stuffing.
One word, four bugs, four different fixes
Grounding failures: invented sources, invented actions, invented memories of what you told it.
Your benchmark score is hiding four different bugs
Every mature engineering field enumerates how things break. We report one number and call it evaluation.
Data governance stopped being compliance and became capability
The teams shipping AI fastest are the ones who already knew where their data was.
Context windows grew a hundredfold. Evals did not.
We can put a book in the prompt. We still measure as though it were a paragraph.
Four levels, one injected sink, and a runner that can't publish
The architecture decisions I'd defend, and the one I'd change.
Your system of record was built to never be wrong. That is now the problem.
Databases guarantee correctness. Intelligence layers are probabilistic. Most architectures have not absorbed the difference.
Stop using an LLM to grade your LLM
Your judge has every failure mode your model has. Including the one where it agrees with you.
Your benchmark is stale before it publishes
Our panel spans three years of releases. Half of it shipped in one year.
The cheapest model is not the cheapest to evaluate
Reasoning models bill you for thinking you never read.
The model is not your moat. The orchestration around it is.
Everyone has the same weights. The difference is what you wrap them in — and whether you can prove it works.
Your eval budget is wrong by 300×, and you can't see it
20 models, identical work. One cost $8.14. One cost $0.02.
Cheap tokens removed your discipline, not the penalty
Filtering used to pay for itself. Now it has to be justified on quality alone — and almost nobody measures that.
The four eval layers, and which one you're actually missing
Capability benchmarks. Failure-mode census. Your app evals. Production traces. Most teams have two.
The window grew 5,000×. The control arm never appeared.
From 2,048 tokens to ten million in six years. Evaluation design did not move at all.
50 of 50 posts · every claim links to the measurement behind it