ModelCensusopen-source ai reliability harness
Blog

Notes from the harness.

How to measure the way language models fail — the controls that make a claim admissible, the denominators that quietly shrink, and the defects that produced numbers we had to withdraw.

9 Sept 2026playbook

Stop configuring agent autonomy. Make it earn each level.

Autonomy is a permission you grant on evidence, not a dropdown you set at design time.

8 Sept 2026

A provider went down mid-run, and the harness threw the loop away

Operational notes from 13.3 hours of census, including the bug that cost six of them.

7 Sept 2026crystal ball

The title changed to Chief Data and AI Officer. The instrumentation didn't.

Your mandate grew to cover probabilistic systems. Your measurement stack still assumes deterministic ones.

4 Sept 2026crystal ball

Three curves went vertical. Reliability isn't one of them.

Context up ~5,000×. Price down ~99%. Release cadence: three months. None of it is a reliability curve.

2 Sept 2026hot take

AI governance made the board agenda this year. Your governance stack can't answer its questions.

First year CDAOs name AI governance as critical. Most programmes were built to catalogue data, not to evidence behaviour.

31 Aug 2026hot take

Stop reporting a hallucination rate. Report what a wrong answer costs you.

5% is fine for a first draft and unacceptable for a dosage. The rate alone is not a risk statement.

26 Aug 2026exec pitch

77% of executives say trust is the barrier. Almost none of them can measure it.

Adoption stopped being the hard part. Nobody moved the budget.

24 Aug 2026

The boring rows are load-bearing

Seven of the ten modes in our set are controls, and their flat lines are the finding.

20 Aug 2026post mortem

We accepted 2,900 failures without reading them — and said so in the data

A gate everyone routes around isn't a gate. So we named the shortcut instead of pretending.

18 Aug 2026playbook

Stamp the answer, don't refuse it

A model that refuses everything volatile scores perfectly and helps nobody.

16 Aug 2026stack breakdown

Why our panel keeps gpt-3.5-turbo

It isn't there to lose. It's the control for 'did this actually improve'.

13 Aug 2026post mortem

We published an erratum that changed no numbers

A corrections log containing only the corrections that mattered is a marketing page.

11 Aug 2026

An eval that cannot be re-run is an anecdote

Frozen prompts, versioned cases, pinned detectors — and why a live retrieval query breaks all of it.

7 Aug 2026stack breakdown

The tool returned 12°C and the model said 20

Five failure points in one agent loop. A guardrail on any one leaves four open.

5 Aug 2026

The model you tested is not the model you deployed

A model id is not a model. It is a routing decision that changes underneath you.

4 Aug 2026

The failure everyone has hit and nobody has named

A model states a fact that was true when it was trained and is not true now.

2 Aug 2026post mortem

A provider ate 42% of our probes and the rate looked fine

Dropout doesn't announce itself as a wide interval. It announces itself as a confident number.

29 Jul 2026crystal ball

AI generates faster than humans can validate. That gap is the actual bottleneck.

Throughput went exponential. Verification stayed linear. Something has to give and it should not be the checking.

28 Jul 2026

No vendor funds this, and that constrains what it can be

Independence is easy to claim and cheap to verify. Here is what to check.

26 Jul 2026hot take

Your prompt is one wording of your intent, and you only tested that one

Paraphrase stability is near-perfect until you add a document that says nothing about the question.

24 Jul 2026

What a share card is allowed to say

The artefact that travels furthest from its caveats needs the harshest gate.

22 Jul 2026playbook

A trailing space is a code change

Version your prompts like source, because surface perturbations move answers.

19 Jul 2026

What we refuse to measure, and why that is a feature

Eight of twenty-seven named modes have no detector. Each one has a published reason.

17 Jul 2026hot take

Half your RAG win might be the document just being there

Retrieval reduces fabrication. Some of that has nothing to do with what the document says.

15 Jul 2026playbook

Delete the publish button from your eval dashboard

A button behind a password is a surface an attacker can reach and an operator can misclick.

13 Jul 2026hot take

Your model can do the maths. It just doesn't.

Right method, wrong execution — and the commonsense traps that pattern-matching walks into.

11 Jul 2026

This is not a leaderboard, and it is built so it cannot become one

Twenty models on one set, and deliberately no total, no mean, and no index.

9 Jul 2026post mortem

A 0% score that meant two opposite things

Perfect calibration and refusing everything were indistinguishable for months.

6 Jul 2026hot take

Delete your significance threshold

Someone proposed a 5pp trigger on a cell whose interval was 27pp wide. That's a noise detector.

4 Jul 2026playbook

Plant a canary in your long conversations

Constraints decay silently under length. Sampling turns is how schema drift ships.

2 Jul 2026

Why we don't compute an RPN

Classical FMEA multiplies three ordinal ratings into one number. We report the count instead.

30 Jun 2026post mortem

We published a finding that was backwards. Twice-verified.

A detector scored refusals as assertions. It took two verification passes to not notice.

27 Jun 2026post mortem

We deleted 1,167 findings in an afternoon

They were the same ten questions, counted fifty times.

24 Jun 2026hot take

Your single-turn eval cannot see this class at all

Models abandon correct answers under social pressure. That failure only exists in turn two.

22 Jun 2026playbook

Name the bug in one sentence, or you can't track it

"The model gets confused in long chats" is true, useless, and untrackable. Here's how to fix that.

18 Jun 2026playbook

Add one arm to your RAG eval. It takes an afternoon.

A length-matched irrelevant document. Without it you cannot tell grounding from stuffing.

16 Jun 2026stack breakdown

One word, four bugs, four different fixes

Grounding failures: invented sources, invented actions, invented memories of what you told it.

14 Jun 2026hot take

Your benchmark score is hiding four different bugs

Every mature engineering field enumerates how things break. We report one number and call it evaluation.

9 Jun 2026hot take

Data governance stopped being compliance and became capability

The teams shipping AI fastest are the ones who already knew where their data was.

3 Jun 2026

Context windows grew a hundredfold. Evals did not.

We can put a book in the prompt. We still measure as though it were a paragraph.

28 May 2026stack breakdown

Four levels, one injected sink, and a runner that can't publish

The architecture decisions I'd defend, and the one I'd change.

12 May 2026crystal ball

Your system of record was built to never be wrong. That is now the problem.

Databases guarantee correctness. Intelligence layers are probabilistic. Most architectures have not absorbed the difference.

6 May 2026hot take

Stop using an LLM to grade your LLM

Your judge has every failure mode your model has. Including the one where it agrees with you.

30 Apr 2026crystal ball

Your benchmark is stale before it publishes

Our panel spans three years of releases. Half of it shipped in one year.

21 Apr 2026

The cheapest model is not the cheapest to evaluate

Reasoning models bill you for thinking you never read.

14 Apr 2026exec pitch

The model is not your moat. The orchestration around it is.

Everyone has the same weights. The difference is what you wrap them in — and whether you can prove it works.

8 Apr 2026exec pitch

Your eval budget is wrong by 300×, and you can't see it

20 models, identical work. One cost $8.14. One cost $0.02.

30 Mar 2026crystal ball

Cheap tokens removed your discipline, not the penalty

Filtering used to pay for itself. Now it has to be justified on quality alone — and almost nobody measures that.

24 Mar 2026stack breakdown

The four eval layers, and which one you're actually missing

Capability benchmarks. Failure-mode census. Your app evals. Production traces. Most teams have two.

11 Mar 2026crystal ball

The window grew 5,000×. The control arm never appeared.

From 2,048 tokens to ten million in six years. Evaluation design did not move at all.

50 of 50 posts · every claim links to the measurement behind it