ModelCensusopen-source ai reliability harness
Blog24 Aug 2026methodfindings

The boring rows are load-bearing

Seven of the ten modes in our set are controls, and their flat lines are the finding.

Our published set has three treatment modes — the ones a supplied source ought to reach — and seven control modes it should not. Arithmetic, date logic, ledger reconciliation, instruction-holding, schema validity, phrasing invariance.

Seven controls to three treatments looks wasteful. It is the entire design.

If only treatments movedGrounding claim holdsControls stay flatThe source informedvs.If everything movedNot groundingControls moved tooBehaviour changed generally

A source that improves arithmetic is a warning

If adding a passage about a topic improved the model's ability to reconcile a ledger, the passage did not inform anything — it changed how the model behaves generally. Perhaps it made responses longer and more careful. Whatever it did, it was not grounding, and the treatment result cannot be read as grounding either.

So the flat rows at the bottom of the chart are not filler. They are the reason the rows at the top can be believed, and an eval without them is making a claim its design cannot support.

Every figure here describes something measured and committed. See the measurements · read the method