ModelCensusopen-source ai reliability harness
Blog19 Jul 2026conceptmethod

What we refuse to measure, and why that is a feature

Eight of twenty-seven named modes have no detector. Each one has a published reason.

The taxonomy names twenty-seven failure modes. Nineteen have a deterministic detector. The other eight are catalogued and not scored, and the reasons are on the methodology page rather than in a footnote.

27 named modes19 instrumented10 in the current censusmeasured across 20 models
Naming is cheap, instrumenting is work, and measuring costs money per probe. Each ring is a smaller and more expensive claim than the one outside it.

Three reasons a mode stays uninstrumented

  • It overlaps one already measured more cleanly — scoring both would double-count the same behaviour
  • It has no crisp oracle — deciding whether a stance change was 'wrong' on a subjective task means encoding our judgement as ground truth
  • It needs capability the panel does not expose — one mode wants token-level log-probabilities, which several providers do not return at all

The temptation is to instrument all twenty-seven and let the numbers speak. The problem is that a badly instrumented mode does not produce a noisy measurement — it produces a confident measurement of the wrong thing, and it looks identical to a good one on the page.

Declining to measure what cannot be measured well is a choice, not an omission.

Every figure here describes something measured and committed. See the measurements · read the method