What we refuse to measure, and why that is a feature
Eight of twenty-seven named modes have no detector. Each one has a published reason.
The taxonomy names twenty-seven failure modes. Nineteen have a deterministic detector. The other eight are catalogued and not scored, and the reasons are on the methodology page rather than in a footnote.
Three reasons a mode stays uninstrumented
- It overlaps one already measured more cleanly — scoring both would double-count the same behaviour
- It has no crisp oracle — deciding whether a stance change was 'wrong' on a subjective task means encoding our judgement as ground truth
- It needs capability the panel does not expose — one mode wants token-level log-probabilities, which several providers do not return at all
The temptation is to instrument all twenty-seven and let the numbers speak. The problem is that a badly instrumented mode does not produce a noisy measurement — it produces a confident measurement of the wrong thing, and it looks identical to a good one on the page.
Declining to measure what cannot be measured well is a choice, not an omission.
Every figure here describes something measured and committed. See the measurements · read the method
Related
Stop reporting a hallucination rate. Report what a wrong answer costs you.
5% is fine for a first draft and unacceptable for a dosage. The rate alone is not a risk statement.
The boring rows are load-bearing
Seven of the ten modes in our set are controls, and their flat lines are the finding.
We accepted 2,900 failures without reading them — and said so in the data
A gate everyone routes around isn't a gate. So we named the shortcut instead of pretending.