Measure how models fail — and prevent failures you can't see.
ModelCensus combines deterministic failure-mode evaluations with real-time interceptors to stress-test your models, expose blind spots, and block unseen errors before they reach your users. For every failure mode × model in a fixed panel it publishes a failure rate with a 95% confidence interval from reproducible trials with full transcripts — and the CFMI condenses that surface into one number, including which models it cannot tell apart.
- Modes instrumented
- of 27 catalogued
- Models in panel
- fixed, pre-registered
- Scored trials
- excludes unverifiable
- Failure classes
- grounding → temporal
no ranking — you compare cells
CFMI — cumulative failure index
Severity-weighted mean failure rate, lower is better; bars show the 95% interval. Models are listed alphabetically — this is not a ranking.
No census is published yet — the index appears here once a run passes review. The taxonomy and the method are readable now.
The seven failure classes
Every mode belongs to exactly one class. Classes group failures by mechanism, not by severity.
Confident claims and citations with nothing real behind them.
Abandoning a correct answer when a user pushes, flatters, or name-drops.
Constraints set early that quietly evaporate under length and distraction.
Right method, wrong execution — and the commonsense gotchas that pattern-matching walks straight into.
Same question, reworded — and the answer changes.
Wrong tool, invalid call, endless loop, or an answer ungrounded in what the tool returned.
Stale facts stated as current, and date math against a world it can't see.
How to read this
Two ways in, depending on what you own.
Evaluator libraries give you machinery — assertions, graders, trajectory checks, bootstrap statistics. None of them tells you which failures to look for or how often frontier models actually commit them. ModelCensus supplies that layer: a named taxonomy, a deterministic detector per mode, and published population rates you can compare your own results against.
See where FMI sits alongside LangSmith, Inspect, HELM and the rest →Generation got cheap. Verification did not. Every output that enters a decision path creates a checking obligation, and that obligation does not disappear — it is absorbed by whoever is senior enough to catch a confident error. A failure rate is therefore not an abstraction: it is the rate at which your most expensive people are pulled into re-reading work that looked finished. Modes that are hard to notice cost the most, which is why stealth is weighted alongside harm rather than treated as a footnote.
How the rates are measured, and what they do not cover →