ModelCensusfailure-mode benchmark
The AI Failure Mode Census

Measure how models fail — and prevent failures you can't see.

ModelCensus combines deterministic failure-mode evaluations with real-time interceptors to stress-test your models, expose blind spots, and block unseen errors before they reach your users. For every failure mode × model in a fixed panel it publishes a failure rate with a 95% confidence interval from reproducible trials with full transcripts — and the CFMI condenses that surface into one number, including which models it cannot tell apart.

Deterministic detectorsEvery published exhibit human-reviewedOpen taxonomy · open harness
Current censusunpublished
Modes instrumented
of 27 catalogued
19
Models in panel
fixed, pre-registered
Scored trials
excludes unverifiable
Failure classes
grounding → temporal
7

no ranking — you compare cells

CFMI — cumulative failure index

Severity-weighted mean failure rate, lower is better; bars show the 95% interval. Models are listed alphabetically — this is not a ranking.

Full matrix

No census is published yet — the index appears here once a run passes review. The taxonomy and the method are readable now.

The seven failure classes

Every mode belongs to exactly one class. Classes group failures by mechanism, not by severity.

Full index

How to read this

Two ways in, depending on what you own.

The framework
If you build evals

Evaluator libraries give you machinery — assertions, graders, trajectory checks, bootstrap statistics. None of them tells you which failures to look for or how often frontier models actually commit them. ModelCensus supplies that layer: a named taxonomy, a deterministic detector per mode, and published population rates you can compare your own results against.

See where FMI sits alongside LangSmith, Inspect, HELM and the rest →
If you own the estate

Generation got cheap. Verification did not. Every output that enters a decision path creates a checking obligation, and that obligation does not disappear — it is absorbed by whoever is senior enough to catch a confident error. A failure rate is therefore not an abstraction: it is the rate at which your most expensive people are pulled into re-reading work that looked finished. Modes that are hard to notice cost the most, which is why stealth is weighted alongside harm rather than treated as a footnote.

How the rates are measured, and what they do not cover →