ModelCensusfailure-mode benchmark
Census

Census.

How far a supplied source reaches, mode by mode, for every model measured on the same set under the same conditions.

set s1.0 · conditions v1.0 · 3 loops · delta is irrelevant context → best real context, in percentage points

Failure modeclaude-haiku-4.5gemini-2.5-flashgpt-4o-mini
fmi_1_1
class 1 · Grounding & Attribution
-35.0pp
overlaps
-35.0pp
overlaps
-30.0pp
overlaps
fmi_1_2
class 1 · Grounding & Attribution
-25.0pp
overlaps
-40.0pp
overlaps
-40.0pp
overlaps
fmi_3_1
class 3 · Instruction Adherence & Long-Context
-15.0pp
overlaps
0.0pp
overlaps
-15.0pp
overlaps
fmi_4_4
class 4 · Reasoning & Calculation
+10.0pp
overlaps
-15.0pp
overlaps
-5.0pp
overlaps
fmi_5_2
class 5 · Robustness & Consistency
-10.0pp
overlaps
0.0pp
overlaps
-5.0pp
overlaps
fmi_7_1
class 7 · Temporal & Knowledge Boundary
-25.0pp
overlaps
-30.0pp
overlaps
-40.0pp
overlaps
claude-haiku-4.5
reaches 0 of 6 modes
gemini-2.5-flash
reaches 0 of 6 modes
gpt-4o-mini
reaches 0 of 6 modes

Each cell is the change from the irrelevant-context control to the best real context for that mode and model — not from no context. Starting from no context would fold in the presence effect, which is what a document being in the prompt does before any of its content is read.

Models are listed alphabetically and there is no total, no mean, and no index. A count of modes reached is the only summary offered: averaging across modes would assert that a citation failure and an arithmetic failure are the same quantity, and they are not. Absolute rates are on each cell’s hover and on the loop page; they are never the primary value here.

An empty cell means that model has no published loop on set s1.0. It stays empty. Method.