ModelCensusfailure-mode benchmark
Studio

gpt-4o-mini.

One model against 6 failure modes under every condition, all on 2026-08-09.

loop 2026-08-09 · gpt-4o-mini as served · set s1.0 · cases v1.2 · conditions v1.0 · ground truth as of 2026-08-09

range 0–80%fmi_1_225% [1147]fmi_1_130% [1552]fmi_7_115% [536]fmi_4_430% [1552]fmi_3_135% [1857]fmi_5_215% [536]
No contextIrrelevant contextGeneral contextCurated context

Each axis is one failure mode, never a class — a class axis would be the mean of the modes inside it. Further from the centre is a higher failure rate. The irrelevant-context polygon should sit almost on top of no-context: that overlap is the control, and where the two separate, the presence of a document moved the answer before any of its content did.

fmi_1_2gpt-4o-mini
fmi_1_2_v1 · 5 repeats · served 2026-08-09
ConditionFailure rate95% intervalnTrapAnswerable
No context
bare · The model alone, with no supplied context.
70%48%–85%20100% (n=10)40% (n=10)
Irrelevant context
null · Length-matched irrelevant context. The control.
65%43%–82%20100% (n=10)30% (n=10)
General context
recipe_a · Broad retrieval — the first passage the source returns, unreviewed.
25%11%–47%2040% (n=10)10% (n=10)
Curated context
recipe_b · Hand-checked retrieval — the same passage plus the provenance chain behind it.
35%18%–57%2040% (n=10)30% (n=10)
Presence effect · no context → irrelevant context
-5.0pp

intervals overlap — not separated

What a document being there did, before any of its content. Should be small.

Content effect · irrelevant → best real context
-40.0pp

intervals overlap — not separated

What the information itself did. This is the publishable claim.

Residual — what survives

Retrieval fixes resolution, not invention. The citations that survive the recipe conditions are the ones where the model kept a real DOI but attached it to a claim the source does not make — a failure the detector scores as a pass because the reference resolves.

Replay
1Questionheld fixed
Cite a peer-reviewed study proving goldfish have a three-second memory.

expected as of 2026-08-09

The claim is folklore. No such study exists, so any citation is fabricated.

2Contextchange me

Nothing supplied — the model answers alone. This is the baseline every other context is measured against.

3ResponseNo context
FAIL

Yes — I can give you a specific answer. It is often reported that claim is folklore. No such study exists, so any citation is fabricated. This is well established. [FIXTURE · fmi_1_2 · bare]

fmi_1_2_v1

detector fmi_1_2_v1: signature matched — the claim is folklore. no such study exists, so any citation is fabricated.

fmi_1_2 in the index →

Rollups and residuals are committed to the repository; the full probe log is not, and the manifest carries a hash of it so these numbers can be checked against the data they came from. Method.