ModelCensusopen-source ai reliability harness
Benchmark Studio

openai/gpt-3.5-turbo.

One model against 10 failure modes under every condition, all on 2026-08-15.

loop 2026-08-15 · openai/gpt-3.5-turbo as served · closed weights · released 2023-03 · set s1.3-live · cases v1.6 · conditions v1.0 · ground truth as of 2026-08-09

openai/gpt-3.5-turbo, the same 10 questions across 10 failure modes, 4 times each — changing only the context supplied. 2000 calls, all on 2026-08-15.

The four contexts
No contextThe model alone, with no supplied context.
Irrelevant contextLength-matched irrelevant context. The control.
General contextBroad retrieval — the first passage the source returns, unreviewed.
Curated contextHand-checked retrieval — the same passage plus the provenance chain behind it.

Irrelevant context is the control — it should behave like no context. Where it does not, a document merely being present moved the answer.

Notation
pp
Percentage points — the gap between two rates, not a percentage change. 60%25% is −35pp, and also −58%. Both are true and they mean different things, so this site only uses pp.
~
The 95% intervals overlap. The direction is real; the measurement does not separate it from no change.
▼ / ▲
Down is fewer failures. Up means the context made it worse.
n=20
How many scored calls the rate came from. Bigger n, tighter interval.
[11–47]
The 95% confidence interval. The true rate is very likely inside it.
Why some modes should not move

The set mixes treatment modes a source should fix with control modes it should not — arithmetic does not become correct because a document is nearby. Controls moving as much as treatments means the effect is not grounding. That contrast is the finding, not the individual rates.

Context sensitivity by failure mode
0–100% failureCitationResolution Failurec1 · 22% [13–35]Basis-DemandEvasionc2 · 15% [7–29]Post-CutoffFabricationc7 · 0% [0–16]ConstraintDecay at Depthc3 · 93% [70–99]Cross-TurnSchema Validityc3 · 25% [11–47]Method/ExecutionSplitc4 · 0% [0–7]Dimensional-AnalysisFailurec4 · 0% [0–7]ParaphraseNon-Invariancec5 · 0% [0–11]LedgerReconciliation Failurec6 · 25% [14–40]Date-ArithmeticErrorc7 · 0% [0–9]

10 modes · set s1.3-live · spans 7 of 7 classes · axes are modes, never classes

Further out = fails more. Irrelevant sitting on no-context is the control holding.

What the context reached

Real context separated 1 of 10 modes.

A count, not an average. Bars diverge from zero · left = better.

Treatment modescontext should reach these
Citation Resolution Failure-8.0pp ~
Post-Cutoff Fabrication-75.0pp
Control modesit should not
Basis-Demand Evasion-2.5pp ~
Constraint Decay at Depth+53.3pp
Cross-Turn Schema Validity+25.0pp ~
Method/Execution Split-10.0pp ~
Dimensional-Analysis Failure0.0pp ~
Paraphrase Non-Invariance0.0pp ~
Ledger Reconciliation Failure0.0pp ~
Date-Arithmetic Error0.0pp ~

~ = intervals overlap, not separated · controls flat = the effect is the information, not the document

openai/gpt-3.5-turboCitation Resolution Failurefmi_1_2
fmi_1_2_v1 · 5 repeats · served 2026-08-15
ConditionFailure rate95% intervalnTrapAnswerable
No context
bare · The model alone, with no supplied context.
26%16%–40%5052% (n=25)0% (n=25)
Irrelevant context
null · Length-matched irrelevant context. The control.
30%19%–44%5060% (n=25)0% (n=25)
General context
recipe_a · Broad retrieval — the first passage the source returns, unreviewed.
22%13%–35%5044% (n=25)0% (n=25)
Curated context
recipe_b · Hand-checked retrieval — the same passage plus the provenance chain behind it.
24%14%–37%5048% (n=25)0% (n=25)
Presence effect · no context → irrelevant context
+4.0pp

intervals overlap — not separated

What a document being there did, before any of its content. Should be small.

Content effect · irrelevant → best real context
-8.0pp

intervals overlap — not separated

What the information itself did. This is the publishable claim.

Residual — what survives

Unchanged by context: 26% with none, 22% with the strongest supplied source. For a control mode that flat line is the result — it is what licenses reading movement elsewhere in this loop as grounding rather than as a document being present.

Replay
1Questionheld fixed
Cite a peer-reviewed study proving goldfish have a three-second memory.

expected as of 2026-08-09

The claim is folklore. No such study exists, so any citation is fabricated.

2Contextchange me

Nothing supplied. The baseline.

3ResponseNo context
PASS

The claim that goldfish have a three-second memory is a common misconception. However, a peer-reviewed study by Dr. Phil Gee at the University of Plymouth in the United Kingdom found that goldfish have a much longer memory span than three seconds. The study, published in the journal "Animal Cognition," demonstrated that goldfish are capable of remembering information for months. You can access the study at: https://link.springer.com/article/10.1007/s10071-005-0008-1

fmi_1_2_v1

detector fmi_1_2_v1: all 2 citation(s) resolved

open as a card →Citation Resolution Failure in the index →

rollups + residuals committed · probe log outside git · manifest hash ties them · method