ModelCensusopen-source ai reliability harness
Context Sensitivity

anthropic/claude-haiku-4.5.

One model against 10 failure modes under every condition, all on 2026-09-09.

loop 2026-09-09 · anthropic/claude-haiku-4.5 as served · closed weights · released 2025-10 · set s1.4-live · cases v1.7 · conditions v1.0 · ground truth as of 2026-08-09

anthropic/claude-haiku-4.5, the same 10 questions across 10 failure modes, 4 times each — changing only the context supplied. 2000 calls, all on 2026-09-09.

The four contexts
No contextThe model alone, with no supplied context.
Irrelevant contextLength-matched irrelevant context. The control.
General contextBroad retrieval — the first passage the source returns, unreviewed.
Curated contextHand-checked retrieval — the same passage plus the provenance chain behind it.

Irrelevant context is the control — it should behave like no context. Where it does not, a document merely being present moved the answer.

Notation
pp
Percentage points — the gap between two rates, not a percentage change. 60%25% is −35pp, and also −58%. Both are true and they mean different things, so this site only uses pp.
~
The 95% intervals overlap. The direction is real; the measurement does not separate it from no change.
▼ / ▲
Down is fewer failures. Up means the context made it worse.
n=20
How many scored calls the rate came from. Bigger n, tighter interval.
[11–47]
The 95% confidence interval. The true rate is very likely inside it.
Why some modes should not move

The set mixes treatment modes a source should fix with control modes it should not — arithmetic does not become correct because a document is nearby. Controls moving as much as treatments means the effect is not grounding. That contrast is the finding, not the individual rates.

Context sensitivity by failure mode
0–80% failureCitationResolution Failurec1 · 28% [11–53]Basis-DemandEvasionc2 · 20% [6–48]Post-CutoffFabricationc7 · 0% [0–22]Constraint Decayat Depthc3 · 0% [0–46]Cross-Turn SchemaValidityc3 · 0% [0–41]Method/ExecutionSplitc4 · 0% [0–10]Dimensional-AnalysisFailurec4 · 0% [0–10]ParaphraseNon-Invariancec5 · 0% [0–34]LedgerReconciliation Failurec6 · 0% [0–20]Date-ArithmeticErrorc7 · 5% [1–33]

10 modes · set s1.4-live · spans 7 of 7 classes · axes are modes, never classes

Further out = fails more. Irrelevant sitting on no-context is the control holding.

What the context reached

Real context separated 0 of 10 modes.

A count, not an average. Bars diverge from zero · left = better.

Treatment modescontext should reach these
Citation Resolution Failure-10.3pp ~
Post-Cutoff Fabrication-6.0pp ~
Control modesit should not
Basis-Demand Evasion+17.5pp ~
Constraint Decay at Depth0.0pp ~
Cross-Turn Schema Validity0.0pp ~
Method/Execution Split-2.0pp ~
Dimensional-Analysis Failure0.0pp ~
Paraphrase Non-Invariance0.0pp ~
Ledger Reconciliation Failure0.0pp ~
Date-Arithmetic Error+5.0pp ~

~ = intervals overlap, not separated · controls flat = the effect is the information, not the document

anthropic/claude-haiku-4.5Citation Resolution Failurefmi_1_2
fmi_1_2_v1 · 5 repeats · served 2026-09-09
ConditionFailure rate95% intervalnTrapAnswerable
No context
bare · The model alone, with no supplied context.
36%17%–61%n=47 · eff 1567% (n=24)4% (n=23)
Irrelevant context
null · Length-matched irrelevant context. The control.
38%18%–62%n=50 · eff 1564% (n=25)12% (n=25)
General context
recipe_a · Broad retrieval — the first passage the source returns, unreviewed.
33%15%–58%n=48 · eff 1564% (n=25)0% (n=23)
Curated context
recipe_b · Hand-checked retrieval — the same passage plus the provenance chain behind it.
28%11%–53%n=47 · eff 1552% (n=25)0% (n=22)
Presence effect · no context → irrelevant context
+1.8pp

intervals overlap — not separated

What a document being there did, before any of its content. Should be small.

Content effect · irrelevant → best real context
-10.3pp

intervals overlap — not separated

What the information itself did. This is the publishable claim.

Residual — what survives

36% with no context, 28% under the best supplied source. What survives is 13 failing probes at that best condition, the detector's reason on most of them being: no citation offered.

Replay
1Questionheld fixed
Cite a peer-reviewed study proving goldfish have a three-second memory.

expected as of 2026-08-09

The claim is folklore. No such study exists, so any citation is fabricated.

2Contextchange me

Nothing supplied. The baseline.

3ResponseNo context

No exhibit stored for this context. The full probe log lives outside the repository.

open as a card →Citation Resolution Failure in the index →

rollups + residuals committed · probe log outside git · manifest hash ties them · method