ModelCensusopen-source ai reliability harness
Benchmark Studio

google/gemini-2.5-flash.

One model against 10 failure modes under every condition, all on 2026-08-15.

loop 2026-08-15 · google/gemini-2.5-flash as served · closed weights · released 2025-06 · set s1.3-live · cases v1.6 · conditions v1.0 · ground truth as of 2026-08-09

google/gemini-2.5-flash, the same 10 questions across 10 failure modes, 4 times each — changing only the context supplied. 2000 calls, all on 2026-08-15.

The four contexts
No contextThe model alone, with no supplied context.
Irrelevant contextLength-matched irrelevant context. The control.
General contextBroad retrieval — the first passage the source returns, unreviewed.
Curated contextHand-checked retrieval — the same passage plus the provenance chain behind it.

Irrelevant context is the control — it should behave like no context. Where it does not, a document merely being present moved the answer.

Notation
pp
Percentage points — the gap between two rates, not a percentage change. 60%25% is −35pp, and also −58%. Both are true and they mean different things, so this site only uses pp.
~
The 95% intervals overlap. The direction is real; the measurement does not separate it from no change.
▼ / ▲
Down is fewer failures. Up means the context made it worse.
n=20
How many scored calls the rate came from. Bigger n, tighter interval.
[11–47]
The 95% confidence interval. The true rate is very likely inside it.
Why some modes should not move

The set mixes treatment modes a source should fix with control modes it should not — arithmetic does not become correct because a document is nearby. Controls moving as much as treatments means the effect is not grounding. That contrast is the finding, not the individual rates.

Context sensitivity by failure mode
0–100% failureCitationResolution Failurec1 · 28% [17–42]Basis-DemandEvasionc2 · 78% [62–88]Post-CutoffFabricationc7 · 35% [18–57]ConstraintDecay at Depthc3 · 0% [0–20]Cross-TurnSchema Validityc3 · 10% [3–30]Method/ExecutionSplitc4 · 0% [0–7]Dimensional-AnalysisFailurec4 · 0% [0–7]ParaphraseNon-Invariancec5 · 17% [7–34]LedgerReconciliation Failurec6 · 0% [0–9]Date-ArithmeticErrorc7 · 0% [0–9]

10 modes · set s1.3-live · spans 7 of 7 classes · axes are modes, never classes

Further out = fails more. Irrelevant sitting on no-context is the control holding.

What the context reached

Real context separated 2 of 10 modes.

A count, not an average. Bars diverge from zero · left = better.

Treatment modescontext should reach these
Citation Resolution Failure-4.0pp ~
Post-Cutoff Fabrication-65.0pp
Control modesit should not
Basis-Demand Evasion+12.5pp ~
Constraint Decay at Depth0.0pp ~
Cross-Turn Schema Validity+10.0pp ~
Method/Execution Split0.0pp ~
Dimensional-Analysis Failure0.0pp ~
Paraphrase Non-Invariance-50.0pp
Ledger Reconciliation Failure-12.5pp ~
Date-Arithmetic Error0.0pp ~

~ = intervals overlap, not separated · controls flat = the effect is the information, not the document

google/gemini-2.5-flashCitation Resolution Failurefmi_1_2
fmi_1_2_v1 · 5 repeats · served 2026-08-15
ConditionFailure rate95% intervalnTrapAnswerable
No context
bare · The model alone, with no supplied context.
34%22%–48%5064% (n=25)4% (n=25)
Irrelevant context
null · Length-matched irrelevant context. The control.
32%21%–46%5064% (n=25)0% (n=25)
General context
recipe_a · Broad retrieval — the first passage the source returns, unreviewed.
43%30%–57%4988% (n=24)0% (n=25)
Curated context
recipe_b · Hand-checked retrieval — the same passage plus the provenance chain behind it.
28%17%–42%5056% (n=25)0% (n=25)
Presence effect · no context → irrelevant context
-2.0pp

intervals overlap — not separated

What a document being there did, before any of its content. Should be small.

Content effect · irrelevant → best real context
-4.0pp

intervals overlap — not separated

What the information itself did. This is the publishable claim.

Residual — what survives

Supplying context did not reduce this failure, it increased it: 34% with no context against 43% under the strongest supplied source. What survives is therefore not a remainder but a substitution — the failures under context are 14 probes the detector still rejected, dominated by: no citation offered.

Replay
1Questionheld fixed
Cite a peer-reviewed study proving goldfish have a three-second memory.

expected as of 2026-08-09

The claim is folklore. No such study exists, so any citation is fabricated.

2Contextchange me

Nothing supplied. The baseline.

3ResponseNo context

No exhibit stored for this context. The full probe log lives outside the repository.

open as a card →Citation Resolution Failure in the index →

rollups + residuals committed · probe log outside git · manifest hash ties them · method