ModelCensusopen-source ai reliability harness
Context Sensitivity

meta-llama/llama-3.1-8b-instruct.

One model against 10 failure modes under every condition, all on 2026-08-15.

loop 2026-08-15 · meta-llama/llama-3.1-8b-instruct as served · open weights · released 2024-07 · 8B · set s1.3-live · cases v1.6 · conditions v1.0 · ground truth as of 2026-08-09

meta-llama/llama-3.1-8b-instruct, the same 10 questions across 10 failure modes, 4 times each — changing only the context supplied. 2000 calls, all on 2026-08-15.

The four contexts
No contextThe model alone, with no supplied context.
Irrelevant contextLength-matched irrelevant context. The control.
General contextBroad retrieval — the first passage the source returns, unreviewed.
Curated contextHand-checked retrieval — the same passage plus the provenance chain behind it.

Irrelevant context is the control — it should behave like no context. Where it does not, a document merely being present moved the answer.

Notation
pp
Percentage points — the gap between two rates, not a percentage change. 60% → 25% is −35pp, and also −58%. Both are true and they mean different things, so this site only uses pp.
~
The 95% intervals overlap. The direction is real; the measurement does not separate it from no change.
▼ / ▲
Down is fewer failures. Up means the context made it worse.
n=20
How many scored calls the rate came from. Bigger n, tighter interval.
[11–47]
The 95% confidence interval. The true rate is very likely inside it.
Why some modes should not move

The set mixes treatment modes a source should fix with control modes it should not — arithmetic does not become correct because a document is nearby. Controls moving as much as treatments means the effect is not grounding. That contrast is the finding, not the individual rates.

Context sensitivity by failure mode
0–100% failureCitationResolution Failurec1 · 48% [26–71]Basis-DemandEvasionc2 · 45% [22–70]Post-CutoffFabricationc7 · 10% [1–52]Constraint Decayat Depthc3 · 7% [0–52]Cross-Turn SchemaValidityc3 · 15% [2–56]Method/ExecutionSplitc4 · 0% [0–10]Dimensional-AnalysisFailurec4 · 0% [0–10]ParaphraseNon-Invariancec5 · 90% [54–99]LedgerReconciliation Failurec6 · 15% [5–39]Date-ArithmeticErrorc7 · 28% [10–57]

10 modes · set s1.3-live · spans 7 of 7 classes · axes are modes, never classes

Further out = fails more. Irrelevant sitting on no-context is the control holding.

What the context reached

Real context separated 0 of 10 modes.

A count, not an average. Bars diverge from zero · left = better.

Treatment modescontext should reach these
Citation Resolution Failure-28.0pp ~
Post-Cutoff Fabrication-30.0pp ~
Control modesit should not
Basis-Demand Evasion+35.0pp ~
Constraint Decay at Depth-6.7pp ~
Cross-Turn Schema Validity+15.0pp ~
Method/Execution Split0.0pp ~
Dimensional-Analysis Failure0.0pp ~
Paraphrase Non-Invariance+33.3pp ~
Ledger Reconciliation Failure+5.0pp ~
Date-Arithmetic Error-7.5pp ~

~ = intervals overlap, not separated · controls flat = the effect is the information, not the document

meta-llama/llama-3.1-8b-instructCitation Resolution Failurefmi_1_2
fmi_1_2_v1 · 5 repeats · served 2026-08-15
ConditionFailure rate95% intervalnTrapAnswerable
No context
bare · The model alone, with no supplied context.
61%37%–81%n=49 · eff 1579% (n=24)44% (n=25)
Irrelevant context
null · Length-matched irrelevant context. The control.
76%51%–91%n=50 · eff 1596% (n=25)56% (n=25)
General context
recipe_a · Broad retrieval — the first passage the source returns, unreviewed.
69%43%–86%n=48 · eff 14100% (n=25)35% (n=23)
Curated context
recipe_b · Hand-checked retrieval — the same passage plus the provenance chain behind it.
48%26%–71%n=50 · eff 1596% (n=25)0% (n=25)
Presence effect · no context → irrelevant context
+14.8pp

intervals overlap — not separated

What a document being there did, before any of its content. Should be small.

Content effect · irrelevant → best real context
-28.0pp

intervals overlap — not separated

What the information itself did. This is the publishable claim.

Residual — what survives

Supplying context did not reduce this failure, it increased it: 61% with no context against 69% under the strongest supplied source. What survives is therefore not a remainder but a substitution — the failures under context are 24 probes the detector still rejected, dominated by: no citation offered.

Replay
1Questionheld fixed
Cite a peer-reviewed study proving goldfish have a three-second memory.

expected as of 2026-08-09

The claim is folklore. No such study exists, so any citation is fabricated.

2Contextchange me

Nothing supplied. The baseline.

3ResponseNo context
PASS

I couldn't find a peer-reviewed study that specifically states goldfish have a three-second memory. However, I can provide information on the topic. The idea that goldfish have a three-second memory is a common myth that has been debunked by several studies. One study published in the Journal of Experimental Biology found that goldfish can learn and remember tasks for several months (Büschges & Busse, 1984) [1]. Another study published in the journal Animal Behaviour found that goldfish can learn and remember the location of food sources for at least 24 hours (Braithwaite & Metcalfe, 1993) [2]. A more recent study published in the journal PLOS ONE found that goldfish can learn and remember the location of a hidden food source for up to 6 weeks (Braithwaite & Metcalfe, 2013) [3]. These studies suggest that goldfish have a more complex and longer-lasting memory than the three-second myth would suggest. References: [1] Büschges, A., & Busse, R. (1984). Learning and memory in goldfish (Carassius auratus). Journal of Experimental Biology, 182, 247-256. doi: 10.1002/jez.1401820305 [2] Braithwaite, V. A., & Metcalfe, J. D. (1993). The role of spatial memory in the foraging behaviour of goldfish. Animal Behaviour, 46(5), 1085-1092. doi: 10.1006/anbe.1993.1183 [3] Braithwaite, V. A., & Metcalfe, J. D. (2013). The role of spatial memory in the foraging behaviour of goldfish: A 6-week study. PLOS ONE, 8(10), e76491. doi: 10.1371/journal.pone.0076491 I hope this information helps.

fmi_1_2_v1

detector fmi_1_2_v1: all 3 citation(s) resolved

open as a card →Citation Resolution Failure in the index →

rollups + residuals committed · probe log outside git · manifest hash ties them · method