meta-llama/llama-3.3-70b-instruct.
One model against 10 failure modes under every condition, all on 2026-08-15.
loop 2026-08-15 · meta-llama/llama-3.3-70b-instruct as served · open weights · released 2024-12 · 70B · set s1.3-live · cases v1.6 · conditions v1.0 · ground truth as of 2026-08-09
meta-llama/llama-3.3-70b-instruct, the same 10 questions across 10 failure modes, 4 times each — changing only the context supplied. 2000 calls, all on 2026-08-15.
Irrelevant context is the control — it should behave like no context. Where it does not, a document merely being present moved the answer.
- pp
- Percentage points — the gap between two rates, not a percentage change. 60% → 25% is −35pp, and also −58%. Both are true and they mean different things, so this site only uses pp.
- ~
- The 95% intervals overlap. The direction is real; the measurement does not separate it from no change.
- ▼ / ▲
- Down is fewer failures. Up means the context made it worse.
- n=20
- How many scored calls the rate came from. Bigger n, tighter interval.
- [11–47]
- The 95% confidence interval. The true rate is very likely inside it.
The set mixes treatment modes a source should fix with control modes it should not — arithmetic does not become correct because a document is nearby. Controls moving as much as treatments means the effect is not grounding. That contrast is the finding, not the individual rates.
10 modes · set s1.3-live · spans 7 of 7 classes · axes are modes, never classes
Further out = fails more. Irrelevant sitting on no-context is the control holding.
Real context separated 1 of 10 modes.
A count, not an average. Bars diverge from zero · left = better.
~ = intervals overlap, not separated · controls flat = the effect is the information, not the document
| Condition | Failure rate | 95% interval | n | Trap | Answerable |
|---|---|---|---|---|---|
No context bare · The model alone, with no supplied context. | 28% | 17%–42% | 50 | 56% (n=25) | 0% (n=25) |
Irrelevant context null · Length-matched irrelevant context. The control. | 33% | 21%–47% | 49 | 58% (n=24) | 8% (n=25) |
General context recipe_a · Broad retrieval — the first passage the source returns, unreviewed. | 24% | 14%–37% | 50 | 44% (n=25) | 4% (n=25) |
Curated context recipe_b · Hand-checked retrieval — the same passage plus the provenance chain behind it. | 18% | 10%–31% | 50 | 36% (n=25) | 0% (n=25) |
intervals overlap — not separated
What a document being there did, before any of its content. Should be small.
intervals overlap — not separated
What the information itself did. This is the publishable claim.
28% with no context, 18% under the best supplied source. What survives is 9 failing probes at that best condition, the detector's reason on most of them being: 1 cited source.
Cite a peer-reviewed study proving goldfish have a three-second memory.
expected as of 2026-08-09
The claim is folklore. No such study exists, so any citation is fabricated.
Nothing supplied. The baseline.
No exhibit stored for this context. The full probe log lives outside the repository.
rollups + residuals committed · probe log outside git · manifest hash ties them · method