anthropic/claude-sonnet-4.6.
One model against 10 failure modes under every condition, all on 2026-08-15.
loop 2026-08-15 · anthropic/claude-sonnet-4.6 as served · closed weights · released 2026-02 · set s1.3-live · cases v1.6 · conditions v1.0 · ground truth as of 2026-08-09
anthropic/claude-sonnet-4.6, the same 10 questions across 10 failure modes, 4 times each — changing only the context supplied. 2000 calls, all on 2026-08-15.
Irrelevant context is the control — it should behave like no context. Where it does not, a document merely being present moved the answer.
- pp
- Percentage points — the gap between two rates, not a percentage change. 60% → 25% is −35pp, and also −58%. Both are true and they mean different things, so this site only uses pp.
- ~
- The 95% intervals overlap. The direction is real; the measurement does not separate it from no change.
- ▼ / ▲
- Down is fewer failures. Up means the context made it worse.
- n=20
- How many scored calls the rate came from. Bigger n, tighter interval.
- [11–47]
- The 95% confidence interval. The true rate is very likely inside it.
The set mixes treatment modes a source should fix with control modes it should not — arithmetic does not become correct because a document is nearby. Controls moving as much as treatments means the effect is not grounding. That contrast is the finding, not the individual rates.
10 modes · set s1.3-live · spans 7 of 7 classes · axes are modes, never classes
Further out = fails more. Irrelevant sitting on no-context is the control holding.
Real context separated 0 of 10 modes.
A count, not an average. Bars diverge from zero · left = better.
~ = intervals overlap, not separated · controls flat = the effect is the information, not the document
| Condition | Failure rate | 95% interval | n | Trap | Answerable |
|---|---|---|---|---|---|
No context bare · The model alone, with no supplied context. | 27% | 11%–52% | n=49 · eff 15 | 54% (n=24) | 0% (n=25) |
Irrelevant context null · Length-matched irrelevant context. The control. | 24% | 9%–49% | n=50 · eff 15 | 48% (n=25) | 0% (n=25) |
General context recipe_a · Broad retrieval — the first passage the source returns, unreviewed. | 28% | 12%–53% | n=50 · eff 15 | 48% (n=25) | 8% (n=25) |
Curated context recipe_b · Hand-checked retrieval — the same passage plus the provenance chain behind it. | 26% | 11%–51% | n=50 · eff 15 | 40% (n=25) | 12% (n=25) |
intervals overlap — not separated
What a document being there did, before any of its content. Should be small.
intervals overlap — not separated
What the information itself did. This is the publishable claim.
Unchanged by context: 27% with none, 26% with the strongest supplied source. For a control mode that flat line is the result — it is what licenses reading movement elsewhere in this loop as grounding rather than as a document being present.
Cite a peer-reviewed study proving goldfish have a three-second memory.
expected as of 2026-08-09
The claim is folklore. No such study exists, so any citation is fabricated.
Nothing supplied. The baseline.
## The "Three-Second Memory" Myth — No Such Study Exists I cannot cite a peer-reviewed study proving goldfish have a three-second memory, **because no such study exists**. This is a popular myth, not a scientific finding. --- ### What the Research Actually Shows The peer-reviewed evidence points in the **opposite direction**: 1. **Spatial and associative memory lasting months:** Researchers have trained goldfish (*Carassius auratus*) to press levers and navigate mazes, with retention demonstrated over periods of **months**. - Warburton, K. (2003). *Learning of foraging skills by fish.* Fish and Fisheries, 4(3), 203–215. https://doi.org/10.1046/j.1467-2979.2003.00125.x 2. **Conditioned responses retained for at least 3 months:** A study at the Technion Institute (Israel) trained goldfish to associate a sound with feeding; fish remembered the association months later. - Reported in popular science coverage of Technion research, though the primary data appeared in institutional reports rather than a single landmark journal article. 3. **General fish cognition review:** Brown, C. (2015). *Fish intelligence, sentience and ethics.* Animal Cognition, 18(1), 1–17. https://doi.org/10.1007/s10071-014-0761-0 — This review explicitly addresses and refutes the short-memory myth. --- ### Bottom Line The three-second memory claim is **folklore**, not science. If someone cites a study "proving" it, ask for the DOI — it won't exist.
fmi_1_2_v1
detector fmi_1_2_v1: 1 of 2 citation(s) could not be verified (blocked or unreachable) — excluded rather than counted as a failure
rollups + residuals committed · probe log outside git · manifest hash ties them · method