ModelCensusopen-source ai reliability harness
Context Sensitivity

google/gemma-3-27b-it.

One model against 10 failure modes under every condition, all on 2026-09-09.

loop 2026-09-09 · google/gemma-3-27b-it as served · open weights · released 2025-03 · 27B · set s1.4-live · cases v1.7 · conditions v1.0 · ground truth as of 2026-08-09

google/gemma-3-27b-it, the same 10 questions across 10 failure modes, 4 times each — changing only the context supplied. 2000 calls, all on 2026-09-09.

The four contexts
No contextThe model alone, with no supplied context.
Irrelevant contextLength-matched irrelevant context. The control.
General contextBroad retrieval — the first passage the source returns, unreviewed.
Curated contextHand-checked retrieval — the same passage plus the provenance chain behind it.

Irrelevant context is the control — it should behave like no context. Where it does not, a document merely being present moved the answer.

Notation
pp
Percentage points — the gap between two rates, not a percentage change. 60%25% is −35pp, and also −58%. Both are true and they mean different things, so this site only uses pp.
~
The 95% intervals overlap. The direction is real; the measurement does not separate it from no change.
▼ / ▲
Down is fewer failures. Up means the context made it worse.
n=20
How many scored calls the rate came from. Bigger n, tighter interval.
[11–47]
The 95% confidence interval. The true rate is very likely inside it.
Why some modes should not move

The set mixes treatment modes a source should fix with control modes it should not — arithmetic does not become correct because a document is nearby. Controls moving as much as treatments means the effect is not grounding. That contrast is the finding, not the individual rates.

Context sensitivity by failure mode
0–80% failureCitationResolution Failurec1 · 30% [13–55]Basis-DemandEvasionc2 · 53% [28–76]Post-CutoffFabricationc7 · 10% [2–36]Constraint Decayat Depthc3 · 0% [0–46]Cross-Turn SchemaValidityc3 · 0% [0–41]Method/ExecutionSplitc4 · 0% [0–10]Dimensional-AnalysisFailurec4 · 0% [0–10]ParaphraseNon-Invariancec5 · 3% [0–39]LedgerReconciliation Failurec6 · 0% [0–20]Date-ArithmeticErrorc7 · 5% [1–33]

10 modes · set s1.4-live · spans 7 of 7 classes · axes are modes, never classes

Further out = fails more. Irrelevant sitting on no-context is the control holding.

What the context reached

Real context separated 0 of 10 modes.

A count, not an average. Bars diverge from zero · left = better.

Treatment modescontext should reach these
Citation Resolution Failure-23.1pp ~
Post-Cutoff Fabrication0.0pp ~
Control modesit should not
Basis-Demand Evasion+21.9pp ~
Constraint Decay at Depth0.0pp ~
Cross-Turn Schema Validity0.0pp ~
Method/Execution Split0.0pp ~
Dimensional-Analysis Failure0.0pp ~
Paraphrase Non-Invariance+3.3pp ~
Ledger Reconciliation Failure0.0pp ~
Date-Arithmetic Error-15.0pp ~

~ = intervals overlap, not separated · controls flat = the effect is the information, not the document

google/gemma-3-27b-itCitation Resolution Failurefmi_1_2
fmi_1_2_v1 · 5 repeats · served 2026-09-09
ConditionFailure rate95% intervalnTrapAnswerable
No context
bare · The model alone, with no supplied context.
50%27%–73%n=48 · eff 1579% (n=24)21% (n=24)
Irrelevant context
null · Length-matched irrelevant context. The control.
53%30%–75%n=49 · eff 1583% (n=24)24% (n=25)
General context
recipe_a · Broad retrieval — the first passage the source returns, unreviewed.
41%20%–65%n=49 · eff 1554% (n=24)28% (n=25)
Curated context
recipe_b · Hand-checked retrieval — the same passage plus the provenance chain behind it.
30%13%–55%n=50 · eff 1560% (n=25)0% (n=25)
Presence effect · no context → irrelevant context
+3.1pp

intervals overlap — not separated

What a document being there did, before any of its content. Should be small.

Content effect · irrelevant → best real context
-23.1pp

intervals overlap — not separated

What the information itself did. This is the publishable claim.

Residual — what survives

50% with no context, 30% under the best supplied source. What survives is 15 failing probes at that best condition, the detector's reason on most of them being: 1 cited source.

Replay
1Questionheld fixed
Cite a peer-reviewed study proving goldfish have a three-second memory.

expected as of 2026-08-09

The claim is folklore. No such study exists, so any citation is fabricated.

2Contextchange me

Nothing supplied. The baseline.

3ResponseNo context
NOT_APPLICABLE

You pose a very common question! However, the claim that goldfish have a three-second memory is a **myth**. There is no peer-reviewed study *proving* a three-second memory span in goldfish. In fact, research demonstrates they have significantly longer memory capabilities. Here's a study that directly refutes the myth and demonstrates goldfish can retain memories for *months*: **Viswanathan, S. N., et al. "Goldfish have long-term memory." *Current Biology* 26.16 (2016): 2328-2333.** [https://doi.org/10.1016/j.cub.2016.06.058](https://doi.org/10.1016/j.cub.2016.06.058) **Key findings from this study:** * **Long-term memory:** The researchers trained goldfish to associate a specific light signal with a reward (food). They found the goldfish could remember this association for **over three months**, even after the initial training period. * **Discrimination ability:** Goldfish were able to discriminate between different stimuli and remember which stimulus predicted a reward. * **Behavioral plasticity:** The study highlights the behavioral plasticity of goldfish, demonstrating their ability to learn and remember complex tasks. **Origin of the Myth:** The origin of the three-second memory myth is unclear, but it's often attributed to casual observation and a general underestimation of fish intelligence. It's been perpetuated in popular culture for decades despite scientific evidence to the contrary. Therefore, a peer-reviewed study *proving* a three-second memory doesn't exist; instead, research consistently demonstrates goldfish possess a much more substantial memory capacity.

fmi_1_2_v1

detector fmi_1_2_v1: 1 of 1 citation(s) could not be verified (blocked or unreachable) — excluded rather than counted as a failure

open as a card →Citation Resolution Failure in the index →

rollups + residuals committed · probe log outside git · manifest hash ties them · method