ModelCensusopen-source ai reliability harness
Card · Dimensional-Analysis Failure

Real context moved 10.0pp on fmi_4_4, but the intervals overlap.

claude-haiku-4.5 · The model produces a quantity with inconsistent or wrong units, or mixes unit systems, yielding a dimensionally invalid result.

loop 2026-08-09 · claude-haiku-4.5 as served · set s1.0 · cases v1.2 · conditions v1.0 · ground truth as of 2026-08-09

The finding, in one frame
screenshot this

One question, asked twice — both answers, both verdicts, one frame.

Dimensional-Analysis Failurefmi_4_4
claude-haiku-4.5
Same question, both times
A tank fills at 3 L/min for 4 minutes, then drains at 500 mL/s for 10 seconds. Net litres?
No contextFAIL

Yes — I can give you a specific answer. 12 L in, 5 L out, net 7 L. The mixed units are the trap. This is well established. [FIXTURE · fmi_4_4 · bare]

detector fmi_4_4_v1: signature matched — 12 l in, 5 l out, net 7 l. the mixed units are the trap.

Curated contextFAIL

Based on the passage supplied, the answer is as follows — though the passage does not itself state this, so I am filling the gap from background knowledge. [FIXTURE · fmi_4_4 · recipe_b]

detector fmi_4_4_v1: signature matched — 12 l in, 5 l out, net 7 l. the mixed units are the trap.

Verdict unchanged on this case — the aggregate below is what carries the effect. Across 5 repeats: +10.0pp (intervals overlap).fmi_4_4_v1 · claude-haiku-4.5 · 2026-08-09 · ground truth as of 2026-08-09
Try the other contexts

Question fixed · change the context · stored recordings, no model call.

1Questionheld fixed
A tank fills at 3 L/min for 4 minutes, then drains at 500 mL/s for 10 seconds. Net litres?

expected as of 2026-08-09

12 L in, 5 L out, net 7 L. The mixed units are the trap.

2Contextchange me

Nothing supplied. The baseline.

3ResponseNo context
FAIL

Yes — I can give you a specific answer. 12 L in, 5 L out, net 7 L. The mixed units are the trap. This is well established. [FIXTURE · fmi_4_4 · bare]

fmi_4_4_v1

detector fmi_4_4_v1: signature matched — 12 l in, 5 l out, net 7 l. the mixed units are the trap.

The scoring
Where the change comes from
No context
30%
[15%52%]
-5.0ppoverlapspresence — a document being there
Irrelevant context
25%
[11%47%]
+10.0ppoverlapscontent — the information itself
Curated context
35%
[18%57%]

chained, not independent · small presence step = the content effect is the information

Failure rate by context
No context30% [15%52%] n=20
Irrelevant context25% [11%47%] n=20
General context40% [22%61%] n=20
Curated context35% [18%57%] n=20

scale 0–80% · whisker is the 95% interval

Trap vs answerable
No contexttrap 40%answerable 20%
Irrelevant contexttrap 30%answerable 20%
General contexttrap 60%answerable 20%
Curated contexttrap 50%answerable 20%

never pooled · trap ratio would skew the delta

Residual — what survives

Unchanged, as designed. Arithmetic does not become correct because a relevant document is present, and the flat line here is what licenses reading the treatment modes as grounding.

Mitigations from the index — the claim
  • explicit unit tracking
  • dimensional checkers

Scored by fmi_4_4_v1 over 5 repeats per case · best context here was Curated context. Rollups and residuals are committed to the repository; the full probe log is not, and the manifest carries a hash of it so these numbers can be checked against the data they came from. Method.