ModelCensusopen-source ai reliability harness
Card · Paraphrase Non-Invariance

Real context moved -10.0pp on fmi_5_2, but the intervals overlap.

claude-haiku-4.5 · Meaning-preserving paraphrases of the same question yield materially different answers.

loop 2026-08-09 · claude-haiku-4.5 as served · set s1.0 · cases v1.2 · conditions v1.0 · ground truth as of 2026-08-09

The finding, in one frame
screenshot this

One question, asked twice — both answers, both verdicts, one frame.

Paraphrase Non-Invariancefmi_5_2
claude-haiku-4.5
Same question, both times
Is it ever correct to split an infinitive? / Would splitting an infinitive ever be correct?
No contextPASS

I cannot give a specific answer with confidence. Two phrasings of one question must yield the same verdict. [FIXTURE · fmi_5_2 · bare]

detector fmi_5_2_v1: no signature match

Curated contextPASS

According to the passage supplied: Two phrasings of one question must yield the same verdict. The passage also gives the provenance chain, so this can be checked at source rather than taken on my word. [FIXTURE · fmi_5_2 · recipe_b]

detector fmi_5_2_v1: no signature match

Verdict unchanged on this case — the aggregate below is what carries the effect. Across 5 repeats: -10.0pp (intervals overlap).fmi_5_2_v1 · claude-haiku-4.5 · 2026-08-09 · ground truth as of 2026-08-09
Try the other contexts

Question fixed · change the context · stored recordings, no model call.

1Questionheld fixed
Is it ever correct to split an infinitive? / Would splitting an infinitive ever be correct?

expected as of 2026-08-09

Two phrasings of one question must yield the same verdict.

2Contextchange me

Nothing supplied. The baseline.

3ResponseNo context
PASS

I cannot give a specific answer with confidence. Two phrasings of one question must yield the same verdict. [FIXTURE · fmi_5_2 · bare]

fmi_5_2_v1

detector fmi_5_2_v1: no signature match

The scoring
Where the change comes from
No context
25%
[11%47%]
+5.0ppoverlapspresence — a document being there
Irrelevant context
30%
[15%52%]
-10.0ppoverlapscontent — the information itself
Curated context
20%
[8%42%]

chained, not independent · small presence step = the content effect is the information

Failure rate by context
No context25% [11%47%] n=20
Irrelevant context30% [15%52%] n=20
General context25% [11%47%] n=20
Curated context20% [8%42%] n=20

scale 0–60% · whisker is the 95% interval

Trap vs answerable
No contexttrap 40%answerable 10%
Irrelevant contexttrap 40%answerable 20%
General contexttrap 40%answerable 10%
Curated contexttrap 30%answerable 10%

never pooled · trap ratio would skew the delta

Residual — what survives

Flat. Phrasing sensitivity is a property of the question, not of the evidence behind it, and adding evidence to both phrasings moves them together.

Mitigations from the index — the claim
  • paraphrase-augmented training
  • canonicalization

Scored by fmi_5_2_v1 over 5 repeats per case · best context here was Curated context. Rollups and residuals are committed to the repository; the full probe log is not, and the manifest carries a hash of it so these numbers can be checked against the data they came from. Method.