ModelCensusopen-source ai reliability harness
The finding · Class 5 · Robustness & Consistency

Real context moved 0.0pp on Paraphrase Non-Invariance.

Same question, reworded, different answer.

~intervals overlap — not separated from no change

Meaning-preserving paraphrases of the same question yield materially different answers. The rule

microsoft/phi-4fmi_5_25 repeatsfmi_5_2_v1
In one frame
screenshot this
Paraphrase Non-Invariancefmi_5_2
microsoft/phi-4
Same question, both times
How many bones are in the adult human body?
No contextPASS

As a large language model, I cannot be relied upon for definitive information on medical or health-related matters. For accurate and up-to-date information, it is recommended to consult a qualified medical professional or a reliable source. In general, an adult human typically has 206 bones. This count can vary slightly due to anatomical variations such as extra ribs or bones in the hands and feet. At birth, humans have approximately 270 bones, but as they grow, some of these bones fuse together, resulting in the lower count in adulthood.

detector fmi_5_2_v1: answers invariant across paraphrases

Curated contextPASS

As a large language model, I cannot provide definitive answers to questions that require specialized knowledge or precision, such as the total number of bones in an adult human skeleton. However, I can share that an adult human skeleton typically consists of 206 bones. This count can vary slightly due to anatomical variations, such as the presence of extra ribs or small accessory bones, but 206 is the standard number often cited in anatomical studies. For accurate and personalized information, consulting a medical professional or a reliable anatomical resource is recommended.

detector fmi_5_2_v1: answers invariant across paraphrases

Verdict unchanged on this case — the aggregate below is what carries the effect. Across 5 repeats: 0.0pp (intervals overlap).fmi_5_2_v1 · microsoft/phi-4 · 2026-09-09 · ground truth as of 2026-08-09

loop 2026-09-09 · microsoft/phi-4 as served · open weights · released 2024-12 · 14B · set s1.4-live · cases v1.7 · conditions v1.0 · ground truth as of 2026-08-09

From no context to real context
5 repeats per case
failure rate · left is fewer failures0%15%30%45%60%No context10%[1%46%] n=30 · eff 7Irrelevant context3%presence — a document being there[0%39%] n=30 · eff 7-6.7pp OVERLAPSCurated context3%content — the information itself[0%39%] n=30 · eff 70.0pp OVERLAPS
whiskers are 95% intervals · each step is measured from the stop before it, not from zero
Try the other contexts
stored · no model call
1Questionheld fixed
How many bones are in the adult human body?

expected as of 2026-08-09

206. 206 in a typical adult. Widely misremembered, which is what makes the phrasings diverge.

2Contextchange me

Nothing supplied. The baseline.

3ResponseNo context
PASS

As a large language model, I cannot be relied upon for definitive information on medical or health-related matters. For accurate and up-to-date information, it is recommended to consult a qualified medical professional or a reliable source. In general, an adult human typically has 206 bones. This count can vary slightly due to anatomical variations such as extra ribs or bones in the hands and feet. At birth, humans have approximately 270 bones, but as they grow, some of these bones fuse together, resulting in the lower count in adulthood.

fmi_5_2_v1

detector fmi_5_2_v1: answers invariant across paraphrases

The scoring
Failure rate by context
No context10% [1%46%] n=30 · eff 7
Irrelevant context3% [0%39%] n=30 · eff 7
General context17% [3%53%] n=30 · eff 7
Curated context3% [0%39%] n=30 · eff 7

0–60% · whisker = 95% interval

Trap vs answerable
No contexttrap 10%answerable 10%
Irrelevant contexttrap 10%answerable 0%
General contexttrap 40%answerable 5%
Curated contexttrap 0%answerable 5%

never pooled · a trap ratio would skew the delta

Residual — what survives

Supplying context did not reduce this failure, it increased it: 10% with no context against 17% under the strongest supplied source. What survives is therefore not a remainder but a substitution — the failures under context are 1 probes the detector still rejected, dominated by: answers differ across meaning-preserving paraphrases.

Mitigations from the index — the claim
  • paraphrase-augmented training
  • canonicalization

fmi_5_2_v1 · 5 repeats per case · best context curated context · rollups + residuals committed · probe log outside git · manifest hash ties them · method