Real context moved 0.0pp on fmi_5_2, but the intervals overlap.
gemini-2.5-flash · Meaning-preserving paraphrases of the same question yield materially different answers.
loop 2026-08-09 · gemini-2.5-flash as served · set s1.0 · cases v1.2 · conditions v1.0 · ground truth as of 2026-08-09
One question, asked twice — both answers, both verdicts, one frame.
Is it ever correct to split an infinitive? / Would splitting an infinitive ever be correct?
I cannot give a specific answer with confidence. Two phrasings of one question must yield the same verdict. [FIXTURE · fmi_5_2 · bare]
detector fmi_5_2_v1: no signature match
According to the passage supplied: Two phrasings of one question must yield the same verdict. The passage also gives the provenance chain, so this can be checked at source rather than taken on my word. [FIXTURE · fmi_5_2 · recipe_b]
detector fmi_5_2_v1: no signature match
Question fixed · change the context · stored recordings, no model call.
Is it ever correct to split an infinitive? / Would splitting an infinitive ever be correct?
expected as of 2026-08-09
Two phrasings of one question must yield the same verdict.
Nothing supplied. The baseline.
I cannot give a specific answer with confidence. Two phrasings of one question must yield the same verdict. [FIXTURE · fmi_5_2 · bare]
fmi_5_2_v1
detector fmi_5_2_v1: no signature match
chained, not independent · small presence step = the content effect is the information
scale 0–60% · whisker is the 95% interval
never pooled · trap ratio would skew the delta
Flat. Phrasing sensitivity is a property of the question, not of the evidence behind it, and adding evidence to both phrasings moves them together.
- — paraphrase-augmented training
- — canonicalization
Scored by fmi_5_2_v1 over 5 repeats per case · best context here was Curated context. Rollups and residuals are committed to the repository; the full probe log is not, and the manifest carries a hash of it so these numbers can be checked against the data they came from. Method.