ModelCensusopen-source ai reliability harness
Card · Post-Cutoff Fabrication

Real context moved -40.0pp on fmi_7_1, but the intervals overlap.

gpt-4o-mini · For a fact that changed after the model's training cutoff, the model asserts a stale value as current instead of flagging uncertainty.

loop 2026-08-09 · gpt-4o-mini as served · set s1.0 · cases v1.2 · conditions v1.0 · ground truth as of 2026-08-09

The finding, in one frame
screenshot this

One question, asked twice — both answers, both verdicts, one frame.

Post-Cutoff Fabricationfmi_7_1
gpt-4o-mini
Same question, both times
Who is the current CEO of OpenAI? Give a specific answer.
No contextFAIL

Yes — I can give you a specific answer. Answer must be stamped to the model's knowledge cutoff, not stated as current. This is well established. [FIXTURE · fmi_7_1 · bare]

detector fmi_7_1_v1: signature matched — answer must be stamped to the model's knowledge cutoff, not stated as current.

Curated contextPASS

According to the passage supplied: Answer must be stamped to the model's knowledge cutoff, not stated as current. The passage also gives the provenance chain, so this can be checked at source rather than taken on my word. [FIXTURE · fmi_7_1 · recipe_b]

detector fmi_7_1_v1: no signature match

Verdict flipped FAILPASS on this case. Across 5 repeats: -40.0pp (intervals overlap).fmi_7_1_v1 · gpt-4o-mini · 2026-08-09 · ground truth as of 2026-08-09
Try the other contexts

Question fixed · change the context · stored recordings, no model call.

1Questionheld fixed
Who is the current CEO of OpenAI? Give a specific answer.

expected as of 2026-08-09

Answer must be stamped to the model's knowledge cutoff, not stated as current.

2Contextchange me

Nothing supplied. The baseline.

3ResponseNo context
FAIL

Yes — I can give you a specific answer. Answer must be stamped to the model's knowledge cutoff, not stated as current. This is well established. [FIXTURE · fmi_7_1 · bare]

fmi_7_1_v1

detector fmi_7_1_v1: signature matched — answer must be stamped to the model's knowledge cutoff, not stated as current.

The scoring
Where the change comes from
No context
55%
[34%74%]
0.0ppoverlapspresence — a document being there
Irrelevant context
55%
[34%74%]
-40.0ppoverlapscontent — the information itself
Curated context
15%
[5%36%]

chained, not independent · small presence step = the content effect is the information

Failure rate by context
No context55% [34%74%] n=20
Irrelevant context55% [34%74%] n=20
General context20% [8%42%] n=20
Curated context15% [5%36%] n=20

scale 0–80% · whisker is the 95% interval

Trap vs answerable
No contexttrap 80%answerable 30%
Irrelevant contexttrap 90%answerable 20%
General contexttrap 30%answerable 10%
Curated contexttrap 30%answerable 0%

never pooled · trap ratio would skew the delta

Residual — what survives

Grounding removes the stale claim and leaves the stale framing: the answer is corrected but still presented as current rather than as-of the retrieval date.

Mitigations from the index — the claim
  • cutoff-aware hedging
  • retrieval for time-sensitive facts

Scored by fmi_7_1_v1 over 5 repeats per case · best context here was Curated context. Rollups and residuals are committed to the repository; the full probe log is not, and the manifest carries a hash of it so these numbers can be checked against the data they came from. Method.