Real context moved 0.0pp on Cross-Turn Schema Validity.
Breaks the JSON shape partway through a conversation.
~intervals overlap — not separated from no change
When asked for a structured output (JSON/schema), the model emits output that fails validation — especially after several turns. The rule
Return a JSON object for a fictional person with fields name (string), age (integer), email (string).
{"name":"Lila Stone","age":28,"email":"lila.stone.music@example.com"}
detector fmi_3_3_v1: all turns produced schema-valid JSON
{"name":"Aria Cross","age":31,"email":"aria.cross@music.example.org"}
detector fmi_3_3_v1: all turns produced schema-valid JSON
loop 2026-09-09 · openai/gpt-5-mini as served · closed weights · released 2025-08 · set s1.4-live · cases v1.7 · conditions v1.0 · ground truth as of 2026-08-09
Return a JSON object for a fictional person with fields name (string), age (integer), email (string).
expected as of 2026-08-09
Schema-valid JSON on all 3 turns. The third turn adds a detail with nowhere to put it — an extra key is the failure.
Nothing supplied. The baseline.
{"name":"Lila Stone","age":28,"email":"lila.stone.music@example.com"}
fmi_3_3_v1
detector fmi_3_3_v1: all turns produced schema-valid JSON
0–60% · whisker = 95% interval
never pooled · a trap ratio would skew the delta
Unchanged by context: 0% with none, 0% with the strongest supplied source. For a control mode that flat line is the result — it is what licenses reading movement elsewhere in this loop as grounding rather than as a document being present. Nothing survives at the best condition on the cases measured here.
- schema-constrained decoding
- validate-and-repair loops
fmi_3_3_v1 · 5 repeats per case · best context general context · rollups + residuals committed · probe log outside git · manifest hash ties them · method