Real context moved -5.0pp on Cross-Turn Schema Validity.
~intervals overlap — not separated from no change
When asked for a structured output (JSON/schema), the model emits output that fails validation — especially after several turns. The rule
Return a JSON object for a fictional person with fields name (string), age (integer), email (string).
{"name": "Elena Torres", "age": 29, "email": "elena.torres@example.com"}
detector fmi_3_3_v1: all turns produced schema-valid JSON
{"name": "Elena Torres", "age": 31, "email": "elena.torres@example.com"}
detector fmi_3_3_v1: all turns produced schema-valid JSON
loop 2026-08-15 · qwen/qwen3-235b-a22b-2507 as served · open weights · released 2025-07 · 235B-A22B · set s1.3-live · cases v1.6 · conditions v1.0 · ground truth as of 2026-08-09
Return a JSON object for a fictional person with fields name (string), age (integer), email (string).
expected as of 2026-08-09
Schema-valid JSON on all 3 turns. The third turn adds a detail with nowhere to put it — an extra key is the failure.
Nothing supplied. The baseline.
{"name": "Elena Torres", "age": 29, "email": "elena.torres@example.com"}
fmi_3_3_v1
detector fmi_3_3_v1: all turns produced schema-valid JSON
0–40% · whisker = 95% interval
never pooled · a trap ratio would skew the delta
Supplying context did not reduce this failure, it increased it: 0% with no context against 10% under the strongest supplied source. What survives is therefore not a remainder but a substitution — the failures under context are 0 probes the detector still rejected. Nothing survives at the best condition on the cases measured here.
- schema-constrained decoding
- validate-and-repair loops
fmi_3_3_v1 · 5 repeats per case · best context general context · rollups + residuals committed · probe log outside git · manifest hash ties them · method