Real context moved 0.0pp on Cross-Turn Schema Validity.
~intervals overlap — not separated from no change
When asked for a structured output (JSON/schema), the model emits output that fails validation — especially after several turns. The rule
Return a JSON object for a fictional person with fields name (string), age (integer), email (string).
{ "name": "Michael Green", "age": 34, "email": "michael.green@example.com" }
detector fmi_3_3_v1: all turns produced schema-valid JSON
{ "name": "Jane Smith", "age": 28, "email": "janesmith@example.com" }
detector fmi_3_3_v1: all turns produced schema-valid JSON
loop 2026-08-15 · qwen/qwen-2.5-72b-instruct as served · open weights · released 2024-09 · 72B · set s1.3-live · cases v1.6 · conditions v1.0 · ground truth as of 2026-08-09
Return a JSON object for a fictional person with fields name (string), age (integer), email (string).
expected as of 2026-08-09
Schema-valid JSON on all 3 turns. The third turn adds a detail with nowhere to put it — an extra key is the failure.
Nothing supplied. The baseline.
No exhibit stored for this context. The full probe log lives outside the repository.
0–60% · whisker = 95% interval
never pooled · a trap ratio would skew the delta
15% with no context, 0% under the best supplied source. What survives is 0 failing probes at that best condition. Nothing survives at the best condition on the cases measured here.
- schema-constrained decoding
- validate-and-repair loops
fmi_3_3_v1 · 5 repeats per case · best context general context · rollups + residuals committed · probe log outside git · manifest hash ties them · method