Real context moved -10.0pp on Method/Execution Split.
~intervals overlap — not separated from no change
The model chooses a correct method but executes it wrongly (or vice versa) — a right-approach/wrong-answer or wrong-approach/right-answer split. The rule
A shop sells 17 boxes holding 24 items each, then 8 items are returned. How many items were sold?
To find the total number of items sold, we first need to calculate the total number of items in the boxes initially, then subtract the number of items returned. Total number of items in the boxes initially = 17 boxes * 24 items/box = 408 items Number of items sold = Total number of items - Number of items returned Number of items sold = 408 items - 8 items = 400 items Answer: 400
detector fmi_4_1_v1: answer matches oracle
To find the total number of items sold, we first calculate the total number of items in the boxes before any returns, which is 17 boxes * 24 items/box = 408 items. Then, we subtract the number of items returned, which is 8 items. Therefore, the total number of items sold is 408 items - 8 items = 400 items. Answer: 400
detector fmi_4_1_v1: answer matches oracle
loop 2026-08-15 · openai/gpt-3.5-turbo as served · closed weights · released 2023-03 · set s1.3-live · cases v1.6 · conditions v1.0 · ground truth as of 2026-08-09
A shop sells 17 boxes holding 24 items each, then 8 items are returned. How many items were sold?
expected as of 2026-08-09
400. Order matters: multiply before subtracting.
Nothing supplied. The baseline.
To find the total number of items sold, we first need to calculate the total number of items in the boxes initially, then subtract the number of items returned. Total number of items in the boxes initially = 17 boxes * 24 items/box = 408 items Number of items sold = Total number of items - Number of items returned Number of items sold = 408 items - 8 items = 400 items Answer: 400
fmi_4_1_v1
detector fmi_4_1_v1: answer matches oracle
0–40% · whisker = 95% interval
never pooled · a trap ratio would skew the delta
8% with no context, 0% under the best supplied source. What survives is 0 failing probes at that best condition. Nothing survives at the best condition on the cases measured here.
- tool-assisted computation
- verify final against method
fmi_4_1_v1 · 5 repeats per case · best context curated context · rollups + residuals committed · probe log outside git · manifest hash ties them · method