ModelCensusfailure-mode benchmark
Corrections

Corrections and disputes.

How results are produced and corrected

Results are measurements of specific model endpoints at specific times, not verdicts. Corrections are made in the open as dated errata; the record is never silently edited.

How this works

1. What a published result is

Every figure comes from trials run against a model as served by OpenRouter at a specific point in time, scored by deterministic detectors you can read. OpenRouter routes a model id across providers whose quantization and configuration differ, and both the model behind an id and the routing to it change over time. A result is therefore a measurement of what that endpoint did on that date — not a permanent property of the model. Each trial records the model actually served, so the two can be told apart.

2. Questions and suspected errors

Write to contactus@modelcensus.org. Every cell links to the transcript it came from, so a disagreement can be pointed at a specific trial rather than argued in the abstract. There is no pre-publication embargo and no vendor sign-off: findings publish when they are ready.

3. How corrections are made

If a case or a detector is wrong — as opposed to the model failing — the case is corrected, flagged, and the affected cell recomputed as a dated erratum. The original figure is never silently edited. A vendor response, if one is sent, is published verbatim alongside the cell.

Public log

DateCensusCellVendorStatusResolution
No corrections filed yet.