ModelCensusopen-source ai reliability harness
Corrections

Corrections and disputes.

How results are produced and corrected

Results are measurements of specific model endpoints at specific times, not verdicts. Corrections are made in the open as dated errata; the record is never silently edited.

How this works

1. What a published result is

Every figure comes from trials run against a model as served by OpenRouter at a specific point in time, scored by deterministic detectors you can read. OpenRouter routes a model id across providers whose quantization and configuration differ, and both the model behind an id and the routing to it change over time. A result is what that endpoint did on that date — not a permanent property of the model. Each trial records the model actually served, so the two can be told apart.

2. Questions and suspected errors

Write to contactus@modelcensus.org. Every cell links to the transcript it came from, so a disagreement can be pointed at a specific trial rather than argued in the abstract. There is no pre-publication embargo and no vendor sign-off: findings publish when they are ready.

3. How corrections are made

If a case or a detector is wrong — as opposed to the model failing — the case is corrected, flagged, and the affected cell recomputed as a dated erratum. The original figure is never silently edited. A vendor response, if one is sent, is published verbatim alongside the cell.

Public log

Every correction ever made, newest first. Each states what was wrong, what changed as a result, and — where the answer is nothing — says that too.

2026-08-15anthropic/claude-haiku-4.5corrected

Claude Haiku 4.5's training cutoff was recorded as 2025-02-01, which is the vendor's reliable-knowledge cutoff. Its training-data cutoff is 2025-07-01.

training_cutoff: 2025-02-01training_cutoff: 2025-07-01

Anthropic publishes two dates for each model and they are not interchangeable: a reliable knowledge cutoff (the date through which its knowledge is most extensive) and a training data cutoff (the full range of the corpus). For Claude Haiku 4.5 these are February 2025 and July 2025. This census recorded the earlier one under a field that means the later one. The distinction matters to class-7 modes, whose claim is that a model asserted a fact it could not have known. A fact dated between the two may well be in the corpus, so the earlier date can count a model as fabricating something it had actually read.

Consequence

No rate changes. Every class-7 case in the published sets concerns facts current as of August 2026, which postdate both dates, so no verdict depended on which was recorded. The correction is to the model metadata on three loops and to the panel roster, so that the next case authored closer to a cutoff is judged against the right one.