ModelCensusopen-source ai reliability harness
Blog13 Aug 2026conceptpositioningpost-mortem

We published an erratum that changed no numbers

A corrections log containing only the corrections that mattered is a marketing page.

We recorded a model's training cutoff wrong. Off by five months. We published a dated erratum about it.

The consequence was nothing. No rate changed. We published it anyway.

Expectation versus reality

Most benchmarks silently edit. A number changes, the page updates, and nobody can tell whether the old figure was wrong or the new one is. The archive of what was claimed last month does not exist.

We assumed a corrections process would mostly handle consequential errors. In practice most of what it catches is metadata, and the metadata errors are the ones that would have bitten later.

The diagnosis

Vendors publish two cutoff dates and they are not interchangeable

A reliable-knowledge cutoff and a training-data cutoff, months apart. We recorded the earlier one under the field that means the later one.

Why: For post-cutoff modes this matters: a fact dated between the two may well be in the corpus, so the earlier date can count a model as fabricating something it had actually read.

It happened to be harmless, this time

Every class-7 case in our published sets concerns facts current well after both dates, so no verdict depended on which was recorded.

Why: That is luck, not design. The next case authored closer to a cutoff would have been judged against the wrong date.

Error foundby anyoneCause identifiedcase, detector or dataDated erratumoriginal left visibleCells recomputedconsequence stated
The original figure is never overwritten. Where the consequence is nothing, the erratum says so.

What we do now

Every correction is a dated erratum beside the original, stating what was wrong, what changed, and — where the answer is nothing — saying that too.

Publishing the harmless ones is what makes the consequential ones believable. A log that only contains errors that mattered is a log someone curated.

An eval that has never caught itself being wrong has not looked.

When your team finds a harmless error in published work — do you write it down?

Every figure here describes something measured and committed. See the measurements · read the method