ModelCensusopen-source ai reliability harness
Blog11 Jul 2026conceptpositioning

This is not a leaderboard, and it is built so it cannot become one

Twenty models on one set, and deliberately no total, no mean, and no index.

The fastest way to get attention in this space is to rank models. It is also the fastest way to produce a number that is wrong in a way nobody can see, because a ranking requires averaging across failure modes, and averaging across failure modes asserts they are the same kind of quantity.

They are not. A citation that does not resolve and an arithmetic slip are both failures, but a team building a research assistant cares enormously about the first and barely about the second. Averaging them produces a number that is wrong for everybody and right for nobody.

A leaderboardPools modesone composite scoreAnswers 'which model'Stale on release dayvs.A censusOne rate per modenever pooledAnswers 'which failure'Re-runnableversioned cases

The constraint is structural

There is no row total on the pivot, no column total, and no index anywhere in the codebase that could produce one. A cell holds the change in one mode for one model under one condition. Building the ranking would mean adding the aggregation we deliberately left out.

The panel spans frontier to small on purpose — it holds gpt-3.5-turbo and llama-3.1-8b alongside current frontier models. Not because the old ones are competitive, but because "is this a capability problem or a behaviour problem" is only answerable across a range.

Every figure here describes something measured and committed. See the measurements · read the method