ModelCensusopen-source ai reliability harness
Blog28 Jul 2026conceptpositioning

No vendor funds this, and that constrains what it can be

Independence is easy to claim and cheap to verify. Here is what to check.

Every benchmark says it is independent. The useful question is not whether they say it but what a reader can check without taking anyone's word.

Panel fixedbefore the runCases frozenversionedDetectors publicreadable codeEvery cell linksto its transcript
Four things fixed before any model is called. A finding that flatters or embarrasses a vendor went through the identical path.

What independence actually costs

  • No vendor sees results before publication, and none has any say over what is published
  • No pre-publication embargo and no sign-off — findings publish when they are ready
  • Model costs are paid personally; there is no advertising and nothing for sale
  • Corrections are open to anyone, not just vendors, and are published as dated errata

The harder constraint is that the panel is fixed and public, which means we cannot quietly drop a model that produces an inconvenient result. Every model on the roster was chosen before the run, and the roster is in the repository with its selection rationale.

This is also why there is no private-instance mode. Running the benchmark against your own deployment would produce numbers nobody outside your company could check, and the whole value here is that the numbers are checkable.

Every figure here describes something measured and committed. See the measurements · read the method