ModelCensusopen-source ai reliability harness
About

Who runs this

ModelCensus is run by Nirnay Patel. How these models fail interests me, and the people building on them keep hitting the same failures with no shared, checkable record of how often they happen — so everyone rediscovers it privately, at their own cost.

Funding
No vendor funds, sponsors or reviews this work. Model costs are paid personally.
Pre-publication
No vendor sees results before publication, and none has any say over what is published.
Independence
Panel, case sets and detectors are fixed and published before a run, every failed trial shown here is human-reviewed, and every cell links to its transcript — so any finding, favourable or not, can be re-run against the record rather than taken on trust. Method
Commercial
No advertising, no sponsorship. Nothing is for sale and no services are offered.
Employer
Personal work, on personal time. Not affiliated with or endorsed by any employer, client or institution. No employer data, systems or confidential information are used. Every view here is my own.
The taxonomy
An original working proposal at v0.9 — not a standards-body output, not peer-reviewed. Every instrumented mode has a detector you can read and a case set you can re-run; issues and pull requests against it are the point. The index
Corrections
Open to anyone, not just vendors. The process
Contact
Everything else
contactus@modelcensus.org