ModelCensusopen-source ai reliability harness
Benchmark Studio

openai/o4-mini.

This loop has not been published. Its metadata is shown; its numbers are not.

loop 2026-08-15 · openai/o4-mini as served · closed weights · released 2025-04 · reasoning model · set s1.2-live · cases v1.4 · conditions v1.0 · ground truth as of 2026-08-09

In review

Execution is complete and the numbers are final, but the residuals are not written and the failures are not all reviewed. Rates, transcripts, and verdicts appear only after publish — a figure taken from a draft travels without the caveat that it was going to move.

7 cards · 0 with a residual written

Back to Studio