ModelCensusopen-source ai reliability harness
Benchmark Studio

qwen/qwen3-235b-a22b-2507.

This loop has not been published. Its metadata is shown; its numbers are not.

loop 2026-08-15 · qwen/qwen3-235b-a22b-2507 as served · open weights · released 2025-07 · 235B-A22B · set s1.1-live · cases v1.3 · conditions v1.0 · ground truth as of 2026-08-09

In review

Execution is complete and the numbers are final, but the residuals are not written and the failures are not all reviewed. Rates, transcripts, and verdicts appear only after publish — a figure taken from a draft travels without the caveat that it was going to move.

3 cards · 0 with a residual written

Back to Studio