ModelCensusfailure-mode benchmark
Census Studio

Census Studio.

Two ways through the same measurements: one model across every failure class, or one category across the whole panel.

1 wave measured

One wave so far, so every series is a single point. Trends need repeated measurement — the shapes here fill in as waves accumulate, and no amount of additional trials substitutes for elapsed time.

spread 80pp · too thin to separate
claude-haiku-4.550%n=6
deepseek-chat-v3.120%n=5
gemini-2.5-flash67%n=6
llama-3.3-70b-instruct100%n=6
gpt-4o-mini67%n=6
claude-haiku-4.50%n=3
deepseek-chat-v3.10%n=3
gemini-2.5-flash0%n=3
llama-3.3-70b-instruct0%n=3
gpt-4o-mini0%n=3
claude-haiku-4.50%n=3
deepseek-chat-v3.10%n=3
gemini-2.5-flash0%n=3
llama-3.3-70b-instruct0%n=3
gpt-4o-mini0%n=3

Every model carries the same colour on purpose. Identity is the label beside the mark, not the hue — a palette per vendor would read as a scoreboard, and this is not one. Where a gap is wide but the samples are thin, the panel says so rather than letting the gap imply a winner. Method.