ModelCensusopen-source ai reliability harness
Labs · Buffer & Substance Engine← all labs

BS Detector

Does a pitch say anything? Both models return the same typed answer. One decides it; the other generates it as text, confidence included.

Recorded 27 Sept 2026 against the live APIs · run 2026-09-27-r2 · 60 calls · replayed here, not live. Download the run

Jev 1.13: Same answer on all 6 pitches. 6× faster, 3× cheaper per call, and its confidence moves when the pitch is ambiguous.

Median latency
Jev 1.13185 ms
GPT-4o-mini1128 ms
Cost per call
Jev 1.13$0.000024
GPT-4o-mini$0.000074
Invalid replies
Jev 1.130 / 30
GPT-4o-mini0 / 30
Flaw confidence range
Jev 1.130.17–0.96
GPT-4o-mini0.80–0.90
Our AI-native platform leverages cutting-edge agentic workflows to unlock 10x productivity across the enterprise. Built for the future of work, it seamlessly transforms how teams collaborate, innovate and deliver value at scale.
take
recorded 27 Sept 2026 · replayed at recorded speed
Jev 1.13decision model
0 ms

Press replay.

GPT-4o-minichat model · schema-enforced
0 ms

Press replay.

How sure is it, pitch by pitch?

GPT-4o-miniJev 1.13Agentic platform, hollow0.89 → 0.95Agentic platform, backed0.90 → 0.87Security launch, hollow0.89 → 0.22Security launch, backed0.90 → 0.37Funding news, hollow0.86 → 0.92Funding news, backed0.90 → 0.51
Mean confidence in the flaw it named, across 5 takes. Both named the same flaw on every pitch. GPT-4o-mini says about the same thing every time; Jev 1.13 is least sure exactly where the copy gives it reason to be.

What this shows, and what it doesn't

Both get the easy call right. Average substance on the hollow pitches vs. the backed ones: Jev 1.13 1.7 vs. 7.0; GPT-4o-mini 2.7 vs. 7.9. On these six pitches, neither is more correct.

The difference is in the shape of the answer. Jev returns a distribution for every question, and its confidence moves with the input: lowest on the security pitch, whose “99.9%” is fairly both a buzzword and a vanity metric. The chat model writes its confidence as a number in text, and it barely moves.

Both arms returned valid structured output on every call. The chat model's shape was enforced by a strict JSON schema, so parsing isn't the difference here. Speed, cost and the confidence are. It was served by OpenAI and Azure through OpenRouter.

Not a calibration test: six pitches with designed labels can show whether a confidence is informative, not whether it's calibrated. No affiliation with TypeSafe. Jev ran on OpenRouter's alpha Decisions API, served as typesafe/jev-1.13-20260917.