BS Detector
Does a pitch say anything? Both models return the same typed answer. One decides it; the other generates it as text, confidence included.
Recorded 27 Sept 2026 against the live APIs · run 2026-09-27-r2 · 60 calls · replayed here, not live. Download the run
Jev 1.13: Same answer on all 6 pitches. 6× faster, 3× cheaper per call, and its confidence moves when the pitch is ambiguous.
Our AI-native platform leverages cutting-edge agentic workflows to unlock 10x productivity across the enterprise. Built for the future of work, it seamlessly transforms how teams collaborate, innovate and deliver value at scale.
Press replay.
Press replay.
How sure is it, pitch by pitch?
What this shows, and what it doesn't
Both get the easy call right. Average substance on the hollow pitches vs. the backed ones: Jev 1.13 1.7 vs. 7.0; GPT-4o-mini 2.7 vs. 7.9. On these six pitches, neither is more correct.
The difference is in the shape of the answer. Jev returns a distribution for every question, and its confidence moves with the input: lowest on the security pitch, whose “99.9%” is fairly both a buzzword and a vanity metric. The chat model writes its confidence as a number in text, and it barely moves.
Both arms returned valid structured output on every call. The chat model's shape was enforced by a strict JSON schema, so parsing isn't the difference here. Speed, cost and the confidence are. It was served by OpenAI and Azure through OpenRouter.
Not a calibration test: six pitches with designed labels can show whether a confidence is informative, not whether it's calibrated. No affiliation with TypeSafe. Jev ran on OpenRouter's alpha Decisions API, served as typesafe/jev-1.13-20260917.