ModelCensusopen-source ai reliability harness
Labs · Tokenomics is an architecture decision← all labs

Semantic Layer vs. Model Upgrade

On messy data, what buys more accuracy per dollar: describing the data, or buying a bigger model?

Recorded 28 Sept 2026 against the live APIs · run 2026-09-28-r1 · 540 calls · replayed here, not live. Download the run

Describing the data took GPT-6 Luna from 56% to 91% correct. Upgrading to Claude Opus 5.5 on the same messy data reached 80%, at 44× the cost per correct answer.

ModelClean dataMessy dataMessy data + semantic layerMedian latencyCost per correct
Jev 1.1378%13%62%266 ms$0.000086
GPT-6 Luna100%56%91%4016 ms$0.000272
Gemini 3.8 Flash100%69%100%7364 ms$0.0064
Claude Opus 5.5100%80%100%5895 ms$0.0085

Accuracy per condition over 45 calls each (15 questions × 3 takes). Cost per correct answer is total spend over right answers, all conditions.

What was shipped revenue in 2026-07, in dollars?

Retail orders table:
order_id,month,region,status,amount_usd
5100,2026-07,East,Shipped,519.81
5101,2026-07,South,Shipped,526.73
5102,2026-08,West,Shipped,140.73
5103,2026-07,West,Shipped,302.25
5104,2026-08,West,Shipped,295.81
5105,2026-07,North,Shipped,797.23
5106,2026-08,North,Cancelled,442.44
5107,2026-08,East,Shipped,635.46
5108,2026-07,South,Shipped,840.30
5109,2026-07,West,Shipped,574.66
5110,2026-07,West,Shipped,683.33
5111,2026-08,East,Cancelled,857.88
5112,2026-08,South,Returned,745.29
5113,2026-07,West,Returned,379.97
5114,2026-07,West,Shipped,535.97
5115,2026-07,South,Shipped,368.40
5116,2026-08,East,Shipped,613.17
5117,2026-08,West,Shipped,503.39
5118,2026-08,West,Shipped,678.19
5119,2026-07,South,Returned,247.66
5120,2026-08,South,Cancelled,718.16
5121,2026-08,North,Shipped,735.93
5122,2026-07,East,Shipped,663.10
5123,2026-08,North,Shipped,816.71
take
recorded 28 Sept 2026 · replayed at recorded speed
Jev 1.13decision model
0 ms

Press replay.

GPT-6 Lunachat model · schema-enforced
0 ms

Press replay.

Gemini 3.8 Flashchat model · schema-enforced
0 ms

Press replay.

Claude Opus 5.5chat model · schema-enforced
0 ms

Press replay.

What describing the data buys each model

Messy dataMessy data + semantic layerJev 1.1313% → 62%GPT-6 Luna56% → 91%Gemini 3.8 Flash69% → 100%Claude Opus 5.580% → 100%
Accuracy on the same questions and the same messy table, before and after adding a short data dictionary.
Accuracy vs. cost per correct answer0%20%40%60%80%100%$0.00001$0.0001$0.001$0.01$0.1cost per correct answer (log)accuracyJev 1.13 · cleanGPT-6 Luna · cleanGPT-6 Luna · + dictionaryClaude Opus 5.5 · messy
One dot per model and condition. The dotted line is the best-value frontier: nothing is both cheaper per right answer and more accurate. Also labelled: GPT-6 Luna with the fix, and Claude Opus 5.5 without it.
Misread a code35Read cents as dollars21Counted duplicate rows14Ignored the status filter7Read seconds as minutes4No answer (provider error)1
Why the 82 wrong answers on the messy data were wrong, all models. Each wrong option is the answer one specific data mistake produces.

What this shows, and what it doesn't

Every chat model got every clean question right; Jev got 78%. Messy data dropped them all: GPT-6 Luna to 56%, Gemini 3.8 Flash to 69%, Claude Opus 5.5 to 80%.

A short data dictionary, 100 to 160 extra input tokens, brought Flash and Opus back to 100% and Luna to 91%. Luna with the dictionary beat Opus without it, for about 44 times less per correct answer.

Messy data is also a token tax. On the messy tables Luna and Flash wrote about 3.4 times the output tokens they wrote on clean ones (Luna 146 to 501, Flash 611 to 2,078), took 1.6 to 2.5 times as long across the three chat models, and still got more wrong. Adding the dictionary cut cost per call by a quarter for Luna, a third for Flash and a tenth for Opus, despite the extra input.

The biggest source of error was codes: 35 of the 82 wrong answers on messy data were a misread status, region or priority code, including all 9 of Opus's. No model can know that st=2 means Returned. That's information, not intelligence, and only the data layer can supply it.

Jev is a decision model, not a calculator: 78% on clean tables and 13% on messy ones, mostly from reading cents as dollars. The right model for summing a table isn't the right model for routing a ticket.

What this can't show: three synthetic tables of about 25 rows (retail orders, SaaS tickets, healthcare claims), each shown clean, messy and messy with a dictionary; 15 questions with four options each; three takes per model. The messy codes were deliberately undocumented, so some messy answers were guesses by design. That's the comparison, not a flaw in the models. It says nothing about text-to-SQL at warehouse scale. One Gemini call failed with a provider error and counts as wrong. No affiliation with any vendor.