Semantic Layer vs. Model Upgrade
On messy data, what buys more accuracy per dollar: describing the data, or buying a bigger model?
Recorded 28 Sept 2026 against the live APIs · run 2026-09-28-r1 · 540 calls · replayed here, not live. Download the run
Describing the data took GPT-6 Luna from 56% to 91% correct. Upgrading to Claude Opus 5.5 on the same messy data reached 80%, at 44× the cost per correct answer.
| Model | Clean data | Messy data | Messy data + semantic layer | Median latency | Cost per correct |
|---|---|---|---|---|---|
| Jev 1.13 | 78% | 13% | 62% | 266 ms | $0.000086 |
| GPT-6 Luna | 100% | 56% | 91% | 4016 ms | $0.000272 |
| Gemini 3.8 Flash | 100% | 69% | 100% | 7364 ms | $0.0064 |
| Claude Opus 5.5 | 100% | 80% | 100% | 5895 ms | $0.0085 |
Accuracy per condition over 45 calls each (15 questions × 3 takes). Cost per correct answer is total spend over right answers, all conditions.
What was shipped revenue in 2026-07, in dollars?
Retail orders table: order_id,month,region,status,amount_usd 5100,2026-07,East,Shipped,519.81 5101,2026-07,South,Shipped,526.73 5102,2026-08,West,Shipped,140.73 5103,2026-07,West,Shipped,302.25 5104,2026-08,West,Shipped,295.81 5105,2026-07,North,Shipped,797.23 5106,2026-08,North,Cancelled,442.44 5107,2026-08,East,Shipped,635.46 5108,2026-07,South,Shipped,840.30 5109,2026-07,West,Shipped,574.66 5110,2026-07,West,Shipped,683.33 5111,2026-08,East,Cancelled,857.88 5112,2026-08,South,Returned,745.29 5113,2026-07,West,Returned,379.97 5114,2026-07,West,Shipped,535.97 5115,2026-07,South,Shipped,368.40 5116,2026-08,East,Shipped,613.17 5117,2026-08,West,Shipped,503.39 5118,2026-08,West,Shipped,678.19 5119,2026-07,South,Returned,247.66 5120,2026-08,South,Cancelled,718.16 5121,2026-08,North,Shipped,735.93 5122,2026-07,East,Shipped,663.10 5123,2026-08,North,Shipped,816.71
Press replay.
Press replay.
Press replay.
Press replay.
What describing the data buys each model
What this shows, and what it doesn't
Every chat model got every clean question right; Jev got 78%. Messy data dropped them all: GPT-6 Luna to 56%, Gemini 3.8 Flash to 69%, Claude Opus 5.5 to 80%.
A short data dictionary, 100 to 160 extra input tokens, brought Flash and Opus back to 100% and Luna to 91%. Luna with the dictionary beat Opus without it, for about 44 times less per correct answer.
Messy data is also a token tax. On the messy tables Luna and Flash wrote about 3.4 times the output tokens they wrote on clean ones (Luna 146 to 501, Flash 611 to 2,078), took 1.6 to 2.5 times as long across the three chat models, and still got more wrong. Adding the dictionary cut cost per call by a quarter for Luna, a third for Flash and a tenth for Opus, despite the extra input.
The biggest source of error was codes: 35 of the 82 wrong answers on messy data were a misread status, region or priority code, including all 9 of Opus's. No model can know that st=2 means Returned. That's information, not intelligence, and only the data layer can supply it.
Jev is a decision model, not a calculator: 78% on clean tables and 13% on messy ones, mostly from reading cents as dollars. The right model for summing a table isn't the right model for routing a ticket.
What this can't show: three synthetic tables of about 25 rows (retail orders, SaaS tickets, healthcare claims), each shown clean, messy and messy with a dictionary; 15 questions with four options each; three takes per model. The messy codes were deliberately undocumented, so some messy answers were guesses by design. That's the comparison, not a flaw in the models. It says nothing about text-to-SQL at warehouse scale. One Gemini call failed with a provider error and counts as wrong. No affiliation with any vendor.