more per right answer: messy data plus a frontier model, against documented data plus the cheapest model, which also got more right
Tokenomics is an architecture decision
Token prices keep falling and AI bills keep rising. I tested the cheapest fix I know against the most expensive one: describe the data, or upgrade the model.
In the same September week, OpenAI priced GPT-6 at half what GPT-5.6 cost and Anthropic launched Opus 5.5 at 40% less to run than Opus 5. A month earlier, Gartner forecast that the cost of running an agentic workflow will rise more than fivefold through 2028 anyway.
Both can be true, because the bill isn't set by the price of a token. It's set by how many tokens the job burns, and that's an architecture decision, not a model one.
So I tested the cheapest architecture fix I know against the most expensive model fix. Three small tables (retail orders, SaaS support tickets, healthcare claims), 15 questions, four models, each table shown three ways: clean, messy, and messy with a short data dictionary. 540 recorded calls.
What the models actually saw
Here's one row from the support-ticket table, the way a model receives it, and the two ways to read it:
How many P1 (urgent) tickets are still open?
GPT-6 Luna, Gemini 3.8 Flash and Claude Opus 5.5 all answered 1, on every run. Luna and Flash said they were 99% sure. Opus said 80–85%.
Add "pr: 3=P1 (urgent), 2=P2, 1=P3 (low)" and all three answered 2, on every run, at 95–100%.
Every model read pr=1 as urgent, because that's what 1 usually means. In this table it meant low. No amount of intelligence recovers a convention nobody wrote down.
Jev, the decision model from my last lab, got it wrong too, but said it was only 11–18% sure. It knew it didn't know.
What was the total paid amount across all claims, in dollars?
$14,175.10, on every run, 99–100% sure. It added three claims that had been loaded twice and flagged dq=Y.
Add "dq: Y marks a duplicate load; exclude dq=Y rows" and it answered $12,432.16, on every run.
Opus got this one right without help: it worked out what dq=Y meant. That's what the upgrade buys you. One line of documentation bought the same thing for the smallest model.
Across all 15 questions
Messy data is a token tax
A confession. I estimated this run's cost from input tokens and was off by 3.4×. The models didn't read more. They thought more, and output is the expensive side of the meter.
The playbook
Build 15–30 questions with known answers on your own data, and price each model by what it spends per right answer. The trap is the per-token dashboard: it ranks Flash cheaper than Opus, and per right answer on messy data, it wasn't.
35 of the 82 wrong answers on messy data were a misread status, region or priority code, including all 9 of Opus's. Don't expect a smarter model to infer them: nothing in the table says st=2 means Returned. That's information, not intelligence.
Cents read as dollars and duplicate rows counted twice caused most of the rest. Then rerun the question set: if accuracy doesn't move, the dictionary is still missing something.
Take the cheapest one that clears your accuracy bar on the described data. At a 90% bar, that was the smallest model on the panel.
Describe the data before you upgrade the model.
What was the last model upgrade your team approved, and did anyone check the column names first?
- ModelCensus Labs — Semantic Layer vs. Model Upgrade: every call, replayable (run 2026-09-28-r1)
- The Register — Gartner: agentic AI inference costs per workflow set to rise fivefold by 2028 (17 Aug 2026)
- Gartner press release — inference costs per agentic workflow (17 Aug 2026)
- Anthropic — Claude Opus 5.5 (22 Sep 2026)
- OpenAI — Introducing GPT-6 Sol and Luna (Sep 2026)
Every figure here describes something measured and committed. See the measurements · read the method