ModelCensusopen-source ai reliability harness
44×

more per right answer: messy data plus a frontier model, against documented data plus the cheapest model, which also got more right

Blog28 Sept 2026tokenomicslabsstrategyenterprise

Tokenomics is an architecture decision

Token prices keep falling and AI bills keep rising. I tested the cheapest fix I know against the most expensive one: describe the data, or upgrade the model.

In the same September week, OpenAI priced GPT-6 at half what GPT-5.6 cost and Anthropic launched Opus 5.5 at 40% less to run than Opus 5. A month earlier, Gartner forecast that the cost of running an agentic workflow will rise more than fivefold through 2028 anyway.

Both can be true, because the bill isn't set by the price of a token. It's set by how many tokens the job burns, and that's an architecture decision, not a model one.

So I tested the cheapest architecture fix I know against the most expensive model fix. Three small tables (retail orders, SaaS support tickets, healthcare claims), 15 questions, four models, each table shown three ways: clean, messy, and messy with a short data dictionary. 540 recorded calls.

GPT-6 Lunacheapest modeldictionary~150 tokens91% right1× per right answerClaude Opus 5.5frontier modelmessy datapr 1 · dq Y80% right44× per right answer
Weight is right answers. In my test, the team that answered messy data with a bigger model paid 44× more per correct answer than the team that documented its data and used the cheapest model, and still got fewer right.

What the models actually saw

Here's one row from the support-ticket table, the way a model receives it, and the two ways to read it:

the roweasy to assumethe dictionarytk 893ticket 893ticket 893pr 1✗ urgent✓ P3 · lowfrt_s 18720✗ 18,720 min✓ 312 minrs 0openopen
Nothing in the row is wrong. The meaning just isn't in it. "1" usually means urgent, and here it means low; frt_s is seconds, not minutes. The priority code is the one all three chat models misread; the units trap mostly caught Jev.
Question · options 1 / 2 / 3 / 5

How many P1 (urgent) tickets are still open?

Messy data

GPT-6 Luna, Gemini 3.8 Flash and Claude Opus 5.5 all answered 1, on every run. Luna and Flash said they were 99% sure. Opus said 80–85%.

+ one line of documentation

Add "pr: 3=P1 (urgent), 2=P2, 1=P3 (low)" and all three answered 2, on every run, at 95–100%.

Every model read pr=1 as urgent, because that's what 1 usually means. In this table it meant low. No amount of intelligence recovers a convention nobody wrote down.

Jev, the decision model from my last lab, got it wrong too, but said it was only 11–18% sure. It knew it didn't know.

Question · claims table

What was the total paid amount across all claims, in dollars?

Messy data · GPT-6 Luna

$14,175.10, on every run, 99–100% sure. It added three claims that had been loaded twice and flagged dq=Y.

+ one line of documentation

Add "dq: Y marks a duplicate load; exclude dq=Y rows" and it answered $12,432.16, on every run.

Opus got this one right without help: it worked out what dq=Y meant. That's what the upgrade buys you. One line of documentation bought the same thing for the smallest model.

Across all 15 questions

messy data+ data dictionaryGPT-6 Luna56% → 91%Gemini 3.8 Flash69% → 100%Claude Opus 5.580% → 100%
Accuracy on the same 15 questions and the same messy tables, before and after a 100–160-token dictionary. On clean tables all three scored 100%.
Luna + dictionary91% · 1×Luna, messy56% · 2×Flash + dictionary100% · 23×Opus + dictionary100% · 32×Opus, messy80% · 44×Flash, messy69% · 49×
Accuracy and cost per correct answer, relative to the cheapest model with the dictionary. Most of the 44× is the price gap between the models; the dictionary is what made the cheap one good enough to use. On the same model it cut cost per right answer about 2× for Luna and about 28% for Opus. The priciest setup of all was the mid-priced model on messy data, thinking its way around codes it couldn't read.

Messy data is a token tax

Luna · clean146Luna · messy501Luna · + dictionary331Flash · clean611Flash · messy2,078Flash · + dictionary1,312
Mean output tokens per answer. On messy tables Luna and Flash wrote about 3.4× what they wrote on clean ones, took up to 2.5× as long, and still got more wrong.

A confession. I estimated this run's cost from input tokens and was off by 3.4×. The models didn't read more. They thought more, and output is the expensive side of the meter.

The playbook

Measurecost per right answerName the codesst=2 means what?Declare unitsand duplicatesThen pick a modelcheapest that passes
1 · Measure cost per correct answer, not per token

Build 15–30 questions with known answers on your own data, and price each model by what it spends per right answer. The trap is the per-token dashboard: it ranks Flash cheaper than Opus, and per right answer on messy data, it wasn't.

2 · Name every code

35 of the 82 wrong answers on messy data were a misread status, region or priority code, including all 9 of Opus's. Don't expect a smarter model to infer them: nothing in the table says st=2 means Returned. That's information, not intelligence.

3 · Declare units and duplicates

Cents read as dollars and duplicate rows counted twice caused most of the rest. Then rerun the question set: if accuracy doesn't move, the dictionary is still missing something.

4 · Only then pick the model

Take the cheapest one that clears your accuracy bar on the described data. At a 90% bar, that was the smallest model on the panel.

Describe the data before you upgrade the model.

What was the last model upgrade your team approved, and did anyone check the column names first?

Every figure here describes something measured and committed. See the measurements · read the method