ModelCensusopen-source ai reliability harness
45

output tokens the decision model wrote on every one of 135 answers, clean data or messy. The chat models ranged from 21 to 5,454.

Blog30 Sept 2026tokenomicslabsstrategyenterprise

Bad data used to cost accuracy. Now it costs tokens.

In traditional ML, a prediction had a fixed price and data quality moved accuracy. In GenAI it moves the bill too. One run showed both paradigms side by side.

In traditional ML, a prediction cost what it cost. You paid for training, and every inference after that had a fixed, small price. Bad data cost you accuracy. That was the whole bill.

I noticed that stopped being true by accident, when my cost estimate for last week's lab came in 3.4× low.

Traditional MLFixed cost per predictionBad data costs accuracyBudget: computeCost known in advancevs.Generative AICost varies per answerBad data costs accuracy and tokensBudget: cost per correct answerCost known after the answer
The same data problem lands on one side of the ledger in traditional ML, and on both in GenAI.

Both paradigms, one run

That lab had a decision model in it, Jev 1.13, which picks from options instead of writing text. Economically, it behaves like traditional ML: a fixed-size output per call. So I had both paradigms answering the same questions, over the same clean and messy tables, in the same run.

Same question, same messy table

How many P1 (urgent) tickets are still open?

Gemini 3.8 Flash · generative

Wrote 710 to 1,254 output tokens per answer. Answered 1, 99–100% sure. Wrong.

Jev 1.13 · decision model

Wrote 45 tokens. Answered 1, 11–18% sure. Also wrong, but cheaply, and it said so.

Neither could know that priority 1 meant low in this table. One spent 16 to 28 times the output reaching that wrong answer, and grew more confident doing it.

clean datamessy dataJev 1.13 (decision)1.0× → 1.0×Claude Opus 5.51.0× → 1.3×GPT-6 Luna1.0× → 2.3×Gemini 3.8 Flash1.0× → 2.9×
Cost per call on messy tables, relative to the same questions on clean ones. The decision model's cost didn't move. Every generative model's did.
Jev 1.131.5×Claude Opus 5.54.5×Gemini 3.8 Flash14.4×GPT-6 Luna19.7×
Dearest call over cheapest call, 135 calls per model. A decision model's cost is set by what you send it. A generative model's is set by how long it thinks, and messy data makes it think longer.

That's the shift in one line: in GenAI, data quality sits on both sides of the ledger.

What changes next

Traditional MLBudget compute. Data quality is an accuracy project.NowBudget tokens. Dashboards show price per token.6–12 monthsFinance asks for cost per correct answer.12–18 monthsData quality gets funded from the inference budget.
A prediction, not a measurement. The second row is already true.

How to prepare

Price outcomes, not tokens

Track cost per correct answer on a fixed question set, per model and per data source. Price per token tells you what the meter charges, not what the job costs.

Make the data-quality case in inference terms

In my test, messy tables raised a chat model's cost per call by up to 2.9×, on top of the wrong answers. That's a number a CFO can fund against.

Use a fixed-cost model where the job is a choice

Routing, triage and classification don't need output that grows with the mess. A decision model's cost stayed flat. Its accuracy still needed good data: 13% on the messy tables.

Bad data used to cost accuracy. Now it costs tokens too.

When did your data-quality budget last get compared with your inference bill?

Every figure here describes something measured and committed. See the measurements · read the method