output tokens the decision model wrote on every one of 135 answers, clean data or messy. The chat models ranged from 21 to 5,454.
Bad data used to cost accuracy. Now it costs tokens.
In traditional ML, a prediction had a fixed price and data quality moved accuracy. In GenAI it moves the bill too. One run showed both paradigms side by side.
In traditional ML, a prediction cost what it cost. You paid for training, and every inference after that had a fixed, small price. Bad data cost you accuracy. That was the whole bill.
I noticed that stopped being true by accident, when my cost estimate for last week's lab came in 3.4× low.
Both paradigms, one run
That lab had a decision model in it, Jev 1.13, which picks from options instead of writing text. Economically, it behaves like traditional ML: a fixed-size output per call. So I had both paradigms answering the same questions, over the same clean and messy tables, in the same run.
How many P1 (urgent) tickets are still open?
Wrote 710 to 1,254 output tokens per answer. Answered 1, 99–100% sure. Wrong.
Wrote 45 tokens. Answered 1, 11–18% sure. Also wrong, but cheaply, and it said so.
Neither could know that priority 1 meant low in this table. One spent 16 to 28 times the output reaching that wrong answer, and grew more confident doing it.
That's the shift in one line: in GenAI, data quality sits on both sides of the ledger.
What changes next
How to prepare
Track cost per correct answer on a fixed question set, per model and per data source. Price per token tells you what the meter charges, not what the job costs.
In my test, messy tables raised a chat model's cost per call by up to 2.9×, on top of the wrong answers. That's a number a CFO can fund against.
Routing, triage and classification don't need output that grows with the mess. A decision model's cost stayed flat. Its accuracy still needed good data: 13% on the messy tables.
Bad data used to cost accuracy. Now it costs tokens too.
When did your data-quality budget last get compared with your inference bill?
- ModelCensus Labs — Semantic Layer vs. Model Upgrade: every call, replayable (run 2026-09-28-r1)
- ModelCensus — Tokenomics is an architecture decision (the same experiment, on cost per correct answer)
- The Register — Gartner: agentic AI inference costs per workflow set to rise fivefold by 2028 (17 Aug 2026)
Every figure here describes something measured and committed. See the measurements · read the method