Cheap tokens removed your discipline, not the penalty
Filtering used to pay for itself. Now it has to be justified on quality alone — and almost nobody measures that.
The weak signal: retrieval pipelines are getting less selective, and nobody is treating that as a decision.
It is a decision. It just used to be made by the invoice.
The forces
In March 2023 GPT-4 launched at $30 per million input tokens and $60 per million output. By 2026 one frontier price index sits at 16 against a March 2023 base of 100 — an 84% fall — while individual model-to-successor comparisons report reductions past 99%.
The exact multiple depends on whether you index by capability tier or by name. The direction is not in dispute.
Old way, new way
When a million tokens cost $30, nobody stuffed a corpus into a prompt. Filtering was cheaper than not filtering, so everybody filtered.
That filtering was doing two jobs — saving money and protecting the answer. Only one of them was ever in the budget.
Cheap tokens removed the first reason and left the second completely untouched. Unfiltered volume still brings distraction, latency and instability. It just stopped appearing on the invoice.
The timeline
Immediate: top-k values chosen under the old economics get quietly raised, because there is no longer a reason not to.
Near-term: quality-per-token degrades while cost-per-answer improves. Both are tracked. Neither dashboard shows the trade.
Dominant: the teams that measured the quality penalty directly are the ones who can tune k on evidence instead of vibes.
What to do now
If your top-k was chosen when tokens were expensive, it was chosen by a constraint that no longer binds. Re-tune it against a measured quality penalty.
Why: The right new value is probably not 'much larger', which is the direction a cost-only argument pushes you.
When did you last change your top-k — and was it because of a measurement or because you could afford to?
Every figure here describes something measured and committed. See the measurements · read the method