The window grew 5,000×. The control arm never appeared.
From 2,048 tokens to ten million in six years. Evaluation design did not move at all.
Here is the shift nobody is pricing in: context capacity went up by three orders of magnitude and the method for evaluating whether context helps is the same one from 2020.
The forces behind it
Architecture improvements made long attention tractable. Inference economics made it affordable. And competitive pressure made window size a headline number, which means it grew faster than the tooling around it.
Old way, new way
What is becoming obsolete: filtering to save money, and the assumption that more retrieved context is better context.
What is becoming valuable: knowing which of your failures context can actually reach. Every failure mode we measure today was measurable at 4,000 tokens. None of them went away.
How this plays out
Immediate: teams stop filtering because they can afford not to. Cost-per-answer improves, quality-per-token gets worse, and both dashboards look fine.
Near-term: the failures that follow get attributed to the model rather than the context pipeline, because nobody is running a control that could separate them.
Dominant: measurement that can attribute a change to content rather than volume becomes the differentiator, because everything else is commoditised.
What to do this quarter
Add one length-matched irrelevant arm to one eval. That is the whole intervention and it takes an afternoon.
Which is doing more work in your system right now — the retrieval, or the fact that something got retrieved?
Every figure here describes something measured and committed. See the measurements · read the method