ModelCensusopen-source ai reliability harness
Blog11 Mar 2026strategycontextcrystal-ball

The window grew 5,000×. The control arm never appeared.

From 2,048 tokens to ten million in six years. Evaluation design did not move at all.

Here is the shift nobody is pricing in: context capacity went up by three orders of magnitude and the method for evaluating whether context helps is the same one from 2020.

May 2020GPT-3 — ~2,048 tokensNov 2022ChatGPT — 4K–8KJul 2023Claude 2 — 100KLate 2023GPT-4 Turbo — 128K2025–261M standard across flagships2026Llama 4 Scout — 10M advertised
Sources differ on GPT-3's launch window — 2,048 or 4,096 — which is worth stating given the size of the multiple.

The forces behind it

Architecture improvements made long attention tractable. Inference economics made it affordable. And competitive pressure made window size a headline number, which means it grew faster than the tooling around it.

Old way, new way

What is becoming obsolete: filtering to save money, and the assumption that more retrieved context is better context.

What is becoming valuable: knowing which of your failures context can actually reach. Every failure mode we measure today was measurable at 4,000 tokens. None of them went away.

Scaled hardWindow size~5,000×Price per tokendown ~99%vs.Did not moveCitations that resolveHolding a rule at depthHedging a stale fact

How this plays out

Immediate: teams stop filtering because they can afford not to. Cost-per-answer improves, quality-per-token gets worse, and both dashboards look fine.

Near-term: the failures that follow get attributed to the model rather than the context pipeline, because nobody is running a control that could separate them.

Dominant: measurement that can attribute a change to content rather than volume becomes the differentiator, because everything else is commoditised.

What to do this quarter

Add one length-matched irrelevant arm to one eval. That is the whole intervention and it takes an afternoon.

Which is doing more work in your system right now — the retrieval, or the fact that something got retrieved?

Every figure here describes something measured and committed. See the measurements · read the method