Context windows grew a hundredfold. Evals did not.
We can put a book in the prompt. We still measure as though it were a paragraph.
Four years ago a long prompt was a few thousand tokens. Now it is a few hundred thousand, and the standard way to evaluate whether that helps is unchanged: ask with context, ask without, compare.
That design was adequate when context was a paragraph. It is not adequate when context is a corpus, because at that size two different things are happening at once and the two-arm comparison adds them together.
Volume has a price beyond the invoice
Every token you add is a token the model has to attend past. Unfiltered volume brings distraction, latency, and — in a finding that surprised us — instability on questions the added text says nothing about. On this panel, paraphrase stability is near perfect when a question is asked plainly and degrades sharply once an irrelevant document is in the prompt.
A retrieval step that returns something irrelevant has not failed to help. It has spent your budget to make the answer worse.
Hold volume constant and vary quality, or hold quality constant and vary volume. Varying both at once produces a number that cannot be attributed to either.
Why: This census holds volume constant — every arm sits at 160–175 tokens — which is the only reason a change can be attributed to what the context says rather than to how much of it there is.
Every figure here describes something measured and committed. See the measurements · read the method