Three curves went vertical. Reliability isn't one of them.
Context up ~5,000×. Price down ~99%. Release cadence: three months. None of it is a reliability curve.
Here is the signal almost nobody is pricing in: every capability curve in this industry is bending upward, and the failure modes have not moved.
Not improved slowly. Not improved unevenly. The same behaviours we measure today were measurable at 4,000 tokens and $60 per million.
Three converging forces
Price fell 84–99% since early 2023 depending how you index it. Cadence is now a frontier release every few months.
Old way, new way
When tokens were expensive, filtering paid for itself. That filtering was doing two jobs — saving money and protecting the answer — and only one was ever in the budget. Cheap tokens removed the first reason and left the second untouched.
How this plays out
Immediate: teams stop filtering because they can afford not to, and quality-per-token gets worse while cost-per-answer gets better. Both numbers look fine on their own dashboards.
Near-term: the failures that follow are attributed to the model rather than to the context pipeline, because nobody is running a control that could tell the difference.
Dominant: whoever has a named, measured taxonomy of how these systems break owns the vocabulary everyone else debugs in. That is a boring, compounding asset, and it does not depend on which model wins.
What I'd do this quarter
Add the null control to one eval. Name the three failures that hurt you most. Record the training cutoff of every model you deploy — you will need it and nobody logs it.
This is an extrapolation and I have labelled it one. A curve continued past its data is a hypothesis with a graph attached.
Which of these three curves do you think breaks first — and what happens to your stack when it does?
Every figure here describes something measured and committed. See the measurements · read the method