ModelCensusopen-source ai reliability harness
Blog4 Sept 2026strategyindustrycrystal-ball

Three curves went vertical. Reliability isn't one of them.

Context up ~5,000×. Price down ~99%. Release cadence: three months. None of it is a reliability curve.

Here is the signal almost nobody is pricing in: every capability curve in this industry is bending upward, and the failure modes have not moved.

Not improved slowly. Not improved unevenly. The same behaviours we measure today were measurable at 4,000 tokens and $60 per million.

Three converging forces

context window (tokens)2K20204K2022100KJul 23128KLate 231M202510M2026log scale
Log scale, or the first four years are a flat line. Roughly 5,000× in six years.

Price fell 84–99% since early 2023 depending how you index it. Cadence is now a frontier release every few months.

Old way, new way

Becoming obsoleteFiltering to save costBenchmark-a-generationWhich-model rankingsvs.Becoming valuableNamed failure modesControls that isolate causeProvenance you can walk

When tokens were expensive, filtering paid for itself. That filtering was doing two jobs — saving money and protecting the answer — and only one was ever in the budget. Cheap tokens removed the first reason and left the second untouched.

How this plays out

Immediate: teams stop filtering because they can afford not to, and quality-per-token gets worse while cost-per-answer gets better. Both numbers look fine on their own dashboards.

Near-term: the failures that follow are attributed to the model rather than to the context pipeline, because nobody is running a control that could tell the difference.

Dominant: whoever has a named, measured taxonomy of how these systems break owns the vocabulary everyone else debugs in. That is a boring, compounding asset, and it does not depend on which model wins.

What I'd do this quarter

Add the null control to one eval. Name the three failures that hurt you most. Record the training cutoff of every model you deploy — you will need it and nobody logs it.

This is an extrapolation and I have labelled it one. A curve continued past its data is a hypothesis with a graph attached.

Which of these three curves do you think breaks first — and what happens to your stack when it does?

Every figure here describes something measured and committed. See the measurements · read the method