ModelCensusopen-source ai reliability harness
Blog26 Aug 2026enterprisepositioninggovernanceexec-pitch

77% of executives say trust is the barrier. Almost none of them can measure it.

Adoption stopped being the hard part. Nobody moved the budget.

An Accenture survey puts it at 77%: executives naming trust, not adoption, as the primary barrier to scaling AI.

Sit with that. The bottleneck moved and most AI budgets did not.

The status quo tax

If trust is the constraint, every pound spent on more capability buys you less than a pound spent on evidence. Most programmes are still allocated as though adoption were the problem — more pilots, more use cases, more surface area to be uncertain about.

The compounding cost is not a failed project. It is a portfolio of systems nobody will let near a customer, each individually defensible, collectively stalled.

Budgeted for adoptionMore pilotsMore use casesTrust unaddressedstalls at the gatevs.Budgeted for trustNamed failure modesRates with intervalsEvidence a risk committee reads

Three columns

Financial

A measured failure-mode baseline across a 20-model panel cost us $25 and thirteen hours of compute.

Why: That is not a rounding error against a stalled programme. The expensive part of evaluation is deciding what to measure, and that is a week of thinking, not a capital request.

Risk

"We tested it and it seemed fine" is not a risk position. "This mode fails 13% of the time, interval 7 to 21, and here is the transcript behind every failure" is.

Why: The second sentence can be taken to a risk committee. The first gets a project deferred and nobody can say why.

Strategic

Trust is the gating factor on deployment velocity, and evidence is the only thing that moves it.

Why: You cannot persuade your way past a risk function. You can show it a rate with an interval and a link to the raw transcript.

The two objections

"Our use case is too specific for generic evals." Correct, and irrelevant. The transferable part is not our rates — it is the controls. A length-matched irrelevant arm, intervals at effective sample size, a trap-versus-answerable split. Copy the method, run it on your cases.

"We don't have the people." You need one person for two weeks to name your three worst failures and build a detector for the easiest one. That produces a number your organisation has never had.

The low-risk first step

Pick the failure that most recently caused an escalation. Name it precisely enough to write a pass/fail check. Measure it on your current model, with a control arm. One number, one fortnight.

If trust is your barrier too — what evidence would actually move your risk function, and do you currently produce any of it?

Every figure here describes something measured and committed. See the measurements · read the method