Stop reporting a hallucination rate. Report what a wrong answer costs you.
5% is fine for a first draft and unacceptable for a dosage. The rate alone is not a risk statement.
Every vendor deck has a hallucination rate on it. Almost none of them means anything on its own.
A 5% error rate is perfectly fine in a brainstorming tool and catastrophic in a clinical one. Same number, opposite decisions.
Why a bare rate accumulates debt
Reporting rate without consequence pushes every conversation toward the wrong optimisation. Teams chase the number down uniformly, spending equally on failures that cost nothing and failures that end careers.
It also makes the number impossible to challenge. If nobody has said what an acceptable rate is for this specific decision, any rate can be argued as fine or as alarming.
What to report instead
Classical FMEA carries severity alongside occurrence for exactly this reason. Our taxonomy publishes severity axes per failure mode — how prevalent, how harmful, and how stealthy it is when it happens.
Stealth is the one people skip and it is often the most important. A failure that announces itself is an inconvenience. A failure that looks exactly like success is a liability, and the two deserve different budgets even at the same rate.
Where a bare rate is fine
Tracking your own progress over time. If you are watching one mode on one system, the rate alone tells you whether last month's change helped.
The moment it crosses an organisational boundary — into a deck, a risk register, a vendor comparison — it needs its severity and its interval attached or it will be read as whatever the reader already believed.
The cost of false confidence is the number. The rate is just one of its inputs.
For your highest-stakes AI decision — what failure rate would actually be acceptable, and has anyone written it down?
Every figure here describes something measured and committed. See the measurements · read the method