the output tokens my failure-mode prompt cost on messy exports. Decks with every figure right: 54 of 60 without it, 49 of 60 with it.
My prompt pack didn't make the decks better
I wrote a prompt that guards against five ways AI gets spreadsheet numbers wrong, and tested it on 240 decks. It didn't help. Two things did.
I set out to build the prompt every finance team should paste before asking AI for a deck: turn this workbook into slides, and don't get the numbers wrong. Five guards, one per failure mode in our index. Map codes from the lookup sheet. Compute in a table first. Convert units. Respect the scope. Write n/a, never guess.
Then I tested it before publishing it. It didn't help.
What I expected, and what happened
On tidy workbooks, every model got every figure right, with or without the prompt. So I rebuilt the test on what a real export looks like: subtotal rows that include intercompany revenue, returns as (45.2), a region renamed since the targets were set, units buried in a note.
That broke the cheaper models. It didn't break them less with my prompt. In my test, 54 of 60 plain decks were fully right against 49 of 60 with it, and the prompt cost 1.6× the output tokens. Twenty decks per cell is small, so I read that as "no better", not "worse".
Why the prompt didn't fix it
Total Q3 external revenue.
Total External Revenue: $11,010.9K ($11.01M across all operating regions)
$12,010.9K. Every region's figure on the next slide was right. The sum was off by exactly a million.
Arithmetic slips like this showed up with and without the prompt. Telling a model to show its working doesn't make it add better.
Midwest's Q3 revenue against its target. The lookup sheet says: Midwest, formerly Central. The target sheet lists Central.
Midwest: $1,832,200 USD | target n/a: workbook target is "Central"; no exact region match
$1,832,200 against $1,910,000: 95.9% of target.
My own guard did this: "never infer a mapping". Two decks with the prompt refused a documented rename. No plain deck did.
- Arithmetic doesn't respond to instructions. The slips appeared in both arms.
- Strict rules get applied too strictly. My mapping guard turned a documented rename into a gap.
- Longer prompts didn't stop misreads. GPT-6 Luna got the new region's revenue wrong (the one with blank Q2 cells) in six decks, five of them with my prompt.
What did work
Claude Opus 5.5 got all 26 figures right on all 40 messy decks. And every wrong deck was caught the moment its figures were checked against the workbook in code. That check is what I've published, not the prompt. For people who don't write code, it's five checks that take ten minutes. For people who do, it's the checker the labs used.
| If you… | Do this |
|---|---|
| Don't write code | Recompute the total in Excel. Spot-check one row and one percentage. Check the scope. Look for figures that can't exist. Check renamed categories. |
| Write code | pip install the ModelCensus toolkit SDK. Compute the expected figures yourself; reconcile() finds each on the slides, unsourced() lists the rest. |
| Choose the tools | For numbers leaders act on, pick the model by testing it on your own exports, not by price per token. |
The prompt files are still public, unchanged, with this result next to them. If I write a v2, it ships only when a lab shows it beats a plain request.
Don't prompt harder. Check the numbers.
How many AI-built decks went to your leadership team last quarter, and how many had their numbers checked by something other than the model that wrote them?
- ModelCensus Labs — Excel to PPTX, tidy workbooks (run 2026-10-08-r1)
- ModelCensus Labs — Excel to PPTX, messy export (run 2026-10-08-r1)
- ModelCensus Toolkit — Excel → PowerPoint: checks and tested prompt files
- ModelCensus toolkit SDK (Python, Apache-2.0)
- The tested prompt, as published: modelcensus_pptxfromexcel_any_v1.txt
Every figure here describes something measured and committed. See the measurements · read the method