Excel to PPTX
Ask a model for a 4-slide Q3 review from a small workbook. How many figures on the slides are right, and does the ModelCensus toolkit prompt change that?
Recorded 8 Oct 2026 against the live APIs · run 2026-10-08-r1 · 120 calls · replayed here, not live. Download the run · Prompt tested: modelcensus_pptxfromexcel_any_v1.txt · Failure mode: Schema-Correspondence Fabrication · Failure mode: Method/Execution Split · Failure mode: Dimensional-Analysis Failure · Failure mode: Constraint Decay at Depth · Failure mode: Unsupported-Claim Confidence
| Figure | GPT-6 Luna · bare | GPT-6 Luna · toolkit | Gemini 3.8 Flash · bare | Gemini 3.8 Flash · toolkit | Claude Opus 5.5 · bare | Claude Opus 5.5 · toolkit |
|---|---|---|---|---|---|---|
| Q3 total | 100% | 100% | 100% | 100% | 100% | 100% |
| Revenue by region (4) | 100% | 100% | 100% | 100% | 100% | 100% |
| Growth vs Q2, incl. n/a (4) | 100% | 100% | 100% | 100% | 100% | 100% |
| Against target (4) | 100% | 100% | 100% | 100% | 100% | 100% |
| Gross margin | 100% | 100% | 100% | 100% | 100% | 100% |
Replay a deck
Pick a workbook and a model. Left is the bare request; right is the same request with the toolkit prompt. Each figure is checked against the workbook in code.
- 1Q3 2026 Total Revenue: $8.07M
- · Q3 2026 (Jul–Sep) actual revenue: $8,073.8K
- · Up 25.5% vs. Q2 2026 ($6,431.2K); North first reported revenue in Q3
- · Excluding North, revenue grew 6.0% ($6,815.0K vs. $6,431.2K)
- · 101.6% of the combined Q3 target of $7,950.0K (+$123.8K)
- 2Q3 Revenue by Region and Growth vs. Q2
- · West: $2,496.8K, down 4.7% vs. Q2 ($2,621.0K)
- · South: $2,414.9K, up 12.5% vs. Q2 ($2,147.2K)
- · East: $1,903.3K, up 14.4% vs. Q2 ($1,663.0K)
- · North: $1,258.8K, new in Q3 (no Q2 revenue, so no growth rate)
- 3Q3 Revenue vs. Target by Region
- · West: $2,496.8K vs. $2,310.0K target, 108.1% (+$186.8K)
- · North: $1,258.8K vs. $1,200.0K target, 104.9% (+$58.8K)
- · East: $1,903.3K vs. $1,860.0K target, 102.3% (+$43.3K)
- · South: $2,414.9K vs. $2,580.0K target, 93.6% (–$165.1K), the only region below target
- · Total: $8,073.8K vs. $7,950.0K target, 101.6%
- 4Q3 Gross Margin: 38.2%
- · Gross profit: $3,084.4K on revenue of $8,073.8K (cost $4,989.4K)
- · Up from 37.7% in Q2 2026
- · By region: East 42.2%, North 40.3%, South 36.1%, West 36.1%
- 1Q3 2026 Total Revenue (Actuals)
- · Total Q3 2026 revenue: $8,073,800
- · Actuals only, Jul–Sep 2026, all 4 regions
- · Excludes $106,700 of Sep forecast for South
- 2Q3 Revenue by Region vs Q2
- · West: $2,496,800 (Q2 $2,621,000), -4.7% QoQ
- · South: $2,414,900 (Q2 $2,147,200), +12.5% QoQ
- · East: $1,903,300 (Q2 $1,663,000), +14.4% QoQ
- · North: $1,258,800; QoQ growth n/a: no Q2 rows for North in workbook
- 3Q3 Revenue vs Q3 Target by Region
- · West: $2,496,800 vs $2,310,000 target, +$186,800 (108.1%)
- · South: $2,414,900 vs $2,580,000 target, -$165,100 (93.6%)
- · East: $1,903,300 vs $1,860,000 target, +$43,300 (102.3%)
- · North: $1,258,800 vs $1,200,000 target, +$58,800 (104.9%)
- · Total: $8,073,800 vs $7,950,000, +$123,800 (101.6%)
- 4Q3 2026 Gross Margin (Actuals)
- · Revenue: $8,073,800
- · Cost: $4,989,400
- · Gross profit: $3,084,400
- · Gross margin: 38.2%
What this shows, and what it doesn't
Every deck got every figure right, with or without the toolkit prompt. On a 35-row workbook with units in the headers and a status legend, all three models mapped the region codes, converted $000, left out the forecast rows (including a forecast true-up inside September) and wrote n/a for the region with no Q2.
The toolkit prompt cost 1.9× the output tokens for no change in accuracy. That is the result, and the prompt file is still published, unchanged, so anyone can check it.
What this can't show: one task shape, ten synthetic workbooks, three models, two takes each. The messy-export lab repeats it on workbooks shaped like real finance exports.