ModelCensusopen-source ai reliability harness
Toolkit

Use cases, tested before they're published.

Each use case goes through Labs first: the method against a plain request, every figure checked in code. What's published is whatever the test supports, including when the answer is that the prompt didn't help.

Excel → PowerPoint

tested in 2 labs · 240 decks · every figure checked in code

We built a prompt that guards against five failure modes and tested it against a plain request. It didn't make the decks more accurate: on messy exports, 54 of 60 plain decks had every figure right against 49 of 60 with the prompt, at 1.6× the output tokens. What did move the result was the model, and checking the numbers.

Decks fully rightTidy · plainTidy · promptMessy · plainMessy · prompt
GPT-6 Luna20 / 2020 / 2018 / 2015 / 20
Gemini 3.8 Flash20 / 2020 / 2016 / 2014 / 20
Claude Opus 5.520 / 2020 / 2020 / 2020 / 20

Tidy workbooks · Messy exports · replay any deck, side by side

  1. 1 · Pick the model for the job

    Claude Opus 5.5 had every figure right on 40 of 40 messy decks. The cheaper models missed on some, mostly arithmetic and blank cells. For numbers leaders will act on, the model was the biggest lever in our tests.

  2. 2 · Ask plainly

    Say the period, what counts (actuals, not forecast) and what to leave out. A plain request did as well as our prompt. One of the prompt's own rules caused errors: told never to infer a mapping, two decks refused a region the workbook documents as renamed.

  3. 3 · Check the deck before it goes out
    Without code: five checks, ten minutes
    • Recompute the total. One SUM in Excel, with the same filters you asked for. Compare it to slide one. fmi_4_1
    • Spot-check one row and one percentage. Pick a region: its revenue, then its growth or margin, worked out yourself. fmi_4_1
    • Check the scope. Right quarter, actuals not forecast, and anything you said to exclude (intercompany, returns) actually excluded. fmi_3_1
    • Look for figures that can't exist. Something new this period has no growth rate. It should say n/a, not a number. fmi_1_1
    • Check renamed or merged categories. A region, product or account that changed name should still be matched to its target. fmi_1_5
    With code: the checker the labs used

    Compute the figures yourself, then let the SDK find each one on the slides and list every number that has no source. It flagged every wrong deck in the messy-export lab.

    pip install "git+https://github.com/npmav18/modelcensus#subdirectory=sdk/python"
    
    from modelcensus_toolkit import checks
    from modelcensus_toolkit.checks import Figure
    
    expected = [Figure("Q3 total", 12_010_900, "money"),
                Figure("Midwest growth", 11.46, "pct", "Midwest")]
    for r in checks.reconcile(deck["slides"], expected):
        print("✓" if r.ok else "✗", r.figure.name, r.found)
    print(checks.unsourced(deck["slides"], expected))
    SDK on GitHub →
The tested prompt files

Kept exactly as tested, so the result can be checked. A published file never changes; an improved prompt would be v2, and would be published only if a lab shows it beats a plain request.

Prompt files: CC BY 4.0, credit “ModelCensus (modelcensus.org)”. SDK: Apache-2.0. Each file is pinned in manifest.json by its sha256.