AI Bookkeeping Accuracy: Benchmarks vs Your Own Test Set

By Jude Lee · · News

Two accountants reviewing transaction listings and a laptop screen in a small firm office

What the new bookkeeping benchmark headlines actually say

In January 2026, a round of coverage reported that frontier AI models had outperformed a human baseline on a bookkeeping-accuracy benchmark. The write-ups I read were secondary — aggregation and commentary, not the organization that built or published the benchmark. I have not independently verified the methodology behind these numbers. Don’t treat them as a procurement decision either.

That isn’t a dodge; it’s the operating rule. Never quote a benchmark you can’t trace to its publisher’s own documentation. When someone brings you the figure — a vendor rep, a partner, a client who read it on LinkedIn — ask for the primary publication and read four things in it: which tasks were scored, who the human comparison group actually was (staff bookkeepers, licensed CPAs, crowdworkers?), whether the ledgers were synthetic or real client data, and whether “accuracy” meant an exact account match or something looser. If whoever is quoting the number can’t produce that document, the number isn’t evidence — it’s a vibe with a decimal point.

And even a well-documented benchmark answers a different question than yours. A benchmark is a controlled test on a fixed dataset. Your firm’s question is: does this tool code transactions correctly on the 40 client files we actually run, under our conventions, at a review cost we can absorb? Those two questions can have opposite answers.

A benchmark tells you a model can do bookkeeping. Only your own test set tells you whether it can do your bookkeeping.

Why a benchmark score won’t predict accuracy on your clients’ books

The gap between lab accuracy and engagement accuracy comes from things benchmarks can’t contain:

That last point is why I think a single number — from a benchmark or a vendor deck — is close to useless for firm decisions.

Build a firm evaluation set you can actually finish

This is the part most firms skip, and it’s the cheapest real evidence available to you.

  1. Sample across clients, not within one

    Pull transactions from a spread of engagement types you serve — say a handful of clients across your main industries. A sample drawn from one well-ruled client will flatter any tool.
  2. Pick a size you'll actually grade

    Something in the range of 150–300 transactions is usually enough to see patterns without eating a week. This is a practical design choice, not a statistically derived threshold; if you need defensible precision, involve someone who can size the sample properly.
  3. Freeze the ground truth

    Have a senior staffer code the sample the way the firm wants it coded, and lock that file. This is your answer key. Expect disagreement among your own staff here — that disagreement is a finding, and it’s often the real cause of “AI inaccuracy.”
  4. Strip or mask identifying data

    Before anything leaves your ledger, decide what can be shared. The IRS’s Publication 4557, Safeguarding Taxpayer Data, sets expectations for protecting client information; the practical implication is that your test set may need to be de-identified or kept inside a vendor environment covered by a written agreement. I compare the configurations in four setups for using client data with AI.
  5. Run every candidate against the same file

    Bank rules, your ledger’s built-in AI, and any agent you’re piloting — identical input, identical grading.
150–300
transactions in a starter eval set (a practical design choice, not a benchmark)
Illustrative
3 tiers
error classes I recommend grading against: immaterial, review-catchable, material
Author recommendation

Grade by error class, not by percentage

Score each miss into one of three buckets:

  1. Immaterial variance — defensible alternative coding, no effect on the financials or return. Arguably not an error at all.
  2. Review-catchable — wrong, but the kind of wrong your existing review step reliably catches (unusual account, odd vendor, missing class).
  3. Material / silent — wrong in a way that flows through to the statements or the return and looks plausible on review.

Tier 3 carries the whole decision, so define it before you grade, not after. Write down two things in advance: a dollar threshold per transaction (set against the engagement’s own materiality, not a house default), and a list of accounts that are automatically tier 3 regardless of amount — fixed assets, sales tax payable, intercompany, owner equity and draws, payroll clearing. Setting the bar after you’ve seen the results invites a taxonomy that quietly flatters whichever tool you already wanted to buy.

A tool with more tier-1 misses and zero tier-3 misses is better than a tool with a higher headline accuracy and two silent material errors. Design your review gates around tier-3 risk specifically — confidence thresholds, dollar thresholds, and account-level rules that force a human look regardless of how certain the model sounds.

Rules, your ledger’s AI, or your own agent — same test, different trade-offs

Built-in AI in QuickBooks or Xero
Zero build cost, already inside the system of record, improves without your involvement. But you can’t inspect why it coded something, can’t encode client-specific policy beyond what the UI exposes, and can’t easily run it against a held-out test set. Vendor capabilities move quickly — re-test after major releases.
Custom agent over your data via MCP
You control the instructions, the evidence it’s allowed to see, and the logging. An assistant connected through MCP can read the chart of accounts, prior-period coding, and vendor history before proposing an entry — and write nothing without approval. Costs real setup time, and you own the maintenance. The failure mode that matters most: because every row runs on the same shared instruction set, an agent can be confidently wrong in exactly the same style across hundreds of transactions. A spot-check of ten rows won’t surface that; a frozen answer key will.

Don’t skip the unglamorous third option: deterministic bank rules. For high-volume, unambiguous vendors, a rule is faster, free, and perfectly auditable. The honest pattern I’d argue for is rules first for the repetitive tail, AI for the ambiguous middle, human for judgment — roughly the argument in bank rules vs offshore vs AI agents. If you’re going the agent route, the mechanics of least-privilege access to the general ledger are covered in connecting an AI assistant to QuickBooks or Xero via MCP.

What rising model accuracy means for CPAs and staff

The “is AI replacing CPAs” question gets a cleaner answer when you grade by error class. Automation is absorbing the classification layer of accounting — matching, coding, routing, drafting. It is not absorbing the attestation, the judgment, or the signature. Xero’s own commentary on practices in Australia, reported via MarketScreener, frames firm AI adoption as growth rather than headcount reduction — one vendor’s framing, but consistent with what the work actually requires.

The professional obligation doesn’t move either. Under the AICPA Code of Professional Conduct, due care and supervision stay with the practitioner regardless of what produced the first draft; confirm the specifics with the AICPA and your state board of accountancy, since requirements vary. Practically: an agent can propose, a human signs.

Modeling the economics with your own numbers

Don’t start from a vendor ROI figure. Build it from your test results:

To see the shape of it, assume one client file at 600 transactions a month, 25 seconds to code a transaction manually, and 8 seconds to review a proposed code. That’s a 17-second delta × 600 = 10,200 seconds, or roughly 2.8 hours a month on that one file. Multiply by your own redeployment rate, repeat across the files that look like this one, then subtract tooling and build cost and any hours clawed back by tier-3 rework. Every input there is a placeholder I made up for the arithmetic — time your own staff on a real file before you believe any of it, and re-run the sums with their numbers.

If the result is marginal, that’s a real answer: plenty of firms should keep rules and a good offshore or in-house process, and the broader ranking in accounting automation examples by AI readiness is a reasonable place to look for higher-yield targets first.

Not sure where to start?

Get a free automation audit: we map your bookkeeping, month-end close, client onboarding, document collection, and AP/AR — and show you what's worth automating before you spend a dollar.

Get a free automation audit