← All insights
AI and trust 29 Aug 2026 · 9 min read

AI mistakes in accounting: the errors that look right

By the Navigator team ·

The dangerous AI failures in accounting are the quiet ones: the total that is 3% off, the "last quarter" that turned out to be the calendar quarter when your year ends in June, the vendor category that was right in March and different in July, and the summary bucket that quietly absorbed three product lines. In the most recent independent benchmark, no general-purpose model passed more than 66% of 101 real accounting tasks, and every one failed at least a third. The errors that survive are the ones that look like a right answer.

What the benchmarks measure

DualEntry Labs published its 2026 Accounting AI Benchmark on 27 August 2026: 101 tasks in eight categories, run against 19 models. The top score was Gemini 3.1 Pro at 66.0%. GPT-4 scored 19.8%. No model exceeded 70%, every model failed at least one-third of the tasks, and the authors wrote that "confidence without accuracy can be more difficult to detect than an obvious gap." That is a test of doing accounting, not of talking about it.

The Digits benchmark, reported by Insightful Accountant on 5 June 2026, is narrower: 2,000 transactions to categorize, 13 frontier models. The best general model got 80.8% right one-shot and 86.8% with a checking layer around it; Digits' own purpose-built system scored 97.8%. Categorization is the easiest accounting task there is, and one in five wrong is what a general model produces on it.

A hallucination is an output the model presents as fact that is not supported by its input. Vectara's leaderboard, as of 11 May 2026, measured the best model at a 1.8% hallucination rate on plain document summaries. That number is low and it is also the wrong test for your books, where the input is a P&L and the failure is a sum, a period or a category, not an invented sentence.

The six AI failures in accounting that look right

The plausible total. You paste a P&L and ask for total operating expenses. The answer is $1,238,400. The QuickBooks report says $1,276,900. The difference is $38,500, about 3%, because a row wrapped in the paste and the model summed what it could see. Nobody checks a number that is close.

The wrong period or basis. On 2 October you ask how "this quarter" went and get the quarter ending 30 September, when your fiscal year ends in June and your bank wants fiscal quarters. Or you ask about profit and the model reads an accrual report while your tax return and your gut both run on cash. Cash basis books a sale when the money arrives; accrual books it when you invoice. The two can differ by a month of revenue and neither is wrong.

The silent bucket. CFO Impulse tested ChatGPT-4o on a set of financials and found an "Other" bucket that had absorbed three industries, discovered only after five rounds of prompting. A summary that is 90% right and 10% "Other" has hidden whatever it could not classify, and the hidden part is usually the interesting part.

The confident category. A $1,840 monthly truck payment coded to auto expense, principal and interest together. On a P&L-only reading that is $22,080 a year of expense when $18,120 of it is principal that belongs on the balance sheet. The model had no way to know, and did not say so. Walton CPA's May 2026 list of AI error types includes classification that changes between periods, so the same vendor can be coded three ways in a year.

The rule that used to be true. Mileage rates, filing thresholds, the point at which an S-corp election pays. Walton CPA flags premature S-corp recommendations made from a single profit figure. A model trained on last year's rules answers with last year's rules, in the same tone it uses for this year's.

The multi-company double count. Your service company bills your install company a $5,000 monthly management fee. Give a model both P&Ls and ask for total revenue, and it reports $4,160,000 when customers paid $4,100,000. The gap is exactly $60,000, twelve fees, counted once as income in one file and again inside the other's total. The fix is an elimination, covered in intercompany eliminations explained, and no general model knows to make it unless you tell it the two companies are yours.

Why the errors take this shape

A language model produces text by predicting the next likely token, and a likely number is one that fits the pattern of the numbers around it. $1,238,400 fits. It does not know the total was wrong, because it was not adding; it was writing something that reads like an addition.

The second reason is that the same prompt can take different paths. Thinking Machines ran one prompt 1,000 times at temperature zero in September 2025 and got 80 different outputs. A 2025 paper on arXiv found near-perfect reproducibility on simple classification and "greater variability" on complex analysis, which is a fair description of the difference between coding a receipt and explaining a margin. Why ChatGPT gives different answers to the same question goes further into it. The practical point is that an answer you got once is one draw from a distribution, and the second draw is a check.

Purpose-built tools narrow this without removing it. VentureBeat's report on Intuit's accounting agent described categorization accuracy up 20 points and users still complaining, which is what an 80% tool improved to a 90% tool feels like from inside a set of books.

What executives believe, and what audits find

Workiva's 2026 Midyear Executive Benchmark Survey, reported by CPA Practice Advisor on 13 August 2026, asked 2,272 executives, 847 of them C-level. 84% had at least some confidence in AI accuracy without human review. 26% said an internal audit had found AI errors that reached the board or an external audience. Only 11% said their data quality was sufficient for AI.

Those are large companies with audit functions, and an owner with three QuickBooks files and an outside bookkeeper has no internal audit, so the errors never get found. The confidence is probably the same.

The five checks

Tie every figure the AI gives you to a QuickBooks report total for the same company and the same period. If it does not match to the dollar, the AI figure is wrong, whatever the explanation.

Before you ask anything about the numbers, ask which company and which period the model is looking at, and what basis. Most wrong answers we see were right answers to a different question.

Ask for the source entry. A figure that cannot be traced to a transaction is a figure that was composed, not computed. If the tool cannot show the entry, treat the number as a draft.

Ask the same question twice, in a fresh chat, and compare. Two different answers tell you the question was hard for the model, which is the thing you most wanted to know.

Never let the same tool compute and check. Common advice is to ask the model to double-check its work, and we think that is close to useless: it is one more draw from the same distribution, delivered with the same confidence. A different tool, a person or the report itself is the check.

Navigator takes a different approach to the same problem. Every number it shows is computed from the QuickBooks Online ledger, read-only, and the AI's job is to explain the number, not produce it. Any figure opens to the company and the entry it came from, so the first check above is one click rather than a spreadsheet, and the management fee in the example is removed by the intercompany elimination on the Pro plan. It cannot tell you whether the entry itself was coded correctly; that is still your bookkeeper. The trial runs 30 days with no card.

What it should never do alone, and what it is good at

No AI should post an entry, file a return or a sales tax report, touch payroll, or delete a transaction without a person looking first. Every one of those is a write to a system of record, and the checks above are for reading. If you connect a model to your books, the read-only connector for ChatGPT and Claude is the version that cannot post, and whether it is safe to upload statements at all is a separate question with a longer answer.

Where it is good is where the cost of a wrong answer is low and the check is easy. Explaining what a line on the balance sheet means. Drafting the note to your bookkeeper that asks why insurance doubled in July. Reading twelve months of expenses and pointing at the one that looks odd, which you then open in QuickBooks and verify. The prompts that work for small business finances are mostly of that kind.

What the checks cannot catch is a rule you did not know had changed, applied confidently to your situation. Only the CPA catches that. The model does not know when it is out of date either.

Questions owners ask

Does ChatGPT make up numbers?

Sometimes, and more often it produces a number that is nearly right. Vectara's hallucination leaderboard, as of May 2026, put the best model's rate on plain summaries at 1.8%. On accounting tasks the failures are higher and quieter: a column summed to a plausible total, a category that changed between months. Check the total against the report, every time.

How accurate is AI at accounting?

Less than it sounds. DualEntry's August 2026 benchmark ran 19 models on 101 real accounting tasks; the top score was 66.0% and none passed 70%. On transaction categorization the Digits benchmark from June 2026 found the best general model at 80.8% one-shot. Good for a first pass, not good enough to post from.

What are the most common AI mistakes in bookkeeping?

The ones that survive review: a total slightly off the report, the wrong period or basis, a summary bucket that absorbs several lines, a confident but wrong category, a rule that changed last year, and, across several companies, a fee counted as revenue in one file and expense in the other. None of them look like a mistake on the page.

How do I check an AI's answer about my financials?

Tie the figure to a QuickBooks report total for the same company and period. Ask which entries it came from and open one. Ask the same question a second time in a fresh chat and compare. If the answer drives a decision, have a person or a different tool recompute it. The tool that produced the number should not be the one checking it.

Can AI replace my bookkeeper?

Not on current evidence. The best model in the most recent independent benchmark failed a third of real accounting tasks, and the errors it makes are the kind a bookkeeper would catch on sight. It can explain a line, draft the question to your bookkeeper, and point at an anomaly you then verify. Posting, filing and payroll stay with a person.

The wider question of what a general model can and cannot do with your books is in can ChatGPT do my accounting. The reason two identical questions get two answers is in why ChatGPT gives different answers to the same question. The prompts worth using once you know the checks are in ChatGPT prompts for small business finances.

If you would rather ask questions of numbers that come from the ledger and open to the entry, the trial connects read-only in about fifteen minutes with no card: navigatorhq.ai.

Navigator Insights by email

Get the next post by email.

One email when a new post goes up: cash, lenders, running several companies and AI. Unsubscribe in one click.

By subscribing you agree to our Privacy Policy.

Published . Last updated . Reviewed by a CFO on the Navigator team.

See every company you own in one place. Every morning.

Navigator connects to QuickBooks Online read-only, consolidates your entities, and answers the questions in these posts on your own numbers. Thirty days free, no card.

Start free trial Run the free Health Check