Skip to content
Research

How accurate are AI-written Google Ads audits? We checked 682 numbers

Mostly right, and wrong exactly where it hurts. In 13 Google Ads audits written by AI, 94% of the numbers matched the export they came from. Almost every wrong one was a number the AI had worked out itself: a total, a sum, a share or a ratio. Nothing in the writing told the right numbers from the wrong ones.

682numbers checked
94%matched the export
3%were wrong
21 of 22wrong numbers were worked out by the AI

How we tested it

We gave three tiers of Anthropic's Claude (Haiku, the small and fast one, then Sonnet and Opus) a Google Ads export and asked for the kinds of audit people actually ask for: a summary for a CEO, a pitch to win the account, a monthly update, a Slack message to a client, a second opinion on another agency's work. Thirteen audits in all. The models wrote without tools: no code and no spreadsheet, only the export pasted in.

How the test worked13 audits, 682 numbers, every one worked out again with code
  1. 13 audit requestsA CEO summary, a pitch, a Slack note and more
  2. 2 made-up exportsOne ends in Google's Total rows, one has none
  3. 3 Claude modelsHaiku, Sonnet and Opus, with no tools
  4. 682 numbersEvery figure the 13 audits stated
  5. Every number checkedWorked out again with code, blind to the model

The exports were made up, in Google's own format, so no real account is in this study. One is Acme Plumbing, a US account that spent $4,973.64 across 18 search terms, with a keyword report from the same account. The other is BrightPath, an Indian test-prep account that spent ₹144,775.20 across 22. Then we checked every number in every audit against its export: 682 numbers, each worked out again with code and labelled correct, wrong, or not something an export can confirm.

What we found

ResultNumbersShare
Matched the export63893.5%
Wrong223.2%
Not something an export can confirm (targets, forecasts, benchmarks)223.2%
Every number, one square eachRed ones were wrong, grey ones weren't something an export can confirm. On a computer, point at a red square to see it.
Matched the exportWrongNot in the data
Claude Haiku: 22 wrong of 201. Claude Sonnet: none wrong of 338. Claude Opus: none wrong of 143.

1. Numbers copied from the export were almost always right

When an audit quoted a figure that sits in the export, such as a search term's cost, clicks or conversions, it was right in every case but one. The models read data well.

2. The errors lived in the numbers the AI worked out

21 of the 22 wrong numbers were figures the AI had to compute: account totals it added up itself, sums of groups of search terms, shares of spend and ratios. Totals were the sharpest case: read from Google's Total row, they were never wrong; added up by the AI, almost one in three was.

How often each kind of number was wrongAll 660 numbers that were claims about the export, by where the figure came from

Two from the test set:

  • "These nine terms burned ₹24,427." The audit listed only eight. Those eight do add up to ₹24,427, but the ninth search term with no conversions was left out of both the list and the total. All nine cost ₹25,287.
  • "Performance Max campaigns took ₹7,115." The four Performance Max rows in the export add up to ₹7,826.85.
"These nine terms burned ₹24,427"The audit's list against every search term in the export with no conversions. Each block is one term; on a computer, point at one to see it.
The audit listed eight terms, and they add up to ₹24,427. The ninth, "how to become ias" at ₹860 (in red), was missing from both the list and the total.

3. One wrong total spreads through the whole audit

The BrightPath export has no Total row, so the audit had to add up 22 rows itself, and it got the account's spend wrong. Everything built on that total went wrong with it: the cost per conversion (₹1,608 instead of ₹1,392), the share of spend wasted, and a claim that the account's average cost per conversion was 18 times its best search term's, which is 15.6 times with the real figures. One slip in the addition, and a chain of confident, wrong numbers.

One wrong total, three more wrong numbersFrom Claude Haiku's BrightPath audit. Struck through: what it said. Green: what the export says.
Total spend₹167,272₹144,775.20
Cost per conversion₹1,608₹1,392.07
Share of spend wasted14.6%17.47%
Average cost per conversion against the best term's18 times15.6 times
The same audit also put the account's clicks at 1,697 instead of 2,099, and its wasted spend at ₹24,427 instead of ₹25,286.85.

4. The smallest model made every error

All 22 wrong numbers came from Haiku: 22 of its 201 numbers, about 1 in 9. Sonnet and Opus got none wrong across 481 numbers between them. This is a small test of one model family, so read it as a direction rather than a law: the faster and cheaper the model, the more you should check what it calculated.

5. Some numbers can't be checked at all

Another 22 numbers were not in the export to begin with: targets, forecasts and industry benchmarks. They aren't errors, but no export can confirm them either, and the bigger models wrote them too (8 of the 22). A good audit says where each one comes from. Most don't.

6. Nothing in the writing flags the wrong ones

The wrong numbers sat in the same confident sentences as the right ones: no hedging, no "about", no change of tone. Reading an AI audit, you cannot tell which 3% to doubt.

What this means if you use AI audits

  • Trust copied figures more than computed ones. A term's cost is probably right. "These terms wasted" and "18 times cheaper" need checking.
  • Re-add the totals first, especially when the export has no Total row. A wrong total quietly breaks every share and average built on it.
  • Check every group sum against the terms it lists. Dropped and doubled rows are the most common slip.
  • Ask for the working on anything that isn't in the export: a target, a forecast or a benchmark.
  • Or let code check it. Our step-by-step method takes about ten minutes by hand.
OpenPPC checks any audit against the export it came from. Paste the audit, add the export, and every number comes back traced, a wrong number (with the real figure) or a wrong label. It runs in your browser and uploads nothing. Open the app

The limits of this test

Thirteen audits, two made-up accounts and one model family. The models had no tools. We have since rerun the test with OpenAI's and Google's models, with and without a Python tool: the tool cut the errors from 14 to 2, when the model used it.

Questions

Can AI write a Google Ads audit?

Yes, and quickly. In our test the numbers AI copied from the export were almost always right. The numbers it worked out itself, like totals, group sums, shares and ratios, are the ones to check before you act.

Which numbers in an AI audit should I check first?

The account totals, then every sum of a group of search terms, then shares of spend and ratios such as cost per conversion. In our test 21 of the 22 wrong numbers were of these kinds.

Did the bigger AI models make mistakes?

Not in our test: Sonnet and Opus got none of their 481 numbers wrong. They did still write numbers no export can confirm, such as targets and forecasts, so those need a source either way.

How can I check an AI audit's numbers quickly?

Get the export the audit was written from, then check the totals, group sums and shares against it. OpenPPC does this automatically in your browser: it marks every number as traced, a wrong number or a wrong label.