How accurate are AI-written Google Ads audits? We checked 682 numbers
Mostly right, and wrong exactly where it hurts. In 13 Google Ads audits written by AI, 94% of the numbers matched the export they came from. Almost every wrong one was a number the AI had worked out itself: a total, a sum, a share or a ratio. Nothing in the writing told the right numbers from the wrong ones.
How we tested it
We gave three tiers of Anthropic's Claude (Haiku, the small and fast one, then Sonnet and Opus) a Google Ads export and asked for the kinds of audit people actually ask for: a summary for a CEO, a pitch to win the account, a monthly update, a Slack message to a client, a second opinion on another agency's work. Thirteen audits in all. The models wrote without tools: no code and no spreadsheet, only the export pasted in.
- 13 audit requestsA CEO summary, a pitch, a Slack note and more
- 2 made-up exportsOne ends in Google's Total rows, one has none
- 3 Claude modelsHaiku, Sonnet and Opus, with no tools
- 682 numbersEvery figure the 13 audits stated
- Every number checkedWorked out again with code, blind to the model
The exports were made up, in Google's own format, so no real account is in this study. One is Acme Plumbing, a US account that spent $4,973.64 across 18 search terms, with a keyword report from the same account. The other is BrightPath, an Indian test-prep account that spent ₹144,775.20 across 22. Then we checked every number in every audit against its export: 682 numbers, each worked out again with code and labelled correct, wrong, or not something an export can confirm.
What we found
| Result | Numbers | Share |
|---|---|---|
| Matched the export | 638 | 93.5% |
| Wrong | 22 | 3.2% |
| Not something an export can confirm (targets, forecasts, benchmarks) | 22 | 3.2% |
1. Numbers copied from the export were almost always right
When an audit quoted a figure that sits in the export, such as a search term's cost, clicks or conversions, it was right in every case but one. The models read data well.
2. The errors lived in the numbers the AI worked out
21 of the 22 wrong numbers were figures the AI had to compute: account totals it added up itself, sums of groups of search terms, shares of spend and ratios. Totals were the sharpest case: read from Google's Total row, they were never wrong; added up by the AI, almost one in three was.
Two from the test set:
- "These nine terms burned ₹24,427." The audit listed only eight. Those eight do add up to ₹24,427, but the ninth search term with no conversions was left out of both the list and the total. All nine cost ₹25,287.
- "Performance Max campaigns took ₹7,115." The four Performance Max rows in the export add up to ₹7,826.85.
3. One wrong total spreads through the whole audit
The BrightPath export has no Total row, so the audit had to add up 22 rows itself, and it got the account's spend wrong. Everything built on that total went wrong with it: the cost per conversion (₹1,608 instead of ₹1,392), the share of spend wasted, and a claim that the account's average cost per conversion was 18 times its best search term's, which is 15.6 times with the real figures. One slip in the addition, and a chain of confident, wrong numbers.
4. The smallest model made every error
All 22 wrong numbers came from Haiku: 22 of its 201 numbers, about 1 in 9. Sonnet and Opus got none wrong across 481 numbers between them. This is a small test of one model family, so read it as a direction rather than a law: the faster and cheaper the model, the more you should check what it calculated.
5. Some numbers can't be checked at all
Another 22 numbers were not in the export to begin with: targets, forecasts and industry benchmarks. They aren't errors, but no export can confirm them either, and the bigger models wrote them too (8 of the 22). A good audit says where each one comes from. Most don't.
6. Nothing in the writing flags the wrong ones
The wrong numbers sat in the same confident sentences as the right ones: no hedging, no "about", no change of tone. Reading an AI audit, you cannot tell which 3% to doubt.
What this means if you use AI audits
- Trust copied figures more than computed ones. A term's cost is probably right. "These terms wasted" and "18 times cheaper" need checking.
- Re-add the totals first, especially when the export has no Total row. A wrong total quietly breaks every share and average built on it.
- Check every group sum against the terms it lists. Dropped and doubled rows are the most common slip.
- Ask for the working on anything that isn't in the export: a target, a forecast or a benchmark.
- Or let code check it. Our step-by-step method takes about ten minutes by hand.
The limits of this test
Thirteen audits, two made-up accounts and one model family. The models had no tools. We have since rerun the test with OpenAI's and Google's models, with and without a Python tool: the tool cut the errors from 14 to 2, when the model used it.
Questions
Can AI write a Google Ads audit?
Yes, and quickly. In our test the numbers AI copied from the export were almost always right. The numbers it worked out itself, like totals, group sums, shares and ratios, are the ones to check before you act.
Which numbers in an AI audit should I check first?
The account totals, then every sum of a group of search terms, then shares of spend and ratios such as cost per conversion. In our test 21 of the 22 wrong numbers were of these kinds.
Did the bigger AI models make mistakes?
Not in our test: Sonnet and Opus got none of their 481 numbers wrong. They did still write numbers no export can confirm, such as targets and forecasts, so those need a source either way.
How can I check an AI audit's numbers quickly?
Get the export the audit was written from, then check the totals, group sums and shares against it. OpenPPC does this automatically in your browser: it marks every number as traced, a wrong number or a wrong label.