GPT vs Gemini for Google Ads audits: we checked 1,371 numbers
OpenAI's GPT-6 Luna got every number right in 12 Google Ads audits. Google's Gemini 3.8 Flash and Gemini 3.1 Pro made 16 errors between them, almost all in figures they worked out themselves, and half of them came from just two wrong totals. Giving the models a Python tool cut the errors from 14 to 2, but having the tool isn't the same as using it: Gemini 3.1 Pro ran it in only two of its six audits.
How we tested it
We gave three AI models the same six Google Ads audit requests, each with a made-up export pasted in, then ran every request again with a Python tool switched on. Then we checked every number the audits stated.
- 6 audit requestsFrom a search terms audit to a Slack note
- 2 made-up exportsOne ends in Google's Total rows, one has none
- 3 models, 2 waysExport pasted in, then again with a Python tool
- 36 auditsWith 1,371 numbers in them
- Every number checkedWorked out again with code, blind to the model
The six requests came from our first test, where Claude wrote 13 audits: a search terms audit, a table of keywords to pause and to scale, the audit section of an agency pitch, a rupee audit comparing Performance Max with Search, a short update for a founder, and a Slack note with five fixes. Both exports are made up, in Google's own format: Acme Plumbing, a US account whose reports end in Google's Total rows, and BrightPath, an Indian test-prep account with 22 rows and no Total row. No real account is in the test. We called the models through their APIs via OpenRouter, with default settings, in October 2026, and let them decide for themselves whether to use the tool. All 36 audits cost $1.02.
37 of the numbers were targets, forecasts or day counts that no export can confirm, so they're left out of the error rates.
What we found
1. The cheapest model made no mistakes
GPT-6 Luna, OpenAI's fast, low-cost model, got all 406 of its checkable numbers right, including the BrightPath totals it had to add up from 22 rows. It cost about a tenth of a cent per audit. Gemini 3.1 Pro cost nearly 50 times as much and made five errors.
In our first test it went the other way: Anthropic's smallest model, Claude Haiku, made every error. You can't tell from the price or the size which model will slip.
2. The errors were in numbers the models worked out
Figures read straight from a row of the export were right every time but two, and those two were real figures with a false claim attached. The other 14 errors were all numbers the models had to work out: totals, sums, shares, costs per conversion and savings. That matches our first test, where 21 of the 22 wrong numbers were worked out by the model.
3. Adding up the export is where it went wrong
BrightPath's export has no Total row, so without tools the models had to add 22 rows in their heads. Three of the seven audits that stated the account's spend got it wrong. With the Python tool, all nine that stated it got it right.
4. One wrong total takes other numbers with it
Eight of the 16 wrong numbers came from Gemini Flash's two wrong totals. Here is the first: every figure built on the total moved with it. Each was correct arithmetic on a wrong starting figure, so nothing looked out of place.
5. Python helps, if the model runs it
Gemini 3.8 Flash ran code 37 times and went from 10 wrong numbers to 1. GPT-6 Luna ran it 15 times and stayed at zero. Gemini 3.1 Pro ran it twice in six audits and wrote the rest in its head, so its drop from 4 errors to 1 is closer to a second try than a test of the tool.
Having a code tool switched on is not the same as using it.
6. Code doesn't check what the model says in words
The two errors that got through with Python were claims written around the numbers, not numbers the code worked out. Gemini 3.1 Pro made the same "highest" mistake without tools, too.
an account-high $7.80 CPCThe highest cost per click in the export was $7.95, for drain cleaning.Gemini 3.8 Flash, with Python
filter out over $1,100 in historically inefficient spendThe search terms on its own negative keyword list cost $880.70.Gemini 3.1 Pro, with Python
Claims like "the highest", "the cheapest" and "would save" need checking whatever the model ran.
What this means if you use ChatGPT or Gemini for audits
- Turn on code, and check it ran. Ask for the calculation behind every total. If there's no working, treat the totals as unchecked.
- Re-add the total yourself when the export has no Total row. It's the number every share and cost per conversion is divided by.
- Check the claims in words: "the highest", "the lowest", "would save". Code that adds up rows doesn't test them.
- Don't pick a model and stop checking. The best result here, no errors in 406 numbers, is one model on 12 audits. Another prompt, export or model update can change it.
The limits of this test
36 audits, two made-up accounts, three models, one run of each. Six audits per model and setting is small: a difference of a few errors between two of them could be chance. We used the APIs, not the ChatGPT or Gemini apps, which add their own instructions and may choose a different model. And we didn't test OpenAI's flagship, GPT-6 Astra, in this round. Read the results as a direction, and check the numbers whichever model you use.
Questions
Is ChatGPT or Gemini better for Google Ads audits?
In our test, OpenAI's GPT-6 Luna made no errors in 406 numbers, while Google's Gemini 3.8 Flash and Gemini 3.1 Pro made 16 between them, mostly in totals they added up without code. It's a small test, so treat it as a direction and check the numbers whichever you use.
Does Python or a code interpreter make AI audits accurate?
It helps when the model uses it: errors fell from 14 to 2 in our test. But one model mostly ignored the tool, and claims written in words, like "the highest" or "would save", slipped through either way.
Is this the same as using the ChatGPT or Gemini apps?
Not exactly. We used the models through their APIs with default settings, the same way for each. The apps add their own instructions and may pick a different model, so results there can differ.
Which numbers should I check first in an AI audit?
The account totals, especially when the export has no Total row. Then every number divided by them, like cost per conversion and shares of spend, and anything described as the highest, the lowest or a saving.