Skip to content
Research

GPT vs Gemini for Google Ads audits: we checked 1,371 numbers

OpenAI's GPT-6 Luna got every number right in 12 Google Ads audits. Google's Gemini 3.8 Flash and Gemini 3.1 Pro made 16 errors between them, almost all in figures they worked out themselves, and half of them came from just two wrong totals. Giving the models a Python tool cut the errors from 14 to 2, but having the tool isn't the same as using it: Gemini 3.1 Pro ran it in only two of its six audits.

1,371numbers checked in 36 audits
0wrong from GPT-6 Luna
14 → 2wrong numbers, without and with Python
8 of 16errors came from two wrong totals

How we tested it

We gave three AI models the same six Google Ads audit requests, each with a made-up export pasted in, then ran every request again with a Python tool switched on. Then we checked every number the audits stated.

How the test worked36 audits, 1,371 numbers, every one worked out again with code
  1. 6 audit requestsFrom a search terms audit to a Slack note
  2. 2 made-up exportsOne ends in Google's Total rows, one has none
  3. 3 models, 2 waysExport pasted in, then again with a Python tool
  4. 36 auditsWith 1,371 numbers in them
  5. Every number checkedWorked out again with code, blind to the model

The six requests came from our first test, where Claude wrote 13 audits: a search terms audit, a table of keywords to pause and to scale, the audit section of an agency pitch, a rupee audit comparing Performance Max with Search, a short update for a founder, and a Slack note with five fixes. Both exports are made up, in Google's own format: Acme Plumbing, a US account whose reports end in Google's Total rows, and BrightPath, an Indian test-prep account with 22 rows and no Total row. No real account is in the test. We called the models through their APIs via OpenRouter, with default settings, in October 2026, and let them decide for themselves whether to use the tool. All 36 audits cost $1.02.

37 of the numbers were targets, forecasts or day counts that no export can confirm, so they're left out of the error rates.

What we found

Every number the audits stated, one square eachRed ones were wrong: 16 of 1,334, in the order the audits wrote them. On a computer, point at a red square to see what it said.
Matched the exportWrong
GPT-6 Luna: 0 wrong of 406. Gemini 3.8 Flash: 10 wrong of 278 with the export pasted in, 1 of 341 with Python. Gemini 3.1 Pro: 4 of 156, then 1 of 153.

1. The cheapest model made no mistakes

GPT-6 Luna, OpenAI's fast, low-cost model, got all 406 of its checkable numbers right, including the BrightPath totals it had to add up from 22 rows. It cost about a tenth of a cent per audit. Gemini 3.1 Pro cost nearly 50 times as much and made five errors.

Price didn't predict accuracyAverage API cost per audit, and wrong numbers across each model's 12 audits

In our first test it went the other way: Anthropic's smallest model, Claude Haiku, made every error. You can't tell from the price or the size which model will slip.

2. The errors were in numbers the models worked out

Figures read straight from a row of the export were right every time but two, and those two were real figures with a false claim attached. The other 14 errors were all numbers the models had to work out: totals, sums, shares, costs per conversion and savings. That matches our first test, where 21 of the 22 wrong numbers were worked out by the model.

The 16 wrong numbers, by kindOne square per wrong number. On a computer, point at one to see it.
8 from Gemini Flash's two wrong totals, and the figures built on them
4 other sums that were off, one by a single rupee
2 savings claims that didn't add up
2 right figures wrongly called the highest

3. Adding up the export is where it went wrong

BrightPath's export has no Total row, so without tools the models had to add 22 rows in their heads. Three of the seven audits that stated the account's spend got it wrong. With the Python tool, all nine that stated it got it right.

BrightPath's total spend, as each audit stated it16 audits gave the figure. Amounts are shown as the audits wrote them, in Indian digit grouping: ₹1,44,775.20 is ₹144,775.20.
Export pasted in: 4 of 7 right
GPT-6 Luna
GPT-6 Luna
Gemini 3.8 Flash
Gemini 3.8 Flash
Gemini 3.8 Flash
Gemini 3.1 Pro
Gemini 3.1 Pro
With a Python tool: 9 of 9 right
GPT-6 Luna
GPT-6 Luna
GPT-6 Luna
Gemini 3.8 Flash
Gemini 3.8 Flash
Gemini 3.8 Flash
Gemini 3.1 Pro
Gemini 3.1 Pro
Gemini 3.1 Pro
Matched the real totalWrong

4. One wrong total takes other numbers with it

Eight of the 16 wrong numbers came from Gemini Flash's two wrong totals. Here is the first: every figure built on the total moved with it. Each was correct arithmetic on a wrong starting figure, so nothing looked out of place.

One wrong total, four more wrong numbersFrom Gemini 3.8 Flash's audit with the export pasted in. Struck through: what it said. Green: what the export says.
Total spend₹1,35,474.90₹1,44,775.20
Blended cost per conversion₹1,302.64₹1,392.07
Wasted spend as a share of total18.67%17.47%
Search spend₹1,27,648.05₹1,36,948.35
Search cost per conversion₹1,239.30₹1,329.60
The wasted spend itself, ₹25,286.85, was right. Divided by the wrong total, its share came out too high. The same audit also put Search's impressions at 20,445 instead of 32,465.

5. Python helps, if the model runs it

Gemini 3.8 Flash ran code 37 times and went from 10 wrong numbers to 1. GPT-6 Luna ran it 15 times and stayed at zero. Gemini 3.1 Pro ran it twice in six audits and wrote the rest in its head, so its drop from 4 errors to 1 is closer to a second try than a test of the tool.

Wrong numbers without and with a Python toolEach model's six audits per setting, next to how often it actually ran code
Export pasted inWith a Python tool

Having a code tool switched on is not the same as using it.

6. Code doesn't check what the model says in words

The two errors that got through with Python were claims written around the numbers, not numbers the code worked out. Gemini 3.1 Pro made the same "highest" mistake without tools, too.

an account-high $7.80 CPCThe highest cost per click in the export was $7.95, for drain cleaning.Gemini 3.8 Flash, with Python
filter out over $1,100 in historically inefficient spendThe search terms on its own negative keyword list cost $880.70.Gemini 3.1 Pro, with Python

Claims like "the highest", "the cheapest" and "would save" need checking whatever the model ran.

What this means if you use ChatGPT or Gemini for audits

  • Turn on code, and check it ran. Ask for the calculation behind every total. If there's no working, treat the totals as unchecked.
  • Re-add the total yourself when the export has no Total row. It's the number every share and cost per conversion is divided by.
  • Check the claims in words: "the highest", "the lowest", "would save". Code that adds up rows doesn't test them.
  • Don't pick a model and stop checking. The best result here, no errors in 406 numbers, is one model on 12 audits. Another prompt, export or model update can change it.
OpenPPC checks any audit against the export it came from, whether ChatGPT, Gemini, Claude or a person wrote it. On these 36 audits it flagged 10 of the 16 wrong numbers, including all 8 from the two wrong totals, and showed 24 of the 1,318 correct numbers (1.8%) as worth a second look. It missed the two one-rupee slips, both "highest" claims and both savings claims. Check an audit now

The limits of this test

36 audits, two made-up accounts, three models, one run of each. Six audits per model and setting is small: a difference of a few errors between two of them could be chance. We used the APIs, not the ChatGPT or Gemini apps, which add their own instructions and may choose a different model. And we didn't test OpenAI's flagship, GPT-6 Astra, in this round. Read the results as a direction, and check the numbers whichever model you use.

Questions

Is ChatGPT or Gemini better for Google Ads audits?

In our test, OpenAI's GPT-6 Luna made no errors in 406 numbers, while Google's Gemini 3.8 Flash and Gemini 3.1 Pro made 16 between them, mostly in totals they added up without code. It's a small test, so treat it as a direction and check the numbers whichever you use.

Does Python or a code interpreter make AI audits accurate?

It helps when the model uses it: errors fell from 14 to 2 in our test. But one model mostly ignored the tool, and claims written in words, like "the highest" or "would save", slipped through either way.

Is this the same as using the ChatGPT or Gemini apps?

Not exactly. We used the models through their APIs with default settings, the same way for each. The apps add their own instructions and may pick a different model, so results there can differ.

Which numbers should I check first in an AI audit?

The account totals, especially when the export has no Total row. Then every number divided by them, like cost per conversion and shares of spend, and anything described as the highest, the lowest or a saving.