Research··6 min read

How to track AI visibility without fooling yourself

Most AI-visibility numbers are wrong — and the worst part is they don't look wrong. Five traps, and the measurement discipline that survives them.

By Pascal Moyon

AI visibility tracking: the number is easy to get wrong

Every week another tool launches promising to score your visibility in AI answers, and most of them will hand you a number that is wrong — not by a little, and not in a way you can see. I spent years building measurement into marketing P&Ls, and the rule I trust most is simple: a number you cannot trust is worse than no number, because it moves budget with false confidence.

AI visibility is the most trap-laden thing I have measured. The traps do not throw errors; they return clean figures that read as truth. Here are the five that matter, and the discipline that survives each one.

Trap 1 — you are seeing barely 40% of the surface

Google now loads its AI Overviews after the page renders. A standard capture arrives, finds an empty shell where the answer will be, and records "no AI Overview here". No error. No warning. A market that does not answer, apparently.

In a camera-and-imaging market our capture logs put the true answer-layer presence at 43%, while a standard capture saw 17% — barely 40% of what was actually there. Every presence figure, every citation share, every "the AI doesn't answer our category" conclusion built on the standard capture is an undercount of unknown size.

The discipline: request the asynchronous answer layer explicitly, every collection. And when a presence number drops hard between two runs, suspect the capture before you believe the market moved — a missing async request looks exactly like a market that went quiet.

Trap 2 — the number moves when nothing changes

Ask the same questions twice on the same morning, with everything identical, and in our capture logs 15 to 20% of them disagree on whether an AI answer even appears. Google decides at serve time — experiment bucket, load, the user's own state. Nothing about you changed; the number did.

This is the difference between weather and climate, and most trackers report the weather. A dashboard announcing "AI visibility up 12% this week" is, more often than not, inside the noise.

The discipline: presence is a rate over repeated captures, never a property of a single keyword. Do not report movement below the noise band — for us, up to 20% on presence — as a change; state the band beside the number every time. And prefer domain-level citation share for trend, because it averages over many answers and holds far steadier than presence.

Trap 3 — the surfaces are not one surface

Google AI Overviews, ChatGPT and Gemini do not cite the same world, and pooling them does not average the truth — it reverses it.

They cite different worlds. In a UK health market:

  • Google AI Overviews lead with YouTube — the single most-cited domain, and close to invisible in the chatbots;
  • ChatGPT leans on Reddit — well over half of its community citations there;
  • Gemini is the only one of the three that reaches for X;
  • People Also Ask is a fourth surface again, and the near-universal one — it appears on almost every result that carries an AI Overview, and on plenty that do not.

A channel that looks dead on one surface leads on another. So a single "AI visibility" score, pooled across surfaces, is not a summary — it is a mistake with a decimal point.

The discipline: measure each surface on its own column and never pool them. A number without a surface named beside it should not leave the building.

Trap 4 — cited is not the same as consulted

In our health-market capture sets, roughly half the sources under an AI answer were not cited to the reader at all — they are grounding the model quietly reached for, not links the answer pointed anyone towards. Counting the two together doubles your apparent visibility and hides whether you actually reached a human.

There is a related tax: the platforms cite themselves. On the Gemini surface, Google's own units were the top "earned" source until we excluded them — leave them in and you are benchmarking against Google, not your rivals.

The discipline: separate what the answer pointed the reader to from what the model merely consulted, and strip the platform's self-citations before you rank anyone. And count on the answer grain — did the model name this domain in this answer — not the URL grain, which simply rewards whoever deep-links the most pages.

Trap 5 — your question list is not the market

This is the quiet one, and the most common. Hand-pick fifty questions you think matter, run them, and you have measured your own question list — not your market. Weight it toward your own brand name and the earned number flatters you, because a model knowing your own story is not the same as it recommending you on a neutral question.

The discipline: the question set is the instrument, so it has to be built like one — enumerated from the market's own structure rather than chosen by hand, weighted so the generic (non-brand) questions carry the earned metric, and then frozen. The moment you edit the questions between runs, the time series breaks and you are back to measuring the instrument. Earned visibility is share on the generic questions; the branded ones answer a different question entirely.

Some take-aways

A trustworthy AI-visibility number is not hard because the maths is hard. It is hard because five things go wrong quietly:

  • the surface hides — you see barely 40% of it;
  • the number wobbles — 15 to 20%, with nothing changed;
  • the channels disagree — never pool them;
  • half the sources never reached the reader;
  • and the easy way to build the test is the one that measures yourself.

Get those five right and the number means something. Get any one wrong and it reads exactly the same — which is the danger.

None of this is an argument against tracking — quite the opposite. It is an argument for tracking it properly, on the discipline above, so the movement you report is real. That is the first half of the picture. The second is why a source gets cited in the first place — which is where your Google rank turns out to matter far more than "SEO is dead" would have you believe.

If you want the framework this measurement feeds, start with the AI-visibility ladder; for the whole-market view, see what market intelligence actually is, or explore the platform.


Method note: figures are Theia measurements across live consumer markets (a UK private-healthcare market and a global camera-and-imaging brand's category), captured with Google's asynchronous answer layer requested explicitly and reported as a rate over repeated captures. Surface-level results are never pooled. We publish the result and the reasoning, never the method that produces them.

Frequently asked

Are AI visibility tools accurate?
Most under-report without telling you. Google loads its AI Overviews after the page, so a standard capture stores an empty shell and reads it as 'no AI answer'. In our capture logs a market with 43% true answer-layer presence showed only 17% unless the asynchronous layer was requested explicitly — the tool was seeing barely 40% of the surface, with no error.
Why do my AI visibility numbers keep changing?
Because presence is decided at serve time. Ask the same questions twice on the same morning, everything identical, and 15-20% of them flip on whether an AI answer even appears. Treat presence as a rate over repeated captures, never a fixed property, and do not call a sub-20% swing a change.
Which AI surface should I measure — ChatGPT, Gemini or Google AI Overviews?
All three, separately, and never pooled — they cite different worlds. In a UK health market, YouTube is the single most-cited domain in Google's AI Overviews and close to invisible in the chatbots, while Reddit supplies well over half of ChatGPT's community citations. A channel that looks dead on one surface leads on another, so a number without a surface named beside it is meaningless.

Subscribe for the next piece.

Bi-weekly research on structured market intelligence. Free.