AI searchAEOModels

Why ChatGPT, Gemini, Claude and Perplexity Disagree About Your Brand

Noma Team7 min read

Ask five AI engines which tools lead your category and you will get five different lists. Not slight reorderings of one list — genuinely different sets, with brands that appear on two and are absent from three.

This surprises people, because we inherited an intuition from search engines, where the top results for a query look broadly similar wherever you look. That intuition does not transfer.

Four reasons the answers differ

1. Different training data, different cut-offs

Each model absorbed a different slice of the web at a different time. A brand that launched, rebranded, or shifted category recently will be represented unevenly: current in the models trained since, stale or absent in the others. This is the cause most teams reach for, and it is real, but it is the least of the four.

2. Different retrieval, different index

When an engine goes and looks something up, what it searches is not the same thing its neighbour searches. The candidate set differs before a single word of the answer is written, and a page that is never fetched can never be cited.

3. Different question rewriting

Engines do not search your literal question. They rewrite it, and they rewrite it differently. “What should I use to track brand mentions in AI?” might become AI brand monitoring tools on one engine and track brand mentions ChatGPT Perplexity on another. Those two searches surface different pages, which is enough on its own to produce different answers.

4. Different composition rules

How many sources to consult, how much to hedge, whether to name commercial products at all: these are product decisions, and they vary. Some engines are noticeably more willing to recommend a specific vendor; others prefer to describe a category and leave the choice open. The same underlying knowledge produces a different answer depending on that posture.

Why the average is the wrong number

The temptation is to blend the five into one visibility score. It makes a cleaner dashboard and it destroys the information you needed.

A brand at 20% blended visibility could be:

  • 20% on all five. A presence problem, evenly spread. The fix is category-wide: more of the questions, better sources.
  • Near 100% on one, near zero on four. A concentration problem. Something one engine reads knows you well and the others have never encountered. The fix is specific and usually cheap: find that source, and get onto its equivalents elsewhere.

Those demand different work, cost different amounts, and take different lengths of time. The blended number cannot distinguish them, and a team acting on it alone is guessing which situation it is in.

What the gaps usually mean

In practice a few patterns recur often enough to be worth pattern-matching against.

Strong on retrieval-heavy engines, weak elsewhere. Your presence is live on the web but thin in training data. Typical of younger brands. It tends to correct itself as models retrain, and can be accelerated by getting accurately described on sources that are widely mirrored.

Strong on training-heavy engines, weak on retrieval. The reverse: you were well covered historically and your current web presence is not being surfaced. Often a structural problem on your own site rather than a reputation one.

Present everywhere but never first. Not a visibility problem at all. The category knows you and nothing distinguishes you inside the answer. No amount of technical AEO fixes this; it is a positioning question wearing a measurement costume.

Does this mean optimising per engine?

Mostly not, and it is worth saying plainly because the alternative is an enormous amount of wasted work.

The things that make a page quotable are the same everywhere: state the answer before the argument, keep claims self-contained, structure the page so it can be parsed, and be accurately represented on the third-party sources engines actually read. Do that once and it pays on all five.

Engine-specific work earns its keep in one situation: a large gap on one engine that persists across weeks and across many prompts. That is worth investigating, because it almost always resolves to a specific source rather than a general weakness.

How to measure it honestly

Two rules make the comparison mean something.

Same prompts, same day, fresh sessions. Personalised or conversation-carrying sessions contaminate the comparison. Each engine should get the identical question with no history.

Repeat before concluding. Answers vary run to run. A single query on a single day tells you almost nothing; the same query across a few weeks tells you where you actually stand. Most of what looks like a dramatic change on one day is variance, and reacting to it is how teams end up rewriting pages for no reason.