NomaNoma
AI visibility

How to Track Your Brand Across ChatGPT, Gemini, Perplexity, Claude and Grok

A step-by-step method for tracking how five AI engines mention your brand: the prompt set, the schedule, what to record, and how to read the results.

Beckett Lindqvist7 min read

Most brands first check their AI visibility the same way: someone asks ChatGPT about the category, spots the brand in the answer, and shares a screenshot. The trouble is that the same question asked tomorrow, or asked on Gemini, may produce a different answer entirely.

Tracking is the discipline of turning those one-off checks into a measurement you can trust across ChatGPT, Gemini, Perplexity, Claude and Grok. This guide covers the method step by step, how to run it in a spreadsheet, where that breaks, and what a tool changes.

One check is an anecdote, not a measurement

A single answer tells you what one engine said once. Tracking means asking the same fixed questions on every engine, on a schedule, and recording the same fields each time, so that changes reflect the market rather than chance.

The whole method fits in seven steps:

  1. Build a buyer-worded prompt set and freeze it.
  2. Run the same prompts on each engine from clean sessions, on a schedule.
  3. Save the full text of every answer.
  4. Record mentions, competitors, position, sources and sentiment.
  5. Keep branded prompts in their own group.
  6. Read each engine on its own, then the combined view.
  7. Judge changes on trends over weeks.

The sections below take each step in turn, with the reason it exists.

How do you build a prompt set for tracking?

Write the questions a real buyer would ask, in the words they would use, and then stop changing them. The prompt set is the measurement instrument, and every number that follows inherits its flaws.

Cover the kinds of question buyers ask

Category prompts (“what is the best tool for [job]”) and comparison prompts (“[competitor] versus [competitor]”, “alternatives to [your brand]”) carry most of the signal. Add problem prompts that describe a symptom without naming the category, and a few branded prompts for accuracy checks.

Use buyer wording, not keywords

Take phrasing from sales calls, support tickets and community threads. Buyer questions are longer than search queries and often contain the constraint that decides the answer, such as a budget, a team size or an integration. The full approach is in which prompts to track.

Freeze the set

Every change breaks comparability with earlier runs. When the set has to grow, add prompts in dated batches and mark the date on your trend lines.

Run the same prompts, the same way, on every engine

Consistency between runs matters more than any individual run. Every difference in how a prompt is asked becomes a difference in the answer that looks like a real change.

Start from a clean session

Log out, or use an account with memory and personalization turned off. Your own history tells the engine what you care about, including your own brand, and that flatters the result.

Paste, do not retype

Small wording changes produce different answers. Keep the prompts in one place and paste them exactly.

Keep a schedule

Run the full set on fixed days. Answers vary between runs, and repetition is the only thing that turns that variation into a readable trend. Note the date and the engine for every run.

What should you record from each answer?

Save the full answer, then record five fields from it. Together they produce every metric worth reporting.

  • Full answer text. The raw material. It lets you re-score, check a surprising result, or see weeks later exactly when an engine started describing you differently.
  • Brand mentioned. Whether your brand is named. Across all answers on an engine, this gives visibility: the share of answers that name you.
  • Competitors named. Every other brand in the answer. This gives share of voice: your share of all brand mentions.
  • Position. Where you appear relative to the other brands named.
  • Cited sources. The URLs the engine links to. These explain why the answer says what it says.
  • Sentiment. How favorably you are described. Score it only when you are named and leave it blank otherwise, so absence never counts as neutral.

Keep branded prompts in their own group

Branded prompts name you almost every time, so blending them into the headline figure inflates visibility. Report them separately.

They still earn their place. Branded prompts are where you catch an engine quoting an old price, describing a discontinued feature or confusing you with another company. Read those answers for accuracy rather than counting mentions.

Read each engine before you read the total

Look at every engine on its own first, then at the combined view. An average across five engines can hide strength on one and absence on another, and those call for different work.

Engines differ in their training data, in how they retrieve current information, and in the sources they prefer, so split results are normal. The reasons are covered in why AI engines disagree. When engines split, three diagnostics help:

  • Gaps. Which prompts name competitors but not you, and on which engines?
  • Sources. Which pages does each engine cite on the prompts where you are missing?
  • Framing. Is sentiment consistent across engines, or is one repeating an outdated complaint?

How long before a trend means something?

Give it several weeks. Daily figures move with ordinary answer-to-answer variation, and only a sustained shift is a signal worth acting on.

When you ship something that should matter, such as a rewritten comparison page or a corrected third-party listing, mark the date. Judge its effect on the weeks that follow, per engine, not on the next morning.

How to run the method in a spreadsheet

A spreadsheet is a sound way to start, and doing it by hand teaches you what the answers actually look like. Set it up once and keep the structure fixed.

  1. Create a prompt tab listing every prompt with its type: category, comparison, problem or branded.
  2. Create a results tab with one row per prompt, per engine, per run. Columns: date, engine, prompt, prompt type, a link to the saved answer, brand named, position, competitors named, cited URLs and sentiment.
  3. Save each full answer in a shared folder and link it from its row.
  4. Build pivot tables for visibility and share of voice by engine and prompt type, filtering branded prompts out of the main view.
  5. Chart each engine separately over time.

Where manual tracking breaks down

The manual method fails on volume and consistency. Every prompt is multiplied by five engines and by every run, so the work grows faster than the prompt list.

  • Volume. A useful prompt set on five engines produces a large number of answers per run. Most teams respond by running less often, which weakens the trend.
  • Drift. Different people, sessions and days introduce small inconsistencies that look like real movement.
  • Subjective scoring. Position and sentiment judged by hand vary between whoever is doing the judging.
  • Skipped fields. Copying cited URLs is slow, so sources are usually the first column left empty, and they are the column that explains everything else.

How a tool automates the same method

A tracking tool runs the same steps on a schedule and does the recording. The method does not change; the manual labor does.

Noma, for example, tracks ChatGPT, Gemini, Perplexity, Claude and Grok. Every tracked prompt is asked on each engine once a day from a fresh session with no personalization, any prompt can be re-run on demand, and full answers are stored. Answers are parsed for brand and competitor mentions, position, cited URLs and sentiment, which is scored from 0 to 100 and only for answers that name the brand. Gap analysis, on Growth and above, lists the prompts where competitors are named and you are not. Self-serve plans track ChatGPT, Gemini, Perplexity and Claude, with Grok as a paid add-on; the free trial runs for 3 days with 25 prompts on all five engines. Plan details are on the pricing page.

What to do first

In rough order of return on effort:

  1. Write and freeze a prompt set in buyer wording, with branded prompts labeled separately.
  2. Run it on all five engines from clean sessions, and save the full answers from the very first run.
  3. Record mentions and competitors, then add position, sources and sentiment.
  4. Read each engine on its own and list the prompts where competitors appear and you do not.
  5. Keep the schedule for several weeks before drawing conclusions, by hand or with a tool.

None of this requires special access to the engines. It requires asking the same questions the same way, often enough, and writing down what comes back. If you want a quick starting picture before building anything, the free brand audit is a reasonable first look.

← Back to the blog