A tracked prompt set is not a list of keywords. It is a measurement instrument, and like any instrument it can be built to tell you the truth or built to tell you what you want to hear.
Almost every misleading AI visibility number comes from the prompt set rather than from the measurement.
The four kinds worth tracking
Category prompts
Best tools for X. What should I use to do Y. Top platforms for Z.
The core of the set. These are where a buyer with no shortlist gets one, and where being absent costs you a place on it. If you track nothing else, track these.
Comparison prompts
X versus Y. Alternatives to X. Is X better than Y for Z.
Late-stage and high-intent. Someone asking these has a shortlist already, and the answer can remove you from it. Include comparisons against competitors, and against yourself — “alternatives to [you]” is a prompt you want to know the answer to.
Problem prompts
How do I stop X happening. Why is Y so slow. What do I do about Z.
The most under-tracked kind and often the most valuable. The buyer describes a symptom without knowing the category exists. Whoever gets named here shapes the shortlist before it forms, and these prompts are frequently uncontested because competitors are all tracking the obvious category terms instead.
Branded prompts
What is [you]. How much does [you] cost. Does [you] support Z.
Keep a handful, and understand what they are for. You will be named in nearly all of them, so they inflate a blended visibility figure. Their value is accuracy: this is where you catch the engine confidently stating a wrong price or a missing feature.
Sizing the set
Enough that one answer flipping does not visibly move the total. In practice that means at least thirty to fifty prompts for a single category.
Below about twenty, run-to-run variance dominates. You will see a five-point swing, go looking for a cause, and find that one answer phrased itself differently on Tuesday. Teams burn real time on this.
A workable split for a first set: roughly half category, a quarter comparison, a fifth problem, and the small remainder branded.
The rule that makes the numbers mean anything
Fix the set, then leave it alone.
Every prompt added or removed breaks comparability with everything measured before it. Adding three prompts where you happen to be strong raises visibility without anything changing in the world, which is the easiest way to produce a number that flatters a quarterly review and means nothing.
When the set does need to grow, add in deliberate batches, record the date, and annotate the trend line at that point. And never remove a prompt because the answer is unflattering — that prompt is the one telling you something.
Where prompts should come from
Not from a keyword tool. Search keywords are compressed, typed to a machine that rewards brevity. Prompts are conversational, longer, and often contain the constraint that decides the answer: a budget, a team size, an integration, an industry.
Better sources: your sales team’s first-call questions, your support inbox, the questions in your own community, and the way churned customers described what they were trying to do. Those are the sentences buyers actually produce.
Segment before you average
Report category and branded prompts separately. Blended, they produce a figure that looks healthy because your branded prompts are carrying it, and that moves for reasons you cannot act on.
The same applies per engine. Five engines will answer the same set differently, and the average of five different situations describes none of them.