Someone forwards a screenshot. ChatGPT recommended three tools in your category, one was your closest competitor, and you were not on the list. By the time it reaches the marketing channel it has become a strategy: we need to beat them in AI.
Slow down. That screenshot is not a finding. It is a hypothesis, and a cheap one to test. The teams that get value from competitive AI intelligence run the test before the project.
A competitor appearing is not a finding
"They appeared and we did not" is one draw from a distribution. Before it justifies a roadmap it has to survive three questions, in this order. Skipping to the third is how content teams rewrite pages that were never the problem.
- Is the gap real, or did you sample an unlucky answer?
- Is the prompt contested or already decided? These need opposite strategies.
- Why do they win — what is the engine actually reading when it names them?
First, establish that the gap is real
Run the prompt repeatedly across several engines and count. Not because repetition is a virtue, but because the alternative is acting on a coin flip. A competitor appearing in 8 of 8 runs across three engines and one that appeared in the single run you screenshotted look identical from a screenshot.
The surface distinction matters most in competitive work. TrackGeo queries provider APIs — OpenAI, Perplexity's sonar models, Gemini, the Anthropic Messages API, Azure OpenAI, and Gemini with search grounding — not the consumer apps. So a logged-in ChatGPT screenshot and a TrackGeo run are not the same measurement, and when they disagree neither is lying. The API run is the one you can repeat, which makes it the one you can trend. The method is in "How to measure whether ChatGPT and other AI platforms recommend your brand" (/blog/measure-ai-recommendations).
Read the volatility, not just the gap
Here is the part almost nobody does, and the highest-leverage read in competitive AI work. Once you have repeated runs, do not only ask whether a competitor appears. Ask how consistently.
For a yes-or-no signal, consistency is the share of runs agreeing with the majority outcome — which means it cannot fall below 50 percent. That floor is the useful part. A competitor appearing in 4 of 8 runs is not a mild signal; it is the maximum instability the signal can express. The model has no settled view. That prompt is genuinely undecided, and undecided is where a marketing team can change something.
- Prompt A — competitor in 8 of 8 runs
- Consistency 100. Entrenched: the model has a settled view and it is not you.
- Prompt B — competitor in 6 of 8 runs
- Consistency 75. Leaning their way, but not certain.
- Prompt C — competitor in 4 of 8 runs
- Consistency 50, the floor. Maximum instability — the answer is up for grabs.
- Instinctive priority
- Prompt A: the biggest, most visible, most annoying gap.
- Better priority
- C, then B. C is one evidence shift from flipping; A needs the whole source landscape rewritten.
- Caveat that changes the answer
- If A is your highest-value purchase-intent prompt and C is incidental, A wins anyway. Volatility ranks targets, it does not choose them.
The instinct is to attack the biggest gap, and the biggest gap is often the worst target. Same effort, wildly different odds.
Then ask why they win
Once the gap survives sampling, the question stops being whether and becomes what the engine is reading. On grounded engines this is inspectable: the answer carries citations, and the citations have a shape. TrackGeo's displacement analysis rebuilds each finding from stored runs and their citations — which competitors were recommended, which sources were cited, what type they were — so every field traces to a real answer rather than a generated guess.
The source type is usually the diagnosis, and it is rarely what the team assumed.
- Third-party roundups and "best X" listicles. The reason is unglamorous: your competitor is on the list and you are not. A placement problem, not a content problem — rewriting your own site will not fix it.
- Review platforms. The engine is reading aggregate user sentiment. Your fix lives in your review pipeline, not your CMS.
- Documentation and technical pages. Their docs answered the question concretely and yours did not answer it at all. Usually the fastest fix here.
- Forum and community threads. Someone asked your exact question three years ago and a competitor got recommended. That thread is now training the answer.
- The competitor's own domain. Hardest to displace — they have made themselves the canonical source. Worth knowing before you commit a quarter.
- No citations at all. On ungrounded engines the recommendation comes from training data, so there is no source to fix. Nothing you publish this month changes it, and treating it as an SEO problem burns the quarter.
Share of voice against mention rate
Two numbers that look interchangeable and answer different questions. Mention rate asks whether you appear. Share of AI voice asks how much of the conversation is yours: your mentions divided by your mentions plus your competitors'.
- Your mentions across the prompt set
- 12
- Tracked competitor mentions
- 48
- Share of AI voice
- 12 / (12 + 48) = 20 percent
- Six months on, your mentions unchanged at 12
- Competitor mentions fall to 28, so share of AI voice rises to 30 percent
- What moved
- Not you. Your mention rate is identical. The category consolidated around fewer names and you survived the cut.
Share of voice is the better competitive metric and the worse performance metric. It moves when your rivals move — what you want on a competitive dashboard, and what you do not want in a report claiming credit for your team's work. Report both, and be precise about which one moved.
Prioritise by value, not by gap size
The last trap is treating every lost prompt as equally worth winning. Losing "what is generative engine optimization" is a bruise. Losing "best alternative to [your competitor]" is a lost deal. TrackGeo blends the opportunity in a loss with the prompt's buyer value in equal parts, so a transactional or competitor-replacement prompt outranks a definitional one.
Combine that with volatility and the target list is defensible in a planning meeting: contested prompts, on grounded engines, where a buyer is close to a decision. A short list, which is the point.
What to actually change
- Get into the specific roundups the engine cited. Pitch the publication that appeared in the citations, not "do some PR" — you have the exact URL.
- Answer the prompt's question literally, on a page, with the constraint in it. If the prompt is "best CRM for a 12-person agency", a page whose first paragraph answers for a 12-person agency is quotable in a way your homepage never will be.
- Publish the comparison the model is reaching for. If it cites a third-party comparison of you and a rival, that is because you never published one with real numbers in it.
- Fix the factual claims first. An outdated pricing or feature claim in a cited source is a bug with a URL — the cheapest fix here and the most frequently ignored.
- Make claims extractable. A specific number in a table beats the same fact buried in positioning prose, because one lifts into an answer and the other does not.
- Re-sample the same prompt set, unedited, against the confidence bar you started with.
What this does not prove
It also does not prove revenue. A competitor recommended more often than you is a plausible influence on buyers, not a quantified loss, and converting it to a pipeline figure invents precision you do not have. The defensible claim is narrower and stronger: on these prompts, at this sample size and confidence, these competitors are recommended more consistently than you, and these are the sources the engines cited.
That is enough to act on, and honest about its edges. For how the rates and confidence are computed, see /methodology; for the competitive and citation views, /features. And "AI visibility vs traditional SEO: what marketing teams need to track" (/blog/ai-visibility-vs-seo) covers which of these numbers belong in the monthly deck.
See where you stand
Run a free audit of your site, or read how we test prompts, sample repeatedly and score confidence.