Methodology

How we measure, and what we refuse to claim

AI visibility is easy to fake and hard to measure. This page is the arithmetic, the sampling rules, and the explicit list of things TrackGeo does not guarantee — so you can judge our numbers instead of trusting them.

A model never picks your score

Scoring, ranking, trends and benchmarking are computed by our own deterministic code from your stored runs. Models do four jobs here: run your prompts, classify sentiment, extract entities, and write prose explanations. Asking a model to rate you out of 100 would produce a number that changes when nothing else did — so we do not do it.

01 — Prompt selection

How prompts are selected

A measurement is only as good as the question behind it. Tracking questions nobody asks produces a flattering number that means nothing.

Prompts enter your workspace one of three ways: you write them, you import them, or you start from suggestions generated for your brand and category when you onboard. You decide what stays.

Every prompt is then classified into one of 13 intent categories — from pricing and buying, through comparison, alternatives and competitor replacement, down to informational — and one of 5 funnel stages from awareness to post-purchase. Those two classifications drive a buyer-value score out of 100, blending buyer intent, conversion potential, revenue relevance, competitor intensity, search demand, funnel weight, product relevance, ICP fit, region and strategic importance. Prompts at 70 and above are tiered high, 45 and above medium.

Region and language are part of the prompt, not a filter applied afterwards. The same question asked in another market is a separate measurement with its own runs, and the two are never averaged into one number.

What the scoring rewards

Commercial intent, not comfort. A prompt is never weighted up because you perform well on it — the highest-value questions in most workspaces are the ones being lost, which is why the portfolio has a dedicated view for exactly that case.

Signals you have not connected fall back to neutral rather than being invented, so a new workspace is scored conservatively rather than optimistically.

What counts as one measurement

One prompt, on one engine, in one region, at one point in time. Everything above that — trends, share of answer, confidence — is built by aggregating those units. Nothing is aggregated across units that were not comparable in the first place.

02 — Sampling

Why one answer proves nothing

Language models are non-deterministic. Ask the same question twice and you can get two different answers naming two different sets of brands. This is not a flaw in the measurement — it is the thing being measured.

The same question, two answers

Identical prompts return different responses run to run. A screenshot of an assistant praising you is evidence that it did so once. It is not evidence that it does so.

Scheduled, not sampled once

Your prompts run on a schedule — daily or weekly — across the engines on your plan. The pattern across those runs is the finding. Any single run inside it is just one draw.

The ground moves underneath

Providers change models, system prompts and grounding without announcing it. Continuous measurement is what tells you the answer changed. A one-off audit cannot, by construction.

The consequence

Because a single run cannot establish a pattern, a finding built on fewer than two runs is never labelled high confidence — regardless of how strong it looks. Consistency is undefined without repetition, so we treat it as undefined rather than rounding it up.

03 — Confidence

How confidence is derived from sample size

Confidence answers one question: how much should this finding be allowed to influence a decision. It is not a vibe, and it is not a model's opinion — it is arithmetic you can check.

Step one — how consistently did it repeat

Each sub-score is the share of runs that agree with the majority outcome, from 0 to 100, where 100 means every run agreed. They are blended by how much each one matters to a business decision.

consistency = mention agreement     × 0.40
            + competitor agreement  × 0.25
            + citation agreement    × 0.20
            + sentiment agreement   × 0.15

Step two — how much evidence backs it

Consistency alone would let one lucky run read as certainty. So it is gated by a sample-size factor that can only ever range from 0.3 to 1.0.

confidence = consistency × sample size factor(runs, engines)
Runs
Contribute up to 0.7 of the factor, saturating at 8 runs.
Engines
Contribute up to 0.3 of the factor, saturating at 3 engines.
Floor
The factor never drops below 0.3, so a single consistent run keeps some confidence — but can never reach full confidence.

Step three — what you see

The score becomes a label, and the label always travels with the sample that produced it.

  • High75 and above
  • Medium50 and above
  • Lowbelow 50

One override outranks the arithmetic: fewer than two runs is always low confidence, whatever the score says.

Every finding also carries the reason in plain language — how many runs, across how many engines, and what the agreement was on each signal. You never have to accept a label without seeing what earned it.

04 — Citations

How citations are captured

Citations are recorded from the run itself, at the moment the answer came back. They are never reconstructed afterwards from a guess about where an answer probably came from.

When a prompt runs, we store the response and what it cited: the source URL, the engine, the position it appeared at, and when it was captured. That record is what every citation number on your dashboard aggregates, and it is what you open when you want to see why a finding says what it says.

Domains are then classified — editorial, review, community or owned — and rolled into a source mix across trusted, community, social, owned and other. That split is what tells you whether the answer is being shaped by coverage you would have to earn or content you could write yourself.

Engines differ in what they expose, and we do not paper over the gap. Some return sources alongside the answer, which is what makes source tracking possible on those surfaces. Where an engine returns no sources, we record what it did return — we do not infer a citation that the engine never gave us.

Which engines return sources

What a stored run keeps

The response
The answer text the assistant actually returned, kept so a finding can be checked rather than believed.
The engine and the moment
Which surface produced it, and when it was captured — because the same prompt on the same engine can answer differently next week.
The cited sources
Every cited URL with its position, classified by domain and source type.
What changed
Mention flips, position deltas, and sources that appeared or dropped since the previous run.

How long runs and responses are retained is yours to control. See security for retention and data handling.

05 — Evidence tiers

Observed versus inferred

Most of this category quietly reports inferred influence as revenue. We separate what we saw from what we modelled, because the difference is the entire value of the measurement.

Observed

The referrer directly identifies an assistant. We saw the session arrive. This is measurement, and it is the only tier we will ever describe as attribution.

Estimated

Modelled influence where the referral chain is incomplete. Useful for direction, labelled as an estimate in the product, and never totalled into an observed figure.

Inferred

Correlation between visibility movement and demand — for example search behaviour moving alongside a mention. Directional, not attributed. It is not proof that AI caused the outcome.

Exposure

How often you appear in answers at all. A leading indicator. It is not traffic, and reporting it as traffic is the single most common overstatement in this category.

The rule

These tiers are labelled separately everywhere they appear — on screen, in exports, and in reports handed to your client. They are never summed into a single headline figure, because adding something you observed to something you modelled produces a number that means neither.

06 — Limits

What TrackGeo does not guarantee

This section is not a disclaimer we hid at the bottom. If a competitor promises you any of the following, they are describing something outside their control.

We do not guarantee rankings, mentions or traffic

Nobody can, and you should be sceptical of anyone who says otherwise. Assistants change models, system prompts and grounding without notice, and their outputs vary run to run. We show you what is happening, what changed, and what evidence supports an action. The outcome is not ours to promise.

Inferred influence is not proven revenue

When visibility moves and demand moves with it, that is a correlation. We label it inferred in the product and in every report, and we keep it separate from directly observed referrals. Any vendor totalling the two into one revenue number is selling you a story.

Provider APIs are not the consumer apps

We measure 6 surfaces through provider APIs — the models behind the assistants. We do not scrape the consumer apps, and we do not read AI Overviews off a search page. What a model returns through its API is a real and defensible signal, but a user's app session can differ from it, and we will not blur that line to make coverage sound broader.

A visibility score is not a search ranking

There is no position one, and there is no single answer everyone sees. Our index is a defined, published formula for comparing your own measurements over time and against competitors measured the same way. It is not a universal truth about your brand, and it is not comparable to another vendor's score.

An action is a recommendation, not a causal claim

Actions are prioritised from evidence we hold, not from a proven causal model of a given assistant. When you mark one implemented, we re-scan and record what actually happened — including when nothing did. The before-and-after stays empty until a real re-scan fills it.

We do not report a number we did not measure

Where a signal is missing, it falls back to neutral or stays empty rather than being seeded with a plausible-looking value. An empty state is information. A fabricated percentage is not.

What we do promise is narrower and more useful: every number traces to a stored response, every finding shows the sample behind it, and anything we inferred says so.

Questions

Methodology, answered plainly

How are prompts selected?

You write them, import them, or start from suggestions generated for your brand and category when you onboard. Each prompt is then classified into one of 13 intent categories and one of 5 funnel stages, and scored out of 100 on buyer value. Region and language are part of the prompt, not a filter on it — the same question in another market is a separate measurement.

Why do repeated runs matter?

Language models are non-deterministic: ask the same question twice and you can get two different answers naming different brands. One response is an anecdote. Running a prompt repeatedly, over time and across engines, is what separates a real pattern from a coincidence — and it is why a finding from fewer than two runs is never labelled high confidence.

How is confidence calculated?

Confidence blends how consistently a finding repeats with how much evidence backs it. Consistency is the share of runs agreeing with the majority outcome, weighted across mention, competitors, citations and sentiment. That is then multiplied by a sample-size factor between 0.3 and 1.0, where runs saturate at 8 and engines at 3. Scores of 75 and above are labelled high, 50 and above medium. Fewer than two runs is always low, because consistency is undefined without repetition.

Does an AI model decide my score?

No. Scoring, ranking, trends and benchmarking are computed by our own deterministic code from your stored runs. Models are used to run prompts, classify sentiment, extract entities and write prose explanations. A model is never asked to pick a score, so the same runs always produce the same number.

What is the difference between observed and inferred?

Observed means the referrer directly identifies an assistant — we saw the session arrive. Inferred means we correlated visibility movement with demand and are showing you a direction, not an attribution. Exposure means you appeared in an answer, which is a leading indicator and not traffic. These are labelled separately in the product and never totalled together.

Can TrackGeo guarantee AI rankings?

No. Assistants change models, prompts and grounding without notice, and their outputs vary run to run. TrackGeo shows you what is happening, what changed, and what evidence supports a recommended action. Any vendor guaranteeing a mention or a ranking in an AI answer is describing something they cannot control.

Judge the numbers for yourself

Run a free audit of your site, or talk to us about what continuous measurement would look like for your brand and market.