maark
Log inStart for free

AI visibility tracking: the metrics that actually matter

Brand mentions, page citations, and competitor presence move differently. The AI visibility metrics worth tracking per engine, from 7,048 measured answers.

Maark teamAI visibility, How it works

AI visibility tracking comes down to three numbers, measured per engine on a fixed prompt set: how often your brand is mentioned in answers to non-branded buying-intent prompts, how often your pages are cited as sources, and how often your competitors appear. Everything else a tracker shows you — composite scores, share of voice, sentiment — is either derived from those three or decoration.

They are genuinely different numbers. In our Brand Invisibility Report — 7,048 answers from five engines, measured May 22 to July 29, 2026, across 21 brands — brands were mentioned in 2.2% of non-branded answers but cited as sources in 3.8%. Engines quoted brand pages as evidence more often than they named the brand. A tracker that collapses mentions and citations into one "visibility" figure hides the most actionable distinction in the data.

Mentions, citations, and competitor presence measure different things

Mention rate answers: are you in the recommendation? It is the number that maps to demand — a buyer who sees your name in an answer can go looking for you.

Citation rate answers: is your content the evidence? A cited-but-not-mentioned brand is doing the category's homework without getting the customer. In our corpus that gap was real on four of the five engines (ChatGPT was the exception, mentioning slightly more than it cites) — which means "we're in the sources" and "we're in the answer" need separate lines on your report.

Competitor presence answers: is the shelf filling in? Today it mostly isn't — in 97.7% of non-branded answers we measured, neither the brand nor any tracked competitor appeared, and the competitor-named-you-absent outcome occurred in just 0.13% of probes. That makes a rising competitor rate one of the more useful early warnings in the whole system: it means your category's empty shelf is being claimed.

Branded and non-branded prompts are different tests

The same study found brands mentioned in 79.9% of prompts that named them, against 2.2% of prompts that described a need. Those aren't two points on one scale; they are two different instruments:

  • The branded set ("is this brand good for X") is a recall and health check. It should stay high. A drop means something broke — crawler access, a reputation problem, a model update — and is worth investigating the week it happens.
  • The non-branded set ("best X for Y") is the discovery metric, the one your content work is supposed to move. It starts near 2% for almost everyone, and it is the number to trend.

Keep the sets separate and label them. Even a few branded prompts mixed into a non-branded set will flatter the rate enough to hide real movement.

Track engines separately — never as one average

Per-engine rates from the report:

| Engine | Non-branded probes | Brand mentioned | Brand pages cited | |---|---:|---:|---:| | Perplexity | 1,626 | 2.0% | 6.3% | | ChatGPT | 1,618 | 1.9% | 1.1% | | Gemini | 1,550 | 2.5% | 5.4% | | Google AI Overviews | 1,447 | 2.1% | 2.3% | | Google AI Mode | 673 | 3.3% | 4.3% |

Mention rates sit in a narrow 1.9–3.3% band — the invisibility problem is structural, not one engine's quirk. But citation behavior spreads almost sixfold, from ChatGPT's 1.1% to Perplexity's 6.3%. An average across engines would tell you nothing about where you are actually reachable, and reachability is what you act on: Perplexity and Gemini reward citable pages fastest, while Google's AI surfaces sit in between. Weight the engines your buyers actually use.

Expect churn — and attribute movement honestly

External research from Authoritas puts AI-answer citation churn at roughly 70% over 2–3 months. Build that into how you read every number:

  • A point-in-time screenshot of "we're in ChatGPT" is close to worthless. Trends on a stable prompt set are the unit of truth.
  • The same prompt can return different answers on different days, so probe repeatedly and read rates, not individual answers.
  • Losses aren't necessarily your fault, and wins aren't permanent. Hold both loosely and re-measure before reacting.
  • Churn is also the opportunity: the answer shelf gets restocked every few months, which is why new editorial content gets real chances to enter.

Build or buy

You can build a serviceable tracker: script API-based probes of each engine on a schedule, log a per-probe record — the raw answer text, mentioned yes or no, cited URLs, competitors named — and normalize citations to hosts so you can see what fills your category's answers.

Operating our own pipeline taught us where the effort actually goes. API probes carry no logged-in personalization, so treat rates as directional for real user sessions. Five engines means five capture paths that each break separately. And classification is the long tail: our own source-type classification covers about 47% of a 70,106-citation graph so far, and it is ongoing work.

If you buy instead, demand: per-engine reporting (never only a blended score), control of your own prompt set, competitor sets you define, citation-level data with raw answers retained for audit, and a refresh cadence measured in days. That list is the shape we built Maark's AI visibility module around, and it is a fair bar for any tool.

What an AI visibility score should mean

One number for the exec summary is fine — if it decomposes. A useful score is built from the non-branded mention rate per engine (weighted by the engines your buyers use), with citation rate as its own line, branded recall as a health check, and a 4–12 week trend. A score that can't be traced back to probes you can read is a vibe.

Benchmark honestly, too. Across our 21-brand fleet, per-engine non-branded mention rates ran 1.9–3.3% in mid-2026. If a tool reports your "visibility" at 40%, check the prompt set for branded contamination before celebrating.

The weekly report worth reading

Six sections, one page:

  1. Non-branded mention rate per engine, with week-over-week deltas.
  2. Citations won and lost — which of your pages entered or left which answers.
  3. Competitor entries — any tracked competitor newly appearing, and on which prompts.
  4. Branded recall check — flag any drop immediately.
  5. Prompt set changes — what was added or retired, so trends stay comparable.
  6. Three actions, maximum — usually: more of whichever format just got cited. The playbook for earning mentions is the companion to this report.

FAQ

How many prompts does an AI visibility tracker need?

Our study tracked 1,085 prompts across 21 brands — roughly 50 per brand. That is enough to cover your core buying-intent phrasings plus variants, and enough volume for rates to be readable. More prompts make trends stabler; fewer than a couple dozen makes every percentage jumpy.

What counts as a good AI visibility score?

Calibrate against the measured baseline: per-engine non-branded mention rates of 1.9–3.3% across our fleet as of July 2026. Sustained performance above that band means you are ahead of the measured field. On branded prompts, expect roughly 80% recall — treat anything well below that as a defect, not a growth project.

How often should you measure AI visibility?

Probe daily, read weekly. With citation churn around 70% over 2–3 months (Authoritas), monthly one-off audits mostly measure noise, and quarterly ones measure a different internet.

Is AI visibility tracking the same as GEO?

No — GEO is the practice of earning presence in AI answers; tracking is the measurement loop that tells you whether it worked. The distinction, and how both relate to classic search work, is laid out in SEO vs GEO.

Do you need to track all five engines?

Track every engine your buyers plausibly use, and always more than one: the per-engine spread in citation behavior means a brand can be reachable on Perplexity and invisible on ChatGPT at the same time. For most teams that means ChatGPT, at least one citation-forward engine, and Google's AI surfaces.

Maark measures this daily across five engines — mentions, citations, and competitor presence on prompt sets you control — join the waitlist to get your own numbers.

Questions about anything here? The help center goes deeper, or talk to a human.