Skip to content
Measure

AI Visibility Tracking: Sampled Answers vs Recorded Visits

TL;DR
  • Most AI visibility trackers run a fixed prompt list on a schedule and report the percentage of answers that name your brand.
  • The same prompt rarely returns the same brand list twice. A SparkToro and Gumshoe.ai study across 2,961 runs found fewer than 1 in 100 returned the same list of brands.
  • The vendor writes the prompts, so the score reflects their list, not your customers' questions.
  • A mention is not a visit. Pew found users clicked a result in 8% of Google visits with an AI summary against 15% without one.
  • Kymo records what AI engines actually did on your own site, crawler by crawler and page by page. It does not track mentions or what an assistant said.

Most AI visibility tracking products sample answers. A smaller group records visits. Those are two different measurements, and only one of them happened on your own site. If someone has quoted you a visibility score and you cannot trace a single dollar back to it, that gap is the reason.

How most AI visibility trackers work

The pattern repeats across the category. The vendor writes a fixed list of prompts, often a few hundred. A script runs them against ChatGPT, Claude, Perplexity or Google's AI surfaces on a schedule. A model reads each answer and checks whether your brand name appears. Mentions get counted, and the count becomes a percentage. Semrush, Ahrefs Brand Radar, Profound and Peec AI all sell a version of this.

The output is a number that goes up or down, sometimes with a rank position attached. It looks like a metric because it is formatted like one.

Three things break that number.

Fault one: the same prompt does not return the same answer

SparkToro and Gumshoe.ai ran 600 volunteers through 12 prompts across ChatGPT, Claude and Google's AI surfaces 2,961 times in January 2026. Fewer than 1 in 100 runs returned the same list of brands. Fewer than 1 in 1,000 returned the same list in the same order. Rand Fishkin described tools that report a ranking position in AI as "full of baloney".

The wobble is not a scripting bug. Thinking Machines Lab ran 1,000 completions of the same prompt at temperature 0 and got 80 unique completions, because server load changes how requests get batched.

Models also move under you. Chen, Zaharia and Zou measured GPT-4 scoring 84% on a prime-number task in March 2023 and 51% on the same task in June 2023. OpenAI rolled back a GPT-4o update on April 29, 2025 because the model had become overly flattering. When the model changes, your score changes for reasons that have nothing to do with your site.

One caveat from the SparkToro work: a percentage averaged across many runs is more stable than a rank position. That makes the number reproducible without making it true.

Fault two: the vendor picks the prompts

The prompt list is written by the company selling you the tool. It reflects their guesses about your category, not the questions your buyers type. A competitor with a long tail of relevant questions can score lower because those questions were never asked.

Prompt sets are also private. You cannot audit which questions produced your score, and you cannot confirm the same list ran against you last month.

Digital Applied took one brand and one dataset and scored it three ways: 20% on mention-based share of voice, 16.8% position-weighted, 31.4% citation-based. Same brand, same data, three answers.

Fault three: a mention is not a visit

Being named in an answer and someone arriving at your site are separate events. Pew Research Center tracked 900 US adults in July 2025. Users clicked a result in 8% of visits to a Google page with an AI summary, against 15% on pages without one. Clicking a link inside the summary happened in 1% of visits.

The surfaces disagree with each other too. Ahrefs compared 540,000 query pairs in December 2025 and found Google AI Mode and AI Overviews cited the same URLs only 13.7% of the time, despite 86% semantic similarity between the queries.

A mention can be real, stable, and still produce zero sessions on your site.

What a recorded measurement looks like

Kymo samples nothing. It runs a server-side receiver on your site and records what AI engines actually did there.

Each crawler hit gets classified into one of three categories from Kymo's own crawler registry. A live fetch made for one person's question is ai_answer, which includes ChatGPT-User (OpenAI's live fetcher). A crawler building a search index for later answers is indexing, which includes Claude-SearchBot (Anthropic's search index crawler). A bulk crawl for training data is training, which includes ClaudeBot (Anthropic's training crawler).

Those three behave differently. Training crawlers send no traffic. Indexing and live-fetch agents can lead to referral traffic. The distinction has its own post: ClaudeBot vs Claude-SearchBot: What Each One Does on Your Site.

What you want to knowSampled answerRecorded visit
Which pages does the engine read?Not visible, the prompt was genericPage-level, per crawler hit
Which question caused the fetch?The vendor's prompt listLikely prompt behind the fetch, in the AI report
Did anyone arrive?Not measuredConfirmed click-through from AI referrals
Did it produce revenue?Not measuredGoals with revenue attached
Can the number move without your site changing?Yes, on every model updateNo, the log entry is fixed

Kymo's AI Traffic page shows which of your pages AI engines read, how many people AI answers sent to each one, and which of those people completed a goal. Everything else is standard analytics: pageviews, unique visitors, referrers, sessions, bounce rate, countries, devices, the real-time map. The Kymo documentation covers setup.

What Kymo does not do

Kymo does not track mentions. It does not record what an assistant said about you in an answer. It cannot report share of voice, and it will not hand you a ranking position in AI. Those numbers require sampling someone else's chat, and sampling is the part that broke.

It also does not follow one person across days. Visitors are counted with a salted server-side hash that rotates at UTC midnight, so "unique visitors" means unique per day. No cookies, no localStorage for tracking. Raw IPs are never stored.

If you are deciding what to allow and what to block, Should You Block AI Training Bots but Allow Answer Bots? walks through the tradeoffs. The The Small Business Guide to AEO covers the content side, and What Is SEO for AI Called? AEO, GEO, AIO Explained sorts out the vocabulary. For a longer playbook, LLM SEO: What It Means and How to Start is the read.

Pricing and limits

Both plans include every feature and differ only by event volume. Solo is $9 a month or $90 a year, up to 10,000 events a month. Studio is $29 a month or $290 a year, up to 100,000 events a month. AI crawler tracking does not count against your event limit. Exceed the limit and the dashboard pauses; your data is not deleted. Both plans start with 14 days free, no card required.

Do this if you run a small site

Run a recorder if you are a solo builder or small team who needs to know whether AI engines actually visit you, rather than whether a synthetic prompt happened to mention you. Sample the mentions too if you have budget for both, because the two numbers answer different questions. Do not pay a monthly fee for a score that moves every time someone else retrains a model.

Stop guessing. Watch.

Most AI visibility tools ask a chatbot and report a guess. Kymo watches your own site instead: which pages AI engines read, how many people their answers send you, and which of those people buy. Every number comes from your server, not from a prompt someone else chose.

Start free → 14-day free trial. No card required.

Published Sep 28, 2026·All posts·The AEO guide