Before You Buy an LLM Visibility Tool, Ask These 8 Questions

- Most llm visibility tools type prompts into an assistant, count brand names in the answers, and sell the percentage as a score.
- Answers move run to run: a 2026 SparkToro and Gumshoe study of 2,961 runs found fewer than 1 in 100 returned the same brand list.
- Anyone selling a fixed rank inside ChatGPT is selling a number the model does not produce.
- Ask about prompt authors, run count, location, raw answers, and the share-of-voice formula before you pay.
- Kymo records what AI crawlers and AI referrals did on your own site, per page. It does not track mentions.
The short answer: before you buy any llm visibility tool, ask who chose the prompts, how many times each one ran, whether you can see the raw answers, and whether any of the result connects to a visit or a sale. If a vendor dodges those four, the demo is decoration.
Most llm visibility tools work by typing prompts into ChatGPT, Claude or Perplexity, counting brand names in the replies, and selling the resulting percentage as a visibility score. Some run that method carefully. Others do not. Canonry published a list of questions worth taking into every sales call, and this post walks through all eight, with the evidence behind each one.
Kymo takes a different route. It does not sample prompts and it does not parse what an assistant said. It records what AI crawlers did on your site, and what traffic came back. That is a narrower claim, and it is one you can check.
Ask whether the tool scrapes the assistant's app or calls the API
These produce different answers. The consumer app carries system prompts, memory, account history and product features the raw API does not. A tool that drives the web UI sees something closer to what a real user sees. A tool that calls the API sees a cleaner, cheaper, more repeatable signal that is not quite the same product.
Neither is wrong. You need to know which one you bought, because that determines how far you can generalise. Get it in writing.
Ask how many runs per prompt it does, and whether it shows the variance
One run per prompt produces one number and no error bar. Ask how many times each prompt ran and whether the tool reports spread.
The evidence is not subtle. A 2026 SparkToro and Gumshoe.ai study had 600 volunteers run 12 prompts through ChatGPT, Claude, and Google's AI Overviews and AI Mode, 2,961 runs in total. Fewer than 1 in 100 runs returned the same list of brands. Fewer than 1 in 1,000 returned the same list in the same order. Visibility percentage across many runs was steadier than position. Gumshoe sells AI tracking, and the study discloses that.
The same instability shows up below the product layer. Thinking Machines Lab ran the same prompt 1,000 times at temperature 0 and got 80 unique completions. The cause was server load changing the batch size. Nobody changed the prompt.
SparkToro's Rand Fishkin has a shorter version of this. He calls a reported ranking position in AI answers 'full of baloney'. The rank deserves it. A stable share across many runs is a more honest number, and only if the runs are disclosed.
Ask who chose the prompts
If the vendor wrote the prompts, the report tells you how your brand performs on their guess at your market. Ask for the list. Then compare it against the questions your buyers actually ask.
Strong prompt sets come from search queries, sales calls, and support tickets. Weak ones come from a category keyword and a spreadsheet. AI visibility, explained covers how to tell the difference.
Kymo skips the guessing step. It reads fetches that already happened, and the AI Traffic page shows which of your pages AI engines read and how many people AI answers sent to each one. Those are real requests, not a vendor's sample.
Ask which location the runs came from
AI answers vary by country, language, and sometimes by region inside a country. A runner sitting in Virginia is not measuring what a buyer in Manchester sees.
Ask where the queries were sent from, whether the tool covers the markets you sell into, and whether it stores location with the result. If the answer is "US only", you know exactly what you are buying.
Ask to see the raw answers
If the deliverable is a score with no transcript, you cannot audit it. Ask to see the actual model output for a few prompts.
This matters because the thing you are measuring moves without your input. Chen, Zaharia and Zou found GPT-4 scored 84% on a prime-number task in March 2023 and 51% on the same task in June 2023. OpenAI rolled back a GPT-4o update on April 29, 2025 because the model had become overly flattering. Neither event had anything to do with any marketer.
Ask which share-of-voice formula is behind the headline number
Share of voice is not one metric. Mention counting, position weighting and citation counting give three different answers from a single dataset. Digital Applied scored one brand on one dataset at 20% on mention-based share of voice, 16.8% position-weighted, and 31.4% citation-based.
Ask which formula the tool uses, and whether you can switch. If the number moves 10 points when the formula changes, the number was never the point.
Ask how it separates a model update from your own change
You shipped three pages in March. The score moved in April. Was that you? Ask how the tool attributes movement, whether it keeps a change log, and whether it holds any control prompts steady.
Google's own surfaces disagree with each other. Ahrefs compared 540,000 query pairs and found AI Mode and AI Overviews cited the same URLs only 13.7% of the time, with 86% semantic similarity. If two Google AI products disagree on which page to cite, one blended score hides the disagreement instead of explaining it.
Ask whether any of it connects to a visit or a sale
This is the question that decides whether the tool is a research toy or a business input. Can it show you a click, a session, a signup, or revenue?
AI summaries do reduce clicks. Pew Research Center tracked 900 US adults in July 2025 and found users clicked a result in 8% of visits to a Google page with an AI summary, against 15% without one, and clicked a link inside the summary in 1% of visits. Small numbers, and worth knowing precisely.
Kymo's receiver classifies crawler user agents into three groups: live fetches made for one person, indexers such as PerplexityBot, and bulk training crawls like Bytespider (ByteDance) and CCBot (Common Crawl). It then tracks confirmed click-throughs from AI referrals and goals with revenue attached, using the approach in Cookieless Analytics, Explained Without the Jargon, where the visitor identifier is a salted server-side hash that rotates at UTC midnight.
| What you are buying | What the number actually is | Where it breaks |
|---|---|---|
| Prompt sampling | Vendor's prompts run N times through an app or API | Answers change run to run, and by country |
| Mention share of voice | Share of sampled answers containing your brand | Formula changes the headline number |
| Citation-based share | Share of sampled answers linking your domain | Different Google products cite different URLs |
| Server-side crawler logging | Requests AI agents made to your site, and referrals back | Covers only your own site and your own traffic |
See an unedited dashboard before you commit
Ask for the demo. Kymo's is public at the Live Kymo demo dashboard, no login required. If a vendor will not show you a real dashboard, that tells you something too.
Two more things to know before you buy. Blocking a training crawler does not remove your site from that operator's answers, because the indexer and the live fetcher are separate agents, and Does Blocking AI Crawlers Hurt Your Visibility? works through that. If you publish an llms.txt file, treat it as a cheap bet rather than a lever. Adoption is partial, no major operator has publicly committed to honouring it, and publishing the file is not the same as a crawler fetching it. Only your server sees which of those happened. How to add an llms.txt file to Webflow if you want to try it in a few minutes.
For the wider vendor comparison, The Best AI Visibility Tools, and What Each One Can't Tell You lays them side by side. For the practical steps that come before tooling, there is The Small Business Guide to AEO.
Pricing, plainly: Kymo is $9 a month or $90 a year on Solo, up to 10,000 events a month, and $29 a month or $290 a year on Studio, up to 100,000 events a month. Every feature is included on both plans. They differ only by event volume. Both start with 14 days free, no card required. AI crawler tracking does not count against your event limit. Exceed the limit and the dashboard pauses. Your data is not deleted.
Stop guessing. Watch.
Most AI visibility tools ask a chatbot and report a guess. Kymo watches your own site instead: which pages AI engines read, how many people their answers send you, and which of those people buy. Every number comes from your server, not from a prompt someone else chose.
Start free → 14-day free trial. No card required.