There Is No Rank in ChatGPT: What AI Rank Trackers Report
- SparkToro and Gumshoe.ai ran 12 prompts through ChatGPT, Claude and Google's AI surfaces 2,961 times. Fewer than 1 in 1,000 runs returned the same brand list in the same order.
- Rand Fishkin called AI tools that report a ranking position "full of baloney", and the data supports him.
- Models change answers run to run even at temperature 0. Thinking Machines Lab got 80 unique completions from 1,000 identical prompts.
- Visibility percentage across many runs is more stable than position, so track presence rather than rank.
- Kymo does not sample prompts or score mentions. It records what AI crawlers actually did on your own site, per page.
There is no rank in ChatGPT. An assistant writes an answer to one question from one person, and that answer changes on the next run. So any AI rank tracker that reports "position 3 in ChatGPT" is showing you a single sample and calling it a metric. The bigger problem is that most of these tools never touch your site at all.
The same prompt gives a different answer every time
SparkToro and Gumshoe.ai tested the sampling problem directly with 600 volunteers running 12 prompts across ChatGPT, Claude, and Google's AI surfaces. That produced 2,961 runs. Fewer than 1 in 100 runs returned the same list of brands. Fewer than 1 in 1,000 returned the same list in the same order. You can read the January 2026 SparkToro study in full.
Rand Fishkin, who co-authored it, is blunt about what that means for AI rank tracking. Tools reporting a position in AI are "full of baloney." He is right. If the order shifts almost every run, a rank is a coin flip with a decimal point.
The study is fair about what does hold up. Visibility percentage, meaning how often a brand appears across many runs, was more stable than position. If your brand showed up in 40% of runs one week and 42% the next, that is a real reading of a noisy system. "Position 3" is not.
Models are not deterministic, and the prompt is not the cause
You might think the same prompt on the same model returns the same answer. It does not. Thinking Machines Lab showed in September 2025 that 1,000 completions of one prompt at temperature 0 produced 80 unique outputs. Server load changed the batch size, and floating point math did the rest.
Model updates push results further. Chen, Zaharia and Zou found that GPT-4 scored 84% on a prime-number task in March 2023 and 51% on the same task in June 2023. Same model name, very different behavior. OpenAI rolled back a GPT-4o update on April 29, 2025 because the newer version was too flattering.
So when an AI rank tracker records a position, it is recording a snapshot of a system that has already changed underneath it.
Position-based AI tracking versus signal-based tracking
| Method | What it does | What it can show | Where it falls short |
|---|---|---|---|
| Prompt sampling | Types prompts into assistants, counts brand mentions, reports a score | Rough visibility percentage across many runs | Position is one sample. No link to real activity on your site |
| Server-side crawler and referral tracking | Records which AI agents fetched pages and which sent traffic back | Which pages AI engines read, and which referrals became visits | Does not record what an assistant said, because no outside server can see that |
| Citation-based share of voice | Counts citations to your domain in AI answers | Higher signal than raw mention counts in published datasets | Different scoring methods produce different numbers |
Digital Applied scored one brand on one dataset in 2026 and got three numbers: 20% on mention-based share of voice, 16.8% on position-weighted share, and 31.4% on citation-based share. Same brand, same data, three scores. That is the shape of the AI rank tracking category right now.
What AI actually does on your site is observable
Kymo takes a different approach. It does not type prompts into assistants. It does not count mentions or score what an assistant said, because no server outside that assistant can see the full answer.
What it does is run a server-side receiver that classifies crawler user agents into three categories: ai_answer (a live fetch for one person's question, like ChatGPT-User (OpenAI's live fetcher)), indexing (crawlers building search indexes for later answers, like PerplexityBot), and training (bulk crawls like GPTBot (OpenAI's training crawler) and CCBot (Common Crawl)). The AI crawler directory lists every agent Kymo recognizes with its category and whether it can send referral traffic.
Then it reports what happened, per page: which pages were fetched, by which category, and which AI referrals turned into visits. That is server-side data. It does not change based on the day's batch sizes.
Kymo is also honest about what it cannot see. It does not track mentions. It does not tell you what an assistant wrote. It does not follow one person across days. It counts what hit your server.
The AI referral click numbers are smaller than most people expect
Pew Research Center studied 900 US adults in July 2025. On Google pages with an AI summary, users clicked any result in 8% of visits, against 15% without a summary. They clicked a link inside the summary itself in 1% of visits.
AI surfaces do not agree with each other either. Ahrefs analyzed 540,000 query pairs in December 2025 and found Google AI Mode and AI Overviews cited the same URLs only 13.7% of the time, despite 86% semantic similarity between answers. Same company, same queries, almost entirely different sources.
You cannot manage this with a rank. You manage it by knowing which of your pages AI engines fetch, and which categories of crawler show up. If your pricing page gets live fetches and a competitor's gets more, that is a page-level problem you can work on.
Publishing an llms.txt file is a cheap bet, not a lever. Adoption is partial, no major operator has publicly committed to honoring it, and only your own server sees whether a crawler fetched the file. Real llms.txt examples scored against the spec is a decent starting point.
How to use AI visibility data without fooling yourself
Treat any single AI answer as one sample, because that is what it is. If a tool reports your brand at "position 2 in ChatGPT", ask what the sample size was, whether the position was averaged across runs, and whether the tool can show you the fetches behind the number.
Visibility percentage across many runs is more defensible than position. Citation-based scoring beats mention counting in most published comparisons. The ground truth is still what your own server sees. A crawler fetch is a fact. A mention count is a model.
If you want a starting point, the free AI visibility check reads a URL from public signals and sends the report by email. No account needed, and the site does not need Kymo installed. To understand what a visibility score can and cannot mean, that is a separate question worth reading about. Working through the whole approach is covered in The Small Business Guide to AEO. Two more posts that pair well: What Is an AI Visibility Checker and Do You Need One? and Does My Business Show Up in ChatGPT? How to Find Out.
Kymo costs $9 a month or $90 a year for Solo, up to 10,000 events a month, and $29 a month or $290 a year for Studio, up to 100,000 events a month. Every feature is on both plans, and AI crawler tracking does not count against your event limit.
Stop guessing. Watch.
Most AI visibility tools ask a chatbot and report a guess. Kymo watches your own site instead: which pages AI engines read, how many people their answers send you, and which of those people buy. Every number comes from your server, not from a prompt someone else chose.
Start free → 14-day free trial. No card required.