Skip to content
Crawlers

Should You Block AI Crawlers? (It Depends Which Kind)

TL;DR
  • Blocking all AI crawlers is a blunt move that also cuts off traffic from ChatGPT, Perplexity, and Claude referrals.
  • Training crawlers like GPTBot and ClaudeBot can be blocked in robots.txt without hurting your visibility in AI answers.
  • Indexing crawlers, like PerplexityBot, need access for your site to be considered a source.
  • Live answer crawlers, like ChatGPT-User, represent real people asking questions about your business.
  • You cannot decide what to block until you know which crawlers actually visit your site.

No, not across the board. The word "AI crawler" lumps together three very different types of bots, and treating them the same means blocking the ones that can send you customers.

The panic is understandable. Your server logs show a flood of unfamiliar user agents, and someone online said robots.txt is the fix. That advice is too simple, and it can cost you the exact traffic you are trying to protect.

The decision to block ai crawlers should be made per category, not as a single policy. Here is what each type does, what blocking it means, and how to decide.

The three kinds of AI crawlers

Kymo classifies AI crawlers into three groups because that is how they actually behave. They are not interchangeable.

ai_answer: A live fetch made for one person's question. When someone asks ChatGPT for a plumber and ChatGPT pulls your site in real time, that is ChatGPT-User. This is the crawler that sends referral traffic back to you when a person clicks your link in the answer.

indexing: Search-index crawlers that build the knowledge base for AI answers. OAI-SearchBot, Claude-SearchBot, and PerplexityBot fall here. They are similar to Googlebot in function. They need access for your content to be considered at all.

training: Bulk crawls for model training data. GPTBot and ClaudeBot are the most common. These do not send you traffic. They do not cite you in the moment. They exist to improve a model that may or may not ever mention your business.

Here is the cheat sheet:

Crawler typeExamplesWhat they doBlocking impact
ai_answerChatGPT-User, Claude-User, Perplexity-UserLive fetch for one person's questionKills direct referral traffic
indexingOAI-SearchBot, Claude-SearchBot, PerplexityBotBuild search index and answer contextSite becomes invisible in AI answers
trainingGPTBot, ClaudeBotBulk data for model trainingZero effect on current traffic

What you actually lose when you block everything

Let us be concrete. If you block all AI crawlers, Google is not affected. Your SEO rankings stay. What changes is discovery through AI answers.

ChatGPT has a web browsing feature. When it uses it, the crawler fetches your page live, and the answer can include a clickable citation. Some of those clicks land on your site. Same for Perplexity, which has been sending referral traffic since 2023. Claude's search feature works the same way.

PerplexityBot is the clearest case. It indexes your pages, and when someone asks a question your content answers, Perplexity links to you. That is a referral. That is a person clicking through. Block PerplexityBot and you vanish from one of the fastest-growing search surfaces on the web.

The Small Business Guide to AEO goes deeper into why this matters, but the core point is simple: AI answers are the new front door. If you block the crawlers that open that door, you are locking it yourself.

The case for blocking training crawlers

GPTBot (OpenAI's training crawler) and ClaudeBot (Anthropic's training crawler) serve a different purpose. They take your content, process it, and contribute to a model that may never cite you. You get no traffic, no direct credit, and no way to measure the return.

There are legitimate reasons to block them:

  • You do not want your proprietary content absorbed into a model you do not control.
  • Your content is licensed or regulated.
  • You simply see no value in giving away data for free.
  • You already have a visibility problem and cannot afford to be a data source without getting citations back.

Blocking training crawlers is a defensible, low-cost decision. It does not stop you from appearing in AI answers. It only stops bulk training access.

What about the data equation

There is a wrinkle. Some indexing crawlers are also used for training. OpenAI has separate user agents for each function, but not every provider is that clean. You may block a training user agent and accidentally block the indexing one too.

Check your robots.txt rules carefully. If you use a wildcard to block all GPT bots, you also block OAI-SearchBot, which is the crawler that lets you appear in ChatGPT's live answers. That is the mistake most people make.

The fix is to be specific. Block exactly the training user agents you want gone, and leave the indexing and answer crawlers alone.

How to decide for your site

Start with data. Open your logs and look at what actually visits. You will likely find that training crawlers are frequent, indexing crawlers are moderate, and ai_answer crawlers are rare but valuable.

That is the pattern on most Kymo dashboards. Training crawlers hit everything. ai_answer crawlers only show up when someone actually asks about your niche. When they do, those few visits often convert because the intent is already there.

The question is not "should I block AI crawlers." The question is "which ones are costing me more than they return."

  • If you have zero AI referral traffic and you care about data privacy, block the training crawlers.
  • If you want to be visible in AI answers, leave indexing crawlers alone.
  • If you have seen clicks from ChatGPT, Perplexity, or Claude, protect that pipeline like any other referral source.

The tools to make the call

Kymo shows you which crawlers visit, which pages they hit, and whether any of them send back confirmed click-throughs. That is the missing piece. Most blocking advice is written blind, because most site owners do not know what their crawler traffic looks like.

You should also read Privacy-First Analytics: What It Actually Means if you are worried about how tracking works in the first place. Kymo uses cookieless identity, so isolating crawlers does not require storing personal data about humans.

The AI crawler directory lists the user agents you will find in your logs, so you can match what you see to the right category.

And if you are not even sure AI answers mention your business, Does My Business Show Up in ChatGPT? How to Find Out walks you through the basics. The free AI visibility check at kymo.in/tools/ai-visibility-checker reads a URL from public signals only, so your site does not need Kymo installed, and the report arrives by email.

The bigger shift underneath this

You are not just deciding about crawlers. You are deciding how to position your business as discovery shifts from search engines to answer engines.

GEO vs SEO: What's Different and What Still Works covers the mechanics. AI SEO Tools for Small Business: What's Worth Paying For covers the budget. The short version is that the fundamentals of good content, clear answers, and fast pages still carry you, but the distribution channel has changed.

Google sends you traffic based on ranks. AI answers send you traffic based on being chosen. You cannot be chosen if you block the chooser from reading you.

Price also matters here. Kymo costs $9 a month for the Solo plan, up to 10,000 events, and $29 a month for Studio, up to 100,000 events. Both include every feature and start with 14 days free, no card required. AI crawler tracking does not count against your event limit, so auditing your crawler traffic is effectively free.

The bottom line

Block training crawlers if you want. Keep indexing crawlers. Do not block answer crawlers unless you actively hate referral traffic.

The data will tell you which is which. Kymo's AI Visibility report shows your pages fetched by AI assistants, the likely prompts behind those fetches, and confirmed click-throughs from AI referrals. That is the difference between guessing and knowing.

Do not touch robots.txt until you have seen that report. The cost of blocking the wrong crawler is invisible until your traffic drops, and by then you will not know why.

A final recommendation

Block GPTBot and ClaudeBot if you get no value from training crawlers, but leave the indexing and answer crawlers untouched if you want any chance of appearing in AI answers. Do that if you have a business that relies on discovery. If you run a members-only site or sell proprietary content, block training crawlers too, and use your robots.txt to draw a clear line between the crawlers that read for the public and the crawlers that read for the machine.

Then start tracking it. Kymo shows you which of these crawlers actually visit your site and what they do with it. You will see the difference between a training crawl that means nothing and an ai_answer fetch that means a potential customer. The documentation explains how the classification works, and you can be live in minutes.

Start free → 14-day free trial. No card required.

Published Aug 11, 2026·All posts·The AEO guide