cohere-ai: Cohere's Training Crawler
What is cohere-ai?
cohere-ai collects web content associated with Cohere. cohere-ai has no operator documentation that Kymo could locate, so the user agent recorded here comes from observed traffic rather than a vendor page.
Key facts
- cohere-ai is operated by Cohere, and its purpose is training: collects content that may be used to train future models.
- Cohere publishes no full user-agent string for cohere-ai, so Kymo matches the token cohere-ai observed in real logs.
- The robots.txt token for cohere-ai is cohere-ai, and Cohere publishes no compliance statement for cohere-ai.
- cohere-ai never sends a visitor to your site, and cohere-ai attaches no citation when it uses your content.
Specification
| Operator | Cohere |
|---|---|
| Purpose | TrainingCollects content that may be used to train future models. Blocking it removes your content from future training runs. It does not affect whether a model can find and cite you today. |
| Agent kind | CrawlerAutonomous crawler. Appears in your logs under its own user-agent. |
| User agent | Not documented by Cohere. Kymo matches this agent on the token cohere-ai, observed in real logs rather than published in operator documentation. |
| robots.txt token | cohere-ai |
| robots.txt compliance | No operator statement located |
| Verification | No published method, user agent is the only signal |
| Sends referral traffic | No |
| Attaches a citation | No |
| Status | Active |
| Legacy Kymo category | training |
| Source | Third-party source, not yet confirmed against operator docsCohere publishes no crawler documentation that Kymo could locate. Everything on this page comes from observed traffic, not from a vendor page. |
| Verified on |
What this looks like in your logs
grep -i 'cohere-ai' access.log
Cohere publishes no full user-agent string for cohere-ai, so the rest of the line varies and only the token cohere-ai is dependable. Kymo matches on that token, case-insensitively, for the same reason.
What cohere-ai does not do
cohere-ai does not send you a visitor, and it attaches no citation, so a read by it can never turn into a click.
cohere-ai does not decide whether Cohere can find you today, because a training run happens long before any question is asked.
Cohere publishes no documentation for cohere-ai at all, so no crawl frequency and no compliance statement exists to quote.
How Kymo classifies it
Kymo files cohere-ai under Cohere as training, classified from the HTTP request on the server rather than from anything running in a browser.
A hit from cohere-ai appears on the Crawlers dashboard and is deliberately left out of the AI Visibility numbers. A training run sends nobody, so a read by it has no click to compare against.
Seen on our own sites
| Last seen | 14 Sep 2026 |
|---|---|
| Days active, last 30 days | 3 |
| Requests, last 30 days | 6 |
Counted by Kymo on kymo.in and remotestack.in, the two sites we run, up to 15 Sep 2026. Two sites is a small sample. Treat these numbers as proof the bot is active now.
Allow it or block it
Documentation is thin. If no primary source can be found, label the user agent as observed rather than documented and say so on the page.
| Allow it if | Block it if |
|---|---|
| You want your content represented in future Cohere models, and you accept that a training run sends no visitor and no link back. | You sell access to your content, or you object to it being used as training data without payment or credit. |
robots.txt directives
Both blocks below address cohere-ai only. Rules for one token never apply to another, even from the same operator.
Allow cohere-ai
User-agent: cohere-ai Allow: /
Block cohere-ai
User-agent: cohere-ai Disallow: /
Questions
How do I know a request claiming to be cohere-ai is genuine?
Cohere publishes no verification method for cohere-ai, so the user-agent string is the only signal available, and anyone can send one. Treat a claimed cohere-ai request as unproven.
Does blocking cohere-ai affect my Google ranking?
No. A Google ranking is decided by Googlebot, which is a separate crawler. cohere-ai is Cohere's training crawler, so blocking it changes what Cohere can do with your content and changes nothing in Google search results.
Will I see cohere-ai in Google Analytics?
No. Google Analytics runs a JavaScript tag in a visitor's browser, and cohere-ai reads your HTML and leaves without running any script. Server-side logging is the only way to record it.
Does cohere-ai send traffic back to my site?
No. cohere-ai takes content and sends nothing back. A hit appears in your logs as a request with no visitor behind it.
Related crawlers
All AI crawlers and control tokens
See this bot in your own logs
cohere-ai takes your content and sends nothing back, so it leaves no trace at all in a browser-based analytics tool.
Kymo reads the HTTP request on your server, so a crawler that never runs JavaScript is still recorded. You see which bots reached the site, which pages they took and how often, next to your human traffic.
Start free → 14-day free trial. No card required.
No account yet? Run a free AI visibility check on your own site.