Omgilibot: Omgili's Training Crawler
What is Omgilibot?
Omgilibot collects web content for a data-licensing corpus that has been resold to model builders. Omgilibot collects content for a data-licensing corpus that has been resold to model builders, so the company using your content is not the one crawling it.
Key facts
- Omgilibot is operated by Omgili, and its purpose is dataset: collects content into a public dataset that many model builders draw from.
- Omgili publishes no full user-agent string for Omgilibot, so Kymo matches the token Omgilibot observed in real logs.
- The robots.txt token for Omgilibot is Omgilibot, and Omgili publishes no compliance statement for Omgilibot.
- Omgilibot never sends a visitor to your site, and Omgilibot attaches no citation when it uses your content.
Specification
| Operator | Omgili |
|---|---|
| Purpose | DatasetCollects content into a public dataset that many model builders draw from. Blocking it has broad downstream effect and no direct citation cost. |
| Agent kind | CrawlerAutonomous crawler. Appears in your logs under its own user-agent. |
| User agent | Not documented by Omgili. Kymo matches this agent on the token Omgilibot, observed in real logs rather than published in operator documentation. |
| robots.txt token | Omgilibot |
| robots.txt compliance | No operator statement located |
| Verification | No published method, user agent is the only signal |
| Sends referral traffic | No |
| Attaches a citation | No |
| Status | Active |
| Legacy Kymo category | training |
| Source | Observed in real logs, no operator documentation locatedOmgili publishes no crawler documentation that Kymo could locate. Everything on this page comes from observed traffic, not from a vendor page. |
| Verified on |
What this looks like in your logs
grep -i 'Omgilibot' access.log
Omgili publishes no full user-agent string for Omgilibot, so the rest of the line varies and only the token Omgilibot is dependable. Kymo matches on that token, case-insensitively, for the same reason.
What Omgilibot does not do
Omgilibot does not send you a visitor, and it attaches no citation, so a read by it can never turn into a click.
Omgilibot does not decide whether any particular product cites you, because it only fills a dataset that other companies choose to use or ignore.
Omgili publishes no documentation for Omgilibot at all, so no crawl frequency and no compliance statement exists to quote.
How Kymo classifies it
Kymo files Omgilibot under Omgili as training, classified from the HTTP request on the server rather than from anything running in a browser.
A hit from Omgilibot appears on the Crawlers dashboard and is deliberately left out of the AI Visibility numbers. A training run sends nobody, so a read by it has no click to compare against.
Seen on our own sites
Not seen on kymo.in or remotestack.in in the 90 days up to 15 Sep 2026.
Allow it or block it
Blocking it removes your content from a corpus sold on to third parties. There is no live citation surface, so there is no direct visibility cost.
| Allow it if | Block it if |
|---|---|
| You accept your content entering a dataset that many companies draw on, not only Omgili. | You want to limit how far your content travels, because a dataset reaches further than any single crawler. |
robots.txt directives
Both blocks below address Omgilibot only. Rules for one token never apply to another, even from the same operator.
Allow Omgilibot
User-agent: Omgilibot Allow: /
Block Omgilibot
User-agent: Omgilibot Disallow: /
Questions
How do I know a request claiming to be Omgilibot is genuine?
Omgili publishes no verification method for Omgilibot, so the user-agent string is the only signal available, and anyone can send one. Treat a claimed Omgilibot request as unproven.
Does blocking Omgilibot affect my Google ranking?
No. A Google ranking is decided by Googlebot, which is a separate crawler. Omgilibot collects content into a dataset that other companies license, and Google does not read that dataset to rank pages.
Will I see Omgilibot in Google Analytics?
No. Google Analytics runs a JavaScript tag in a visitor's browser, and Omgilibot reads your HTML and leaves without running any script. Server-side logging is the only way to record it.
Does Omgilibot send traffic back to my site?
No. Omgilibot takes content and sends nothing back. A hit appears in your logs as a request with no visitor behind it.
Related crawlers
All AI crawlers and control tokens
See this bot in your own logs
Omgilibot copies your content into a dataset other companies then use, and a browser-based analytics tool records none of it.
Kymo reads the HTTP request on your server, so a crawler that never runs JavaScript is still recorded. You see which bots reached the site, which pages they took and how often, next to your human traffic.
Start free → 14-day free trial. No card required.
No account yet? Run a free AI visibility check on your own site.