AI crawler list: user agents, purposes and robots.txt rules
Most AI companies run more than one bot: one feeds search answers, one collects training data, and one fetches a page when a person asks. Treat them separately and you can stay visible in AI search without offering your content for model training.
Three kinds of AI crawler
It helps to sort AI bots by job before deciding what to allow.
- Search crawlers build the index that answer engines quote and link to. Examples:
OAI-SearchBotfor ChatGPT search,Claude-SearchBotfor Claude,PerplexityBotfor Perplexity, plusGooglebotandBingbotfor Google and Microsoft surfaces. - Training crawlers and controls decide whether your content may be used to train foundation models. Examples:
GPTBot,ClaudeBot,CCBot, and the control tokensGoogle-ExtendedandApplebot-Extended. - User-initiated fetchers visit a page because someone asked a question inside the product. Examples:
ChatGPT-User,Claude-UserandPerplexity-User.
If your goal is to appear in AI answers, the first group matters most. Blocking a training crawler is a legitimate content decision, and on its own it does not remove you from the corresponding search product.
The reference table
Tokens are the names you use in robots.txt. Descriptions follow each operator's own documentation, reviewed on 2026-10-06.
| Operator and token | Category | What the operator says it is for | robots.txt |
|---|---|---|---|
OpenAI OAI-SearchBot | Search | Surfacing websites in ChatGPT's search features | Respected |
OpenAI GPTBot | Training | Content that may be used to train OpenAI's foundation models | Respected |
OpenAI ChatGPT-User | User fetch | Certain user actions in ChatGPT and custom GPTs | May not apply |
OpenAI OAI-AdsBot | Ads review | Checking the safety of pages submitted as ChatGPT ads | Not a crawl-control bot |
Anthropic Claude-SearchBot | Search | Improving the relevance and accuracy of Claude's search results | Respected |
Anthropic ClaudeBot | Training | Content that could contribute to model training | Respected, with Crawl-delay |
Anthropic Claude-User | User fetch | Fetching pages when a person asks Claude a question | Respected |
Perplexity PerplexityBot | Search | Surfacing and linking websites in Perplexity; not used for foundation models | Respected |
Perplexity Perplexity-User | User fetch | Visiting a page to answer a user's question | Generally ignored |
Google Googlebot | Search | Google Search, including eligibility for AI Overviews and AI Mode | Respected |
Google Google-Extended | Training control | A token that controls Gemini training and grounding use; it does not crawl | Token only |
Apple Applebot | Search | Search features in Spotlight, Siri and Safari | Respected |
Apple Applebot-Extended | Training control | Opting out of training Apple's foundation models; it does not crawl | Token only |
Common Crawl CCBot | Open dataset | The public Common Crawl corpus, which many model builders use | Respected |
Microsoft's Bingbot indexes the web for Bing. Bing Webmaster Tools now reports when indexed pages are cited in Microsoft Copilot and Bing's AI-generated answers, so Bing access matters for AI visibility too.
A robots.txt for "visible in AI search, no training"
The example below allows the AI search crawlers, declines the training crawlers and control tokens, and leaves ordinary search engines alone. Replace the private path and the sitemap URL with your own.
# AI search and answer engines
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /
Disallow: /admin/
# Model training
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /
# Everyone else, including Googlebot and Bingbot
User-agent: *
Allow: /
Disallow: /admin/
Sitemap: https://example.com/sitemap.xmlOne detail causes many mistakes: a crawler follows only the most specific group that names it. Once you give OAI-SearchBot its own group, the rules under User-agent: * no longer apply to it. Repeat every private-path Disallow inside each named group, as the example does for /admin/.
Common mistakes
- Blocking the wrong OpenAI bot. Adding
GPTBotto a block list keeps content out of training. It does not hide the site from ChatGPT search, and blockingOAI-SearchBotto stop training removes you from search answers instead. - Expecting Google-Extended to control AI Overviews. Google says Google-Extended has no effect on Search inclusion or ranking. Search features, including AI Overviews, depend on Googlebot and the usual snippet controls such as
nosnippetandmax-snippet. - Forgetting the CDN. A crawler that robots.txt allows can still be stopped by a firewall rule, a bot challenge or an AI crawler block at your CDN. Check those settings separately; see our Cloudflare AI Crawl Control guide.
- Treating robots.txt as access control. robots.txt states a preference. User-initiated fetchers may not follow it, and a misbehaving bot can ignore it. Content that must stay private needs authentication.
- Expecting instant results. OpenAI says changes can take about 24 hours to apply. Other operators publish no timing at all.
How to check who is really visiting
User-agent strings are easy to fake, so verify the source IP before trusting a log line. OpenAI publishes IP lists for each bot, for example https://openai.com/searchbot.json and https://openai.com/gptbot.json. Perplexity publishes https://www.perplexity.com/perplexitybot.json, Anthropic publishes https://claude.com/crawling/bots.json, and Common Crawl publishes https://index.commoncrawl.org/ccbot.json.
Compare those lists with your server or CDN logs, and look for HTTP 403 responses or challenge pages served to search crawlers. If you use Cloudflare, AI Crawl Control lists the AI crawlers that visited, how often, and how many times each one ignored your robots.txt.
When you are ready to write rules, our robots.txt rule generator builds a fragment from your choices in the browser. It does not fetch your site or prove that a crawler can get through.
Sources
- OpenAI: Overview of OpenAI crawlers
- Anthropic: Does Anthropic crawl the web, and how can site owners block the crawler?
- Perplexity: Perplexity crawlers
- Google: Google's common crawlers, including Google-Extended
- Apple: About Applebot
- Common Crawl: CCBot
- Microsoft Bing: Introducing AI Performance in Bing Webmaster Tools
All sources were reviewed on 2026-10-06. Operators change their crawlers and wording, so check the linked documentation before changing a live robots.txt file.
Frequently asked questions
Does blocking GPTBot remove my site from ChatGPT search?
No. GPTBot is OpenAI's training crawler. ChatGPT search uses OAI-SearchBot. OpenAI says sites that opt out of OAI-SearchBot are not shown in ChatGPT search answers, though they can still appear as navigational links.
Does Google-Extended affect AI Overviews or AI Mode?
No. Google describes Google-Extended as a product token, not a separate crawler, and says it does not affect inclusion or ranking in Google Search. AI Overviews and AI Mode use pages that Googlebot crawled and that are eligible to show a snippet.
Can robots.txt stop a user-initiated AI fetch?
Not reliably. OpenAI says robots.txt rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores robots.txt because a person asked for the page. Protect private content with authentication or server rules instead.
How quickly do robots.txt changes take effect?
It depends on the operator. OpenAI says it can take about 24 hours for its systems to adjust after a robots.txt update. Check your server or CDN logs after any change rather than assuming it applied immediately.