Field guide / Updated 2026-10-06

AI crawler list: user agents, purposes and robots.txt rules

Most AI companies run more than one bot: one feeds search answers, one collects training data, and one fetches a page when a person asks. Treat them separately and you can stay visible in AI search without offering your content for model training.

Three kinds of AI crawler

It helps to sort AI bots by job before deciding what to allow.

  • Search crawlers build the index that answer engines quote and link to. Examples: OAI-SearchBot for ChatGPT search, Claude-SearchBot for Claude, PerplexityBot for Perplexity, plus Googlebot and Bingbot for Google and Microsoft surfaces.
  • Training crawlers and controls decide whether your content may be used to train foundation models. Examples: GPTBot, ClaudeBot, CCBot, and the control tokens Google-Extended and Applebot-Extended.
  • User-initiated fetchers visit a page because someone asked a question inside the product. Examples: ChatGPT-User, Claude-User and Perplexity-User.

If your goal is to appear in AI answers, the first group matters most. Blocking a training crawler is a legitimate content decision, and on its own it does not remove you from the corresponding search product.

The reference table

Tokens are the names you use in robots.txt. Descriptions follow each operator's own documentation, reviewed on 2026-10-06.

Operator and tokenCategoryWhat the operator says it is forrobots.txt
OpenAI OAI-SearchBotSearchSurfacing websites in ChatGPT's search featuresRespected
OpenAI GPTBotTrainingContent that may be used to train OpenAI's foundation modelsRespected
OpenAI ChatGPT-UserUser fetchCertain user actions in ChatGPT and custom GPTsMay not apply
OpenAI OAI-AdsBotAds reviewChecking the safety of pages submitted as ChatGPT adsNot a crawl-control bot
Anthropic Claude-SearchBotSearchImproving the relevance and accuracy of Claude's search resultsRespected
Anthropic ClaudeBotTrainingContent that could contribute to model trainingRespected, with Crawl-delay
Anthropic Claude-UserUser fetchFetching pages when a person asks Claude a questionRespected
Perplexity PerplexityBotSearchSurfacing and linking websites in Perplexity; not used for foundation modelsRespected
Perplexity Perplexity-UserUser fetchVisiting a page to answer a user's questionGenerally ignored
Google GooglebotSearchGoogle Search, including eligibility for AI Overviews and AI ModeRespected
Google Google-ExtendedTraining controlA token that controls Gemini training and grounding use; it does not crawlToken only
Apple ApplebotSearchSearch features in Spotlight, Siri and SafariRespected
Apple Applebot-ExtendedTraining controlOpting out of training Apple's foundation models; it does not crawlToken only
Common Crawl CCBotOpen datasetThe public Common Crawl corpus, which many model builders useRespected

Microsoft's Bingbot indexes the web for Bing. Bing Webmaster Tools now reports when indexed pages are cited in Microsoft Copilot and Bing's AI-generated answers, so Bing access matters for AI visibility too.

A robots.txt for "visible in AI search, no training"

The example below allows the AI search crawlers, declines the training crawlers and control tokens, and leaves ordinary search engines alone. Replace the private path and the sitemap URL with your own.

# AI search and answer engines
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /
Disallow: /admin/

# Model training
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /

# Everyone else, including Googlebot and Bingbot
User-agent: *
Allow: /
Disallow: /admin/

Sitemap: https://example.com/sitemap.xml

One detail causes many mistakes: a crawler follows only the most specific group that names it. Once you give OAI-SearchBot its own group, the rules under User-agent: * no longer apply to it. Repeat every private-path Disallow inside each named group, as the example does for /admin/.

Common mistakes

  • Blocking the wrong OpenAI bot. Adding GPTBot to a block list keeps content out of training. It does not hide the site from ChatGPT search, and blocking OAI-SearchBot to stop training removes you from search answers instead.
  • Expecting Google-Extended to control AI Overviews. Google says Google-Extended has no effect on Search inclusion or ranking. Search features, including AI Overviews, depend on Googlebot and the usual snippet controls such as nosnippet and max-snippet.
  • Forgetting the CDN. A crawler that robots.txt allows can still be stopped by a firewall rule, a bot challenge or an AI crawler block at your CDN. Check those settings separately; see our Cloudflare AI Crawl Control guide.
  • Treating robots.txt as access control. robots.txt states a preference. User-initiated fetchers may not follow it, and a misbehaving bot can ignore it. Content that must stay private needs authentication.
  • Expecting instant results. OpenAI says changes can take about 24 hours to apply. Other operators publish no timing at all.

How to check who is really visiting

User-agent strings are easy to fake, so verify the source IP before trusting a log line. OpenAI publishes IP lists for each bot, for example https://openai.com/searchbot.json and https://openai.com/gptbot.json. Perplexity publishes https://www.perplexity.com/perplexitybot.json, Anthropic publishes https://claude.com/crawling/bots.json, and Common Crawl publishes https://index.commoncrawl.org/ccbot.json.

Compare those lists with your server or CDN logs, and look for HTTP 403 responses or challenge pages served to search crawlers. If you use Cloudflare, AI Crawl Control lists the AI crawlers that visited, how often, and how many times each one ignored your robots.txt.

When you are ready to write rules, our robots.txt rule generator builds a fragment from your choices in the browser. It does not fetch your site or prove that a crawler can get through.

Sources

All sources were reviewed on 2026-10-06. Operators change their crawlers and wording, so check the linked documentation before changing a live robots.txt file.

Frequently asked questions

Does blocking GPTBot remove my site from ChatGPT search?

No. GPTBot is OpenAI's training crawler. ChatGPT search uses OAI-SearchBot. OpenAI says sites that opt out of OAI-SearchBot are not shown in ChatGPT search answers, though they can still appear as navigational links.

Does Google-Extended affect AI Overviews or AI Mode?

No. Google describes Google-Extended as a product token, not a separate crawler, and says it does not affect inclusion or ranking in Google Search. AI Overviews and AI Mode use pages that Googlebot crawled and that are eligible to show a snippet.

Can robots.txt stop a user-initiated AI fetch?

Not reliably. OpenAI says robots.txt rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores robots.txt because a person asked for the page. Protect private content with authentication or server rules instead.

How quickly do robots.txt changes take effect?

It depends on the operator. OpenAI says it can take about 24 hours for its systems to adjust after a robots.txt update. Check your server or CDN logs after any change rather than assuming it applied immediately.