Cloudflare AI Crawl Control: allow AI search, block training
Cloudflare sits in front of a large share of websites, so its settings often decide whether AI crawlers ever reach a page. The good news is that Cloudflare treats crawlers individually. You can keep AI search crawlers in and training crawlers out, as long as you know which setting does what.
The three settings that matter
Three Cloudflare features can affect AI crawlers. They do different things, and it is easy to switch on one while expecting the effect of another.
| Setting | What it does | Effect on AI search crawlers |
|---|---|---|
| AI Crawl Control, per crawler | Allow, block or (in closed beta) charge each AI crawler. Blocking creates a WAF custom rule that answers with 403 or 402. | Decides access. Blocking a search crawler here stops it, whatever robots.txt says. |
| Managed robots.txt | Prepends Cloudflare-maintained rules that disallow known AI training crawlers, ahead of your own robots.txt. | Expresses a preference. The documented list targets training crawlers, not the AI search crawlers. |
| Bot Fight Mode and other bot protection | Challenges or slows traffic that looks automated. | Can interfere with any crawler if misconfigured. Review its effect in your logs. |
The dashboard groups AI bots by category, such as "AI Crawler", "AI Search" and "AI Assistant". Use those labels to decide which ones to block.
A setup for "visible in AI search, no training"
- Open AI Crawl Control for your domain and review the crawler list.
- Keep the AI search crawlers allowed, for example
OAI-SearchBot,PerplexityBotandClaude-SearchBot, along withGooglebot,BingbotandApplebot. - Block the training crawlers you do not want, for example
GPTBot,ClaudeBot,CCBotandBytespider, or turn on the managed robots.txt to state that preference for the documented training crawlers. - Decide on the user-initiated assistants, such as
ChatGPT-User,Claude-UserandPerplexity-User. They fetch a page when a person asks about it. Blocking them can reduce how often your page is used in those answers. - Write your own robots.txt rules to match, so the policy is visible to every crawler, not only those Cloudflare enforces. Our robots.txt rule generator builds a fragment in your browser.
Google-Extended and Applebot-Extended are control tokens rather than crawlers. They only work through robots.txt and do not appear as separate traffic.
How the managed robots.txt merges with yours
When the managed setting is on and your site already serves a robots.txt file, Cloudflare prepends its managed rules and returns both as one file. According to Cloudflare's documentation, the managed list disallows Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagent, and adds content signals stating that search use is allowed and AI training is not.
Two consequences follow:
- Your own rules for the same bots may now conflict with the prepended block. Fetch
https://your-site.com/robots.txtafter enabling it and read the combined result. - The managed file is a preference. Cloudflare says compliance with robots.txt is voluntary and recommends AI Crawl Control when you need enforcement.
Check the result
- Fetch your robots.txt and confirm the final text matches your intent.
- In AI Crawl Control, review which crawlers visited, how often, and whether any ignored your robots.txt. The violations column helps you decide whether a preference is enough or a block is needed.
- In your logs, confirm that search crawlers receive HTTP 200 responses rather than 403, 402 or challenge pages.
- Use Google Search Console and Bing Webmaster Tools to confirm key pages remain indexed after any change.
Common mistakes
- Blocking every crawler in the "AI" categories, including search crawlers, and then wondering why ChatGPT search or Perplexity never cites the site.
- Turning on the managed robots.txt and believing it blocks access. It only asks.
- Relying on user-agent strings alone. Verify a bot against its operator's published IP list before allowing it through a strict firewall rule.
- Forgetting staging or alternate hostnames that serve the same content with different settings.
For the full list of AI crawlers and what each operator says it does, see our AI crawler reference.
Sources
- Cloudflare: Manage AI crawlers
- Cloudflare: Managed robots.txt
- Cloudflare blog: Content Independence Day
- OpenAI: Overview of OpenAI crawlers
Reviewed on 2026-10-06. Cloudflare's dashboard labels change over time; the linked documentation describes the current behavior.
Frequently asked questions
Does Cloudflare block AI crawlers by default?
Since July 2025 Cloudflare has asked new domains during setup whether to allow AI crawlers, and its announcement describes blocking as the new default. Older domains keep their existing settings. Check AI Crawl Control for your zone rather than assuming either way.
Does Cloudflare's managed robots.txt block ChatGPT search?
Not according to Cloudflare's documentation. The managed file disallows training-oriented crawlers such as GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot, Bytespider, Amazonbot and meta-externalagent. OAI-SearchBot, Claude-SearchBot and PerplexityBot are not in that list.
What is the difference between robots.txt and blocking in AI Crawl Control?
robots.txt states a preference that well-behaved crawlers follow voluntarily. Blocking in AI Crawl Control creates a WAF rule that refuses the request, with a 403 or 402 response. Cloudflare suggests using the two together: robots.txt to state the policy and AI Crawl Control to enforce it.