AI Crawlers in robots.txt: 18.8% of the Top 10,000 Sites Block One, Mostly Training Bots, Not AI Search

We parsed robots.txt and llms.txt on the 10,000 most popular sites. 18.8% block at least one AI crawler; GPTBot is blocked 2.5x as often as OpenAI's search crawler; 37 sites block only Anthropic's retired token. Breakdowns by rank and TLD, plus the CSV.

English Implementation & Architecture SummaryBase URL: https://api.callaiapi.com/v1

We parsed robots.txt and llms.txt on the 10,000 most popular sites. 18.8% block at least one AI crawler; GPTBot is blocked 2.5x as often as OpenAI's search crawler; 37 sites block only Anthropic's retired token. Breakdowns by rank and TLD, plus the CSV.

Auth Header: Bearer <YOUR_API_KEY>•Billing: TRON USDT from $1 · Pay-as-you-go•Compatibility: 100% OpenAI SDK Drop-in

More sites are blocking AI crawlers in robots.txt, but AI crawlers are not one thing. Some collect training data; others fetch pages live for AI search such as ChatGPT search and Perplexity. Block the wrong one and you may fail to stop training while removing yourself from AI search.

On 28 September 2026 we fetched robots.txt and llms.txt from the top 10,000 sites in the Tranco ranking and evaluated, per RFC 9309, whether 15 AI crawlers and 2 search-engine crawlers may fetch site content.

Key findings

  • Of 4,592 sites with a robots.txt, 18.8% block at least one AI crawler while allowing Googlebot. Only 2% block all 15.
  • Training crawlers are blocked most: CCBot 14.3%, GPTBot 13.4%, Bytespider 13.3%, ClaudeBot 12.3%.
  • AI search crawlers are blocked far less: OpenAI's OAI-SearchBot 5.3%, two-fifths of GPTBot's rate. Of the 615 sites blocking GPTBot, 373 (60.7%) still allow OAI-SearchBot.
  • 9.4% of sites "block training, allow search": at least one training crawler blocked, every AI search crawler allowed.
  • Bigger sites block more: 28.1% of the top 1,000 versus 16.7% of those ranked 5,001-10,000. German .de sites reach 40%.
  • 37 sites block only anthropic-ai, a token Anthropic has retired, while leaving ClaudeBot, the crawler actually in use, unblocked. Examples include espn.com, science.org and trustpilot.com.
  • 10.7% of reachable sites serve an llms.txt; at least 30 of those files were generated by the Yoast SEO WordPress plugin.

Block rate by crawler

Base: the 4,592 sites with a robots.txt. "Blocked" means the homepage or ordinary content pages are disallowed; sites that also block Googlebot are excluded.

CrawlerOperatorPurposeBlocked
CCBotCommon CrawlTraining14.3%
GPTBotOpenAITraining13.4%
BytespiderByteDanceTraining13.3%
ClaudeBotAnthropicTraining12.3%
meta-externalagentMetaTraining10.6%
Google-ExtendedGoogleTraining10.5%
Applebot-ExtendedAppleTraining9.8%
PerplexityBotPerplexitySearch8.9%
ChatGPT-UserOpenAIUser7.4%
OAI-SearchBotOpenAISearch5.3%
Claude-SearchBotAnthropicSearch5.2%
GooglebotGoogle—3.3%

By rank

RankSites with robots.txtBlock any AI crawlerBlock GPTBot
1–1,00050628.1%19.2%
1,001–5,0001,91018.7%13.6%
5,001–10,0002,17616.7%11.9%

Common mistakes

  • Using the old token: anthropic-ai is retired; to block Anthropic's training crawler, name ClaudeBot.
  • Homepage-only access: Disallow: / plus Allow: /$ opens the homepage and closes every content page, which is effectively a full block.
  • Assuming a GPTBot block removes you from ChatGPT: GPTBot is for training, ChatGPT search uses OAI-SearchBot, and pages a user asks ChatGPT to open are fetched as ChatGPT-User. Treat them separately.
  • Order does not matter: under RFC 9309 a crawler uses the group that names it and only falls back to * if none does.
# One way to block training but allow AI search
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: CCBot
Disallow: /

User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /

User-agent: *
Disallow: /admin/

Methodology and limits

  • Sites: the top 10,000 registrable domains of Tranco list 64X3X, retrieved 28 September 2026. 3,687 were unreachable (mostly CDN or API hosts without a website, or 5xx responses) and are excluded.
  • Fetching: 28 September 2026, https first then http, following redirects, User-Agent CallAI-Research/1.0.
  • Evaluation follows RFC 9309: a named group overrides *, groups for the same crawler merge, the longest match wins, Allow wins ties, and * and $ wildcards are supported.
  • robots.txt is voluntary: a block does not guarantee compliance, and some sites also filter by User-Agent at the server or CDN, which this study does not measure.
  • llms.txt counts when the file returns 2xx, is not HTML and starts with "# ".
站内推荐实用工具

Check your robots.txt

Free robots.txt checker: see whether Googlebot, GPTBot, ClaudeBot and other crawlers may fetch any path.

Check now