SearchLink

Guides ยท Updated 2026-09-28

robots.txt for AI search: which bots to allow (2026)

Many sites block AI crawlers to keep their content out of training, and by accident also block the bots that let AI answers cite and link to them. The two are separate, and you can allow one while blocking the other.

The short answer

If you want to be recommended and linked in AI answers, allow the search crawlers. Blocking the training crawlers is your choice and does not remove you from AI search.

robots.txt nameCompanyWhat it doesBlock it and
OAI-SearchBotOpenAIFinds pages to show and link in ChatGPT searchYou stop appearing in ChatGPT search answers
ChatGPT-UserOpenAIVisits a page when a ChatGPT user asks about itChatGPT cannot open your page for a user
GPTBotOpenAICollects content for training modelsKept out of training; search is not affected
PerplexityBotPerplexitySurfaces and links sites in Perplexity answers (Perplexity says it is not used for training)You stop appearing in Perplexity answers
Perplexity-UserPerplexityVisits a page to answer a user's questionPerplexity cannot read your page for a user
Claude-SearchBotAnthropicImproves Claude's search resultsYour pages may show up less in Claude's search answers
Claude-UserAnthropicVisits a page when a Claude user asksClaude cannot open your page for a user
ClaudeBotAnthropicCollects content that may be used for trainingKept out of future training
GooglebotGoogleGoogle Search, including AI Overviews and AI ModeYou disappear from Google, AI features included
Google-ExtendedGoogleA control token only: whether Gemini apps and Vertex AI may use your content. Not a separate crawlerGoogle Search and AI Overviews are not affected
BingbotMicrosoftBing Search, which also feeds Microsoft CopilotYou disappear from Bing and Copilot answers

A robots.txt you can copy

This allows every search crawler and blocks only training. Remove the training lines if you are happy to be used for training.

User-agent: *
Allow: /

# Keep content out of AI training (search is not affected)
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

Sitemap: https://yourwebsite.com/sitemap.xml

Three mistakes we see often

  1. Blocking everything with one line. A leftover User-agent: * / Disallow: / from a staging site hides you from Google and every AI engine at once.
  2. A firewall that blocks what robots.txt allows. Bot protection (for example Cloudflare Bot Fight Mode) can turn away AI search crawlers even when robots.txt lets them in. Google usually gets through; smaller crawlers often do not.
  3. Thinking Google-Extended controls AI Overviews. It does not. AI Overviews and AI Mode are part of Google Search and follow Googlebot.

Allowed is step one

Being allowed in only means AI engines can read you. Whether they actually cite you depends on whether your page answers the question directly, near the top, with facts and dates. SearchLink checks that for the searches your site really gets in Google, and tells you which sites the AI cited instead.

Check your own site in 30 seconds

Free, no sign-up. See which search and AI crawlers your robots.txt lets in, and the 3 things to fix first.

Sources: OpenAI "Overview of OpenAI Crawlers", Perplexity "Perplexity Crawlers", Anthropic support "Does Anthropic crawl data from the web", Google Search Central "Google's common crawlers". Checked 2026-09-28.