Guides ยท Updated 2026-09-28
robots.txt for AI search: which bots to allow (2026)
Many sites block AI crawlers to keep their content out of training, and by accident also block the bots that let AI answers cite and link to them. The two are separate, and you can allow one while blocking the other.
The short answer
If you want to be recommended and linked in AI answers, allow the search crawlers. Blocking the training crawlers is your choice and does not remove you from AI search.
| robots.txt name | Company | What it does | Block it and |
|---|---|---|---|
OAI-SearchBot | OpenAI | Finds pages to show and link in ChatGPT search | You stop appearing in ChatGPT search answers |
ChatGPT-User | OpenAI | Visits a page when a ChatGPT user asks about it | ChatGPT cannot open your page for a user |
GPTBot | OpenAI | Collects content for training models | Kept out of training; search is not affected |
PerplexityBot | Perplexity | Surfaces and links sites in Perplexity answers (Perplexity says it is not used for training) | You stop appearing in Perplexity answers |
Perplexity-User | Perplexity | Visits a page to answer a user's question | Perplexity cannot read your page for a user |
Claude-SearchBot | Anthropic | Improves Claude's search results | Your pages may show up less in Claude's search answers |
Claude-User | Anthropic | Visits a page when a Claude user asks | Claude cannot open your page for a user |
ClaudeBot | Anthropic | Collects content that may be used for training | Kept out of future training |
Googlebot | Google Search, including AI Overviews and AI Mode | You disappear from Google, AI features included | |
Google-Extended | A control token only: whether Gemini apps and Vertex AI may use your content. Not a separate crawler | Google Search and AI Overviews are not affected | |
Bingbot | Microsoft | Bing Search, which also feeds Microsoft Copilot | You disappear from Bing and Copilot answers |
A robots.txt you can copy
This allows every search crawler and blocks only training. Remove the training lines if you are happy to be used for training.
User-agent: *
Allow: /
# Keep content out of AI training (search is not affected)
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
Sitemap: https://yourwebsite.com/sitemap.xml
Three mistakes we see often
- Blocking everything with one line. A leftover
User-agent: * / Disallow: /from a staging site hides you from Google and every AI engine at once. - A firewall that blocks what robots.txt allows. Bot protection (for example Cloudflare Bot Fight Mode) can turn away AI search crawlers even when robots.txt lets them in. Google usually gets through; smaller crawlers often do not.
- Thinking Google-Extended controls AI Overviews. It does not. AI Overviews and AI Mode are part of Google Search and follow Googlebot.
Allowed is step one
Being allowed in only means AI engines can read you. Whether they actually cite you depends on whether your page answers the question directly, near the top, with facts and dates. SearchLink checks that for the searches your site really gets in Google, and tells you which sites the AI cited instead.
Check your own site in 30 seconds
Free, no sign-up. See which search and AI crawlers your robots.txt lets in, and the 3 things to fix first.
Sources: OpenAI "Overview of OpenAI Crawlers", Perplexity "Perplexity Crawlers", Anthropic support "Does Anthropic crawl data from the web", Google Search Central "Google's common crawlers". Checked 2026-09-28.