AI crawler guide
Every AI crawler, who runs it, and what it feeds.
AI assistants reach your site through named crawlers. Some build the search index behind live answers, some fetch a page the moment a person asks, and some collect training data. Each entry links to its operator's own documentation and the robots.txt lines to allow or block it.
The list
AI crawlers by what they feed.
16 crawlers and tokens from OpenAI, Anthropic, Perplexity, Google, Apple, Amazon, Common Crawl, ByteDance and Meta. Open any token for its user agent, the exact robots.txt lines and how it applies them.
Search crawlers
These build the search indexes behind AI answers. Allowing them makes your pages eligible to appear and be linked when someone asks.
| Token | Operator | What it feeds | Kind | Docs | Verified |
|---|---|---|---|---|---|
| OAI-SearchBot | OpenAI | ChatGPT search results | Crawler | OpenAI docs for OAI-SearchBot (opens in a new tab) | |
| Claude-SearchBot | Anthropic | Claude search results | Crawler | Anthropic docs for Claude-SearchBot (opens in a new tab) | |
| PerplexityBot | Perplexity | Perplexity search results | Crawler | Perplexity docs for PerplexityBot (opens in a new tab) | |
| Amzn-SearchBot | Amazon | search in Amazon products and services | Crawler | Amazon docs for Amzn-SearchBot (opens in a new tab) |
Assistant fetchers
These fetch a page at the moment a person asks the assistant about it, so the answer can quote current content.
| Token | Operator | What it feeds | Kind | Docs | Verified |
|---|---|---|---|---|---|
| ChatGPT-User | OpenAI | user actions in ChatGPT and Custom GPTs | Crawler | OpenAI docs for ChatGPT-User (opens in a new tab) | |
| Claude-User | Anthropic | live answers in Claude | Crawler | Anthropic docs for Claude-User (opens in a new tab) | |
| Perplexity-User | Perplexity | live answers in Perplexity | Crawler | Perplexity docs for Perplexity-User (opens in a new tab) | |
| Amzn-User | Amazon | live answers in Alexa | Crawler | Amazon docs for Amzn-User (opens in a new tab) |
Training crawlers and tokens
These govern whether your content is used to train AI models. Training access is your call and is separate from live answers.
| Token | Operator | What it feeds | Kind | Docs | Verified |
|---|---|---|---|---|---|
| GPTBot | OpenAI | OpenAI generative AI foundation models | Crawler | OpenAI docs for GPTBot (opens in a new tab) | |
| ClaudeBot | Anthropic | Anthropic's Claude model training | Crawler | Anthropic docs for ClaudeBot (opens in a new tab) | |
| Google-Extended | Gemini model training and grounding | Control token | Google docs for Google-Extended (opens in a new tab) | ||
| Applebot-Extended | Apple | Apple generative AI model training | Control token | Apple docs for Applebot-Extended (opens in a new tab) | |
| Amazonbot | Amazon | Amazon products, services and AI models | Crawler | Amazon docs for Amazonbot (opens in a new tab) | |
| CCBot | Common Crawl | the Common Crawl open web dataset | Crawler | Common Crawl docs for CCBot (opens in a new tab) | |
| Bytespider | ByteDance | ByteDance AI models | Crawler | No official page | |
| Meta-ExternalAgent | Meta | Meta AI model training and products | Crawler | Meta docs for Meta-ExternalAgent (opens in a new tab) |
robots.txt snippet
Allow answer bots, keep training off.
Paste this into the robots.txt at the root of your domain. It opens every search crawler and assistant fetcher above and closes the training crawlers and tokens.
# AI crawler access, generated by Suede AI Agent Studio
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Perplexity-User
Allow: /
User-agent: Amzn-SearchBot
Allow: /
User-agent: Amzn-User
Allow: /
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: Amazonbot
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: Bytespider
Disallow: /
User-agent: Meta-ExternalAgent
Disallow: /
A group that names a crawler replaces your * group for that crawler, so add any * Disallow lines you still want applied under each group. The AI visibility fix kit writes these groups for your site with your * Disallow lines already merged in.
Prefer to allow training too?
# AI crawler access, generated by Suede AI Agent Studio
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Perplexity-User
Allow: /
User-agent: Amzn-SearchBot
Allow: /
User-agent: Amzn-User
Allow: /
User-agent: GPTBot
Allow: /
User-agent: ClaudeBot
Allow: /
User-agent: Google-Extended
Allow: /
User-agent: Applebot-Extended
Allow: /
User-agent: Amazonbot
Allow: /
User-agent: CCBot
Allow: /
User-agent: Bytespider
Allow: /
User-agent: Meta-ExternalAgent
Allow: /
Precedence
How a crawler picks its rules.
Every crawler above applies the same three rules from the robots.txt standard (RFC 9309).
- A named group beats *. A group whose
User-agentline names the crawler is the only group it follows. With no named group, it followsUser-agent: *. - The longest matching rule wins. Inside that group, the Allow or Disallow line with the longest matching path decides.
- Allow wins a tie. When an Allow and a Disallow match with equal length, the path is allowed.
Suede's own robots.txt invites these crawlers, so every public page here is open to AI search and assistants.
Last updated .
See which crawlers can read your site. Line by line.
The free AI visibility check reads your robots.txt, names the exact line that decides each crawler above, and ranks your fixes. The $9.99 fix kit then writes the robots.txt lines, llms.txt and business schema for your site.