AI crawler guide

Every AI crawler, who runs it, and what it feeds.

AI assistants reach your site through named crawlers. Some build the search index behind live answers, some fetch a page the moment a person asks, and some collect training data. Each entry links to its operator's own documentation and the robots.txt lines to allow or block it.

The list

AI crawlers by what they feed.

16 crawlers and tokens from OpenAI, Anthropic, Perplexity, Google, Apple, Amazon, Common Crawl, ByteDance and Meta. Open any token for its user agent, the exact robots.txt lines and how it applies them.

Assistant fetchers

These fetch a page at the moment a person asks the assistant about it, so the answer can quote current content.

TokenOperatorWhat it feedsKindDocsVerified
ChatGPT-UserOpenAIuser actions in ChatGPT and Custom GPTsCrawlerOpenAI docs for ChatGPT-User (opens in a new tab)
Claude-UserAnthropiclive answers in ClaudeCrawlerAnthropic docs for Claude-User (opens in a new tab)
Perplexity-UserPerplexitylive answers in PerplexityCrawlerPerplexity docs for Perplexity-User (opens in a new tab)
Amzn-UserAmazonlive answers in AlexaCrawlerAmazon docs for Amzn-User (opens in a new tab)

Training crawlers and tokens

These govern whether your content is used to train AI models. Training access is your call and is separate from live answers.

TokenOperatorWhat it feedsKindDocsVerified
GPTBotOpenAIOpenAI generative AI foundation modelsCrawlerOpenAI docs for GPTBot (opens in a new tab)
ClaudeBotAnthropicAnthropic's Claude model trainingCrawlerAnthropic docs for ClaudeBot (opens in a new tab)
Google-ExtendedGoogleGemini model training and groundingControl tokenGoogle docs for Google-Extended (opens in a new tab)
Applebot-ExtendedAppleApple generative AI model trainingControl tokenApple docs for Applebot-Extended (opens in a new tab)
AmazonbotAmazonAmazon products, services and AI modelsCrawlerAmazon docs for Amazonbot (opens in a new tab)
CCBotCommon Crawlthe Common Crawl open web datasetCrawlerCommon Crawl docs for CCBot (opens in a new tab)
BytespiderByteDanceByteDance AI modelsCrawlerNo official page
Meta-ExternalAgentMetaMeta AI model training and productsCrawlerMeta docs for Meta-ExternalAgent (opens in a new tab)

robots.txt snippet

Allow answer bots, keep training off.

Paste this into the robots.txt at the root of your domain. It opens every search crawler and assistant fetcher above and closes the training crawlers and tokens.

Allow answer bots, keep training off
# AI crawler access, generated by Suede AI Agent Studio

User-agent: OAI-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: Claude-User
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Perplexity-User
Allow: /

User-agent: Amzn-SearchBot
Allow: /

User-agent: Amzn-User
Allow: /

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: Amazonbot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Bytespider
Disallow: /

User-agent: Meta-ExternalAgent
Disallow: /

A group that names a crawler replaces your * group for that crawler, so add any * Disallow lines you still want applied under each group. The AI visibility fix kit writes these groups for your site with your * Disallow lines already merged in.

Prefer to allow training too?
Allow every AI crawler
# AI crawler access, generated by Suede AI Agent Studio

User-agent: OAI-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: Claude-User
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Perplexity-User
Allow: /

User-agent: Amzn-SearchBot
Allow: /

User-agent: Amzn-User
Allow: /

User-agent: GPTBot
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: Google-Extended
Allow: /

User-agent: Applebot-Extended
Allow: /

User-agent: Amazonbot
Allow: /

User-agent: CCBot
Allow: /

User-agent: Bytespider
Allow: /

User-agent: Meta-ExternalAgent
Allow: /

Precedence

How a crawler picks its rules.

Every crawler above applies the same three rules from the robots.txt standard (RFC 9309).

  1. A named group beats *. A group whose User-agent line names the crawler is the only group it follows. With no named group, it follows User-agent: *.
  2. The longest matching rule wins. Inside that group, the Allow or Disallow line with the longest matching path decides.
  3. Allow wins a tie. When an Allow and a Disallow match with equal length, the path is allowed.

Suede's own robots.txt invites these crawlers, so every public page here is open to AI search and assistants.

Last updated .

See which crawlers can read your site. Line by line.

The free AI visibility check reads your robots.txt, names the exact line that decides each crawler above, and ranks your fixes. The $9.99 fix kit then writes the robots.txt lines, llms.txt and business schema for your site.

Inside Agent Studio

Good work starts with a clear flow.

Give an agent a job. Connect the steps. See how the work becomes a service someone else can call.

Input
Reason
Output

A workflow you can inspect

Make the steps visible.

Start with a template or describe the job. Connect inputs, reasoning, and output on the canvas.

Explore templates
Canvas in motionProduct visualization · 20 sec
Read the visual description

A stylized organization chart opens into a workflow canvas. Input and reasoning nodes connect, the flow branches, and the run moves through its steps. The film ends with Suede AI Agent Studio. Music only; no spoken narration. Illustrative interface, not a live run.