OpenAI · Training crawler guide
GPTBot: who runs it and how to allow or block it
GPTBot is the crawler OpenAI uses to make its generative AI foundation models more useful and safe.
GPTBot at a glance
- Operator
- OpenAI
- Purpose
- Training It collects content for AI model training.
- Kind
- Crawler. Visits your pages under its own user agent and reads the robots.txt group that names GPTBot.
- What it feeds
- OpenAI generative AI foundation models
- What OpenAI says about robots.txt
- OpenAI states that disallowing GPTBot indicates a site's content should not be used in training generative AI foundation models.
- User agent string
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot- Verified
- against OpenAI's official docs (opens in a new tab).
robots.txt lines
Allow or block GPTBot.
Paste one group into the robots.txt at the root of your domain. A group that names GPTBot replaces your * group for GPTBot, so copy any * Disallow lines you still want applied into it.
User-agent: GPTBot
Allow: /
User-agent: GPTBot
Disallow: /
# AI crawler access, generated by Suede AI Agent Studio
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Perplexity-User
Allow: /
User-agent: Amzn-SearchBot
Allow: /
User-agent: Amzn-User
Allow: /
User-agent: GPTBot
Disallow: /
Opens the crawlers behind live answers in ChatGPT, Claude and Perplexity while GPTBot stays out.
Precedence
How GPTBot reads your robots.txt
Major crawlers apply the robots.txt standard (RFC 9309) the same way. Three rules decide every path.
- A named group beats *. When a group says
User-agent: GPTBot, GPTBot follows that group and ignores the * group entirely. With no named group, the * group applies; with neither, every path is allowed. - The longest matching rule wins. Inside the group, the Allow or Disallow line with the longest matching path decides, so
Allow: /checkout/helpoverridesDisallow: /checkoutfor the help pages.*matches any run of characters and$marks the end of the path. - Allow wins a tie. When an Allow and a Disallow match with the same length, the path is allowed.
User-agent: *
Disallow: /
User-agent: GPTBot
Disallow: /checkout
Allow: /checkout/help
Disallow: /*.pdf$
| Path | Result | Deciding rule |
|---|---|---|
/ | Allowed | No rule in the GPTBot group matches, so it is allowed |
/checkout/cart | Blocked | line 5: Disallow: /checkout |
/checkout/help | Allowed | line 6: Allow: /checkout/help |
/menu.pdf | Blocked | line 7: Disallow: /*.pdf$ |
Line 2 still blocks every crawler without a group of its own. GPTBot reads only lines 4 to 7, which is why the homepage stays open to it.
Your site
See which line decides GPTBot on your site.
FAQ
GPTBot, answered.
What is GPTBot?
GPTBot is the crawler OpenAI uses to make its generative AI foundation models more useful and safe. GPTBot is run by OpenAI and collects content for AI model training.
How do I allow GPTBot in robots.txt?
Add a group that names it: a line reading "User-agent: GPTBot" followed by "Allow: /". A group that names GPTBot replaces your * group for GPTBot, so repeat any Disallow lines from your * group that should still apply to it. To block it instead, use "Disallow: /" in that group.
Does blocking GPTBot remove me from AI answers?
Blocking GPTBot governs whether your content is used for OpenAI generative AI foundation models. Live answers come from answer bots such as OAI-SearchBot and ChatGPT-User, which follow their own robots.txt groups, so you can block GPTBot and keep those open. OpenAI states that disallowing GPTBot indicates a site's content should not be used in training generative AI foundation models.
Related crawlers
Crawlers to set alongside it.
Every crawler, its operator and its docs are in the full AI crawler list.
Last updated .
Keep GPTBot where you want it. Every Monday.
Run the free check for the full crawler table and ranked fixes, unlock the $9.99 fix kit for paste-ready files, then let the AI Visibility Watch template read your robots.txt weekly and flag any answer bot that becomes blocked.