Google · Training crawler guide
Google-Extended: who runs it and how to allow or block it
Google-Extended is a standalone robots.txt product token that controls whether content Google crawls may be used to train future Gemini models and for grounding in Gemini Apps and the Vertex AI API.
Google-Extended at a glance
- Operator
- Purpose
- Training It collects content for AI model training.
- Kind
- Robots.txt control token. Honored by Google's main crawler; no separate visits.
- What it feeds
- Gemini model training and grounding
- What Google says about robots.txt
- Google states that Google-Extended has no separate HTTP user agent, crawling is done with existing Google user agents, and the token does not affect a site's inclusion in Google Search.
- User agent string
- Google-Extended has no user agent of its own. Name it in robots.txt and Google's main crawler applies the rule.
- Verified
- against Google's official docs (opens in a new tab).
robots.txt lines
Allow or block Google-Extended.
Paste one group into the robots.txt at the root of your domain. A group that names Google-Extended replaces your * group for Google-Extended, so copy any * Disallow lines you still want applied into it.
User-agent: Google-Extended
Allow: /
User-agent: Google-Extended
Disallow: /
# AI crawler access, generated by Suede AI Agent Studio
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Perplexity-User
Allow: /
User-agent: Amzn-SearchBot
Allow: /
User-agent: Amzn-User
Allow: /
User-agent: Google-Extended
Disallow: /
Opens the crawlers behind live answers in ChatGPT, Claude and Perplexity while Google-Extended stays out.
Precedence
How Google-Extended reads your robots.txt
Major crawlers apply the robots.txt standard (RFC 9309) the same way. Three rules decide every path.
- A named group beats *. When a group says
User-agent: Google-Extended, Google's crawler follows that group and ignores the * group entirely. With no named group, the * group applies; with neither, every path is allowed. - The longest matching rule wins. Inside the group, the Allow or Disallow line with the longest matching path decides, so
Allow: /checkout/helpoverridesDisallow: /checkoutfor the help pages.*matches any run of characters and$marks the end of the path. - Allow wins a tie. When an Allow and a Disallow match with the same length, the path is allowed.
User-agent: *
Disallow: /
User-agent: Google-Extended
Disallow: /checkout
Allow: /checkout/help
Disallow: /*.pdf$
| Path | Result | Deciding rule |
|---|---|---|
/ | Allowed | No rule in the Google-Extended group matches, so it is allowed |
/checkout/cart | Blocked | line 5: Disallow: /checkout |
/checkout/help | Allowed | line 6: Allow: /checkout/help |
/menu.pdf | Blocked | line 7: Disallow: /*.pdf$ |
Line 2 still blocks every crawler without a group of its own. Google-Extended reads only lines 4 to 7, which is why the homepage stays open to it. Google's crawler reads the Google-Extended group to decide how content may be used; its regular crawling follows its own group.
Your site
See which line decides Google-Extended on your site.
FAQ
Google-Extended, answered.
What is Google-Extended?
Google-Extended is a standalone robots.txt product token that controls whether content Google crawls may be used to train future Gemini models and for grounding in Gemini Apps and the Vertex AI API. Google-Extended is run by Google and collects content for AI model training. It is a robots.txt control token: Google's main crawler honors it and it makes no separate visits.
How do I allow Google-Extended in robots.txt?
Add a group that names it: a line reading "User-agent: Google-Extended" followed by "Allow: /". A group that names Google-Extended replaces your * group for Google-Extended, so repeat any Disallow lines from your * group that should still apply to it. To block it instead, use "Disallow: /" in that group.
Does blocking Google-Extended remove me from AI answers?
Blocking Google-Extended governs whether your content is used for Gemini model training and grounding. Live answers come from answer bots such as OAI-SearchBot, Claude-SearchBot and PerplexityBot, which follow their own robots.txt groups, so you can block Google-Extended and keep those open. Google states that Google-Extended has no separate HTTP user agent, crawling is done with existing Google user agents, and the token does not affect a site's inclusion in Google Search.
Related crawlers
Crawlers to set alongside it.
Every crawler, its operator and its docs are in the full AI crawler list.
Last updated .
Keep Google-Extended where you want it. Every Monday.
Run the free check for the full crawler table and ranked fixes, unlock the $9.99 fix kit for paste-ready files, then let the AI Visibility Watch template read your robots.txt weekly and flag any answer bot that becomes blocked.