Google · Training crawler guide

Google-Extended: who runs it and how to allow or block it

Google-Extended is a standalone robots.txt product token that controls whether content Google crawls may be used to train future Gemini models and for grounding in Gemini Apps and the Vertex AI API.

Google-Extended at a glance

Operator
Google
Purpose
Training It collects content for AI model training.
Kind
Robots.txt control token. Honored by Google's main crawler; no separate visits.
What it feeds
Gemini model training and grounding
What Google says about robots.txt
Google states that Google-Extended has no separate HTTP user agent, crawling is done with existing Google user agents, and the token does not affect a site's inclusion in Google Search.
User agent string
Google-Extended has no user agent of its own. Name it in robots.txt and Google's main crawler applies the rule.
Verified
against Google's official docs (opens in a new tab).

robots.txt lines

Allow or block Google-Extended.

Paste one group into the robots.txt at the root of your domain. A group that names Google-Extended replaces your * group for Google-Extended, so copy any * Disallow lines you still want applied into it.

Allow Google-Extended
User-agent: Google-Extended
Allow: /

Block Google-Extended
User-agent: Google-Extended
Disallow: /

Allow answer bots, block Google-Extended
# AI crawler access, generated by Suede AI Agent Studio

User-agent: OAI-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: Claude-User
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Perplexity-User
Allow: /

User-agent: Amzn-SearchBot
Allow: /

User-agent: Amzn-User
Allow: /

User-agent: Google-Extended
Disallow: /

Opens the crawlers behind live answers in ChatGPT, Claude and Perplexity while Google-Extended stays out.

Precedence

How Google-Extended reads your robots.txt

Major crawlers apply the robots.txt standard (RFC 9309) the same way. Three rules decide every path.

  1. A named group beats *. When a group says User-agent: Google-Extended, Google's crawler follows that group and ignores the * group entirely. With no named group, the * group applies; with neither, every path is allowed.
  2. The longest matching rule wins. Inside the group, the Allow or Disallow line with the longest matching path decides, so Allow: /checkout/help overrides Disallow: /checkout for the help pages. * matches any run of characters and $ marks the end of the path.
  3. Allow wins a tie. When an Allow and a Disallow match with the same length, the path is allowed.
Example robots.txt
User-agent: *
Disallow: /

User-agent: Google-Extended
Disallow: /checkout
Allow: /checkout/help
Disallow: /*.pdf$

What Google-Extended does with this file
PathResultDeciding rule
/AllowedNo rule in the Google-Extended group matches, so it is allowed
/checkout/cartBlockedline 5: Disallow: /checkout
/checkout/helpAllowedline 6: Allow: /checkout/help
/menu.pdfBlockedline 7: Disallow: /*.pdf$

Line 2 still blocks every crawler without a group of its own. Google-Extended reads only lines 4 to 7, which is why the homepage stays open to it. Google's crawler reads the Google-Extended group to decide how content may be used; its regular crawling follows its own group.

Your site

See which line decides Google-Extended on your site.

Opens the free AI visibility check with your URL filled in. It reads your robots.txt and shows the exact line that decides Google-Extended and every other AI crawler, with ranked fixes and a site-specific fix kit to unlock for $9.99.

FAQ

Google-Extended, answered.

What is Google-Extended?

Google-Extended is a standalone robots.txt product token that controls whether content Google crawls may be used to train future Gemini models and for grounding in Gemini Apps and the Vertex AI API. Google-Extended is run by Google and collects content for AI model training. It is a robots.txt control token: Google's main crawler honors it and it makes no separate visits.

How do I allow Google-Extended in robots.txt?

Add a group that names it: a line reading "User-agent: Google-Extended" followed by "Allow: /". A group that names Google-Extended replaces your * group for Google-Extended, so repeat any Disallow lines from your * group that should still apply to it. To block it instead, use "Disallow: /" in that group.

Does blocking Google-Extended remove me from AI answers?

Blocking Google-Extended governs whether your content is used for Gemini model training and grounding. Live answers come from answer bots such as OAI-SearchBot, Claude-SearchBot and PerplexityBot, which follow their own robots.txt groups, so you can block Google-Extended and keep those open. Google states that Google-Extended has no separate HTTP user agent, crawling is done with existing Google user agents, and the token does not affect a site's inclusion in Google Search.

Related crawlers

Every crawler, its operator and its docs are in the full AI crawler list.

Last updated .

Keep Google-Extended where you want it. Every Monday.

Run the free check for the full crawler table and ranked fixes, unlock the $9.99 fix kit for paste-ready files, then let the AI Visibility Watch template read your robots.txt weekly and flag any answer bot that becomes blocked.

Inside Agent Studio

Good work starts with a clear flow.

Give an agent a job. Connect the steps. See how the work becomes a service someone else can call.

Input
Reason
Output

A workflow you can inspect

Make the steps visible.

Start with a template or describe the job. Connect inputs, reasoning, and output on the canvas.

Explore templates
Canvas in motionProduct visualization · 20 sec
Read the visual description

A stylized organization chart opens into a workflow canvas. Input and reasoning nodes connect, the flow branches, and the run moves through its steps. The film ends with Suede AI Agent Studio. Music only; no spoken narration. Illustrative interface, not a live run.