Home / Blog

Reference

AI crawlers and user agents

Thirteen crawlers that read the web on behalf of AI systems, what each one is for, the exact robots.txt token, and what blocking it actually costs you. Kept current and free to cite.

AI crawlers and user agents
The table

Every agent, operator and token

The token in the last column is the exact string to use after User-agent in robots.txt. It is case-insensitive but the spelling matters.

%s
CrawlerOperatorPurposerobots.txt token
The distinction that matters

Training crawlers and retrieval crawlers are not the same decision

Blocking a training crawler keeps your content out of a future model. Blocking a retrieval crawler keeps you out of answers being generated right now. Most accidental blocks hit the second kind.

%s
A working example

A robots.txt that welcomes AI answers

User-agent: *
Allow: /

# AI crawlers, explicitly welcome
User-agent: GPTBot
Allow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: Google-Extended
Allow: /

Sitemap: https://example.com/sitemap.xml

If you would rather stay out of training but remain citable, allow the retrieval agents and disallow the training ones. That is a defensible position and the table above tells you which is which.

How we work

Every number on this site has a source

Search volumes are dated and attributed. Method and limits are published. Where a figure was never measured, the page says so rather than filling the gap with something plausible.

That is an unusual amount of exposure for an agency site. It is also the only defensible position for a business whose product is measurement.

Reuse

Free to cite and quote

Use any of this with attribution. No permission needed, no email required. If you find an error, tell us and we will fix it and say so.

Suggested citation

AI crawlers and user agents reference. Bridging Associates, 2026. https://growwithba.in/blog/ai-crawlers-reference

Updated 3 September 2026. Previous versions are not silently rewritten; corrections are dated.

Sources

Primary sources

Operator documentation. Crawler names and behaviour change, so this page carries a date and is corrected rather than quietly edited.

Free scan See what ChatGPT, Gemini, Perplexity and AI Overviews say about your brand

Twenty buying prompts, four engines, you against three named competitors. Report in one working day. No call needed.

FAQs

AI crawlers: common questions

Does blocking GPTBot remove me from ChatGPT?

No. GPTBot is the training crawler. ChatGPT-User and OAI-SearchBot handle live fetching and search, and they are separate tokens.

Does Google-Extended affect my Search ranking?

No. It controls use in Gemini and for training. Search ranking and AI Overviews eligibility are governed by Googlebot.

Should a brand block any of these?

If your content is the product, possibly. If you want to be recommended, almost certainly not. Many brands are blocking by accident through an inherited robots.txt with a restrictive wildcard.

How often does this list change?

Operators add and rename agents a few times a year. This page is dated and corrected when they do.

Free scan

See what the engines say about you

Twenty buying prompts, four engines, your brand against three named competitors. Report in one working day.

  • 20 buying prompts, your category
  • ChatGPT, Gemini, Perplexity, AI Overviews
  • You against three named competitors
  • Report in one working day, yours to keep
1 · Your brand2 · About you

No call needed. Report in one working day. Nothing is sent until step 2.

WhatsApp