Skip to content
Notes

Short read · SEO

Which AI bots should a site let in?

Each big AI company sends several bots, and at most one of them collects pages for training. The rule I use, a test of 16 bots on my own sites, and a client's site that kept ChatGPT search out for at least six weeks.

Max Kiriienko
Tech Lead SEO & Marketing · 8 min read

Читати українською

In September we went through a month of server logs for happymonday.ua, a Ukrainian career platform. OpenAI’s search crawler, OAI-SearchBot, had asked for pages again and again. Every one of its requests was refused. We checked the addresses against the list OpenAI publishes: this was the real bot, not someone using its name.

robots.txt allowed it when we checked. A rule on the server turned it away by name. At the same time, pages opened fine when a person in ChatGPT asked about them. That is a different bot, ChatGPT-User. So people could reach the site from ChatGPT, but the crawler that feeds ChatGPT search couldn’t read a single page. Apple’s search crawler, Applebot, was turned away the same way.

The logs covered August 9 to September 8, and live checks showed the same refusal until the fix: at least six weeks. The client’s developer removed the rule on September 23. On September 26 we checked again: both bots get the page. Once the logs showed it, the team treated it as a bug, not a choice. That’s the point of this note: a site can end up with a choice about AI bots that nobody made. Mine had one too.

I can’t show yet what the fix brought. The comparison of visits from ChatGPT is planned for a month after it.

More people seem to be meeting these names. In the US, Ahrefs shows no searches for “what is gptbot?” before January 2026. Since June it gets about 500 a month.

US searches per month: “gptbot” and “what is gptbot?”

“gptbot” stayed between 282 and 690 searches a month for two years. “what is gptbot?” had none before January 2026 and has had 480–566 a month since June.

Ahrefs estimates. A query that appears out of nowhere may partly be AI tools searching for someone; the export can't tell. Source: Ahrefs export, US, October 2026

Three jobs, not three companies

OpenAI and Anthropic send three bots each, Perplexity two. OpenAI also has a fourth, OAI-AdsBot, but it only visits pages submitted as ads in ChatGPT. A rule that blocks “OpenAI” throws out the search bot together with the training one. So I sort bots by what they do with the page.

1 · Search

Builds an index for answers that link back to you. OAI-SearchBot, Claude-SearchBot, PerplexityBot, DuckAssistBot, and the classic search crawlers: Googlebot, Bingbot, Applebot. Let them in. This is how you get cited.

2 · A person's request

Opens your page because someone in a chat asked about it. ChatGPT-User, Claude-User, Perplexity-User, MistralAI-User. Let them in. There is a reader on the other end.

3 · Training

Collects pages to train future models. GPTBot, ClaudeBot, CCBot, Meta-ExternalAgent, Amazonbot. Plus two names that never visit and only live in robots.txt: Google-Extended and Applebot-Extended. Your call. It doesn't decide whether you show up in search.

OpenAI, Anthropic, Google and Apple each document training separately from search. OpenAI says outright that you can allow OAI-SearchBot and block GPTBot.

The third group is the only real decision. Closing it costs nothing in search, by the companies’ own documentation. Google says Google-Extended “does not impact a site’s inclusion in Google Search”. Apple says pages that block Applebot-Extended still appear in Siri, Spotlight and Safari. AI Overviews are part of Google Search: they use pages Googlebot has indexed, not a separate AI crawler.

One catch: Google-Extended covers more than training. Google also uses it to decide whether the Gemini app may pull your pages into its answers. Closing it keeps you in Google Search, but it can keep you out of the Gemini app.

The 16 bots, one by one

What each company says its bot does, and what blocking it means. Checked against their documentation on October 4, 2026.

BotWhoseJobIf you block it
OAI-SearchBotOpenAISearchYou don't appear in ChatGPT search answers.
ChatGPT-UserOpenAIA person's requestChatGPT can't open your page for a user. OpenAI says robots.txt may not apply to it.
GPTBotOpenAITrainingYour pages stay out of training. Search is not affected.
Claude-SearchBotAnthropicSearchYou may show up less in Claude's search results.
Claude-UserAnthropicA person's requestClaude can't fetch your page when a user asks for it.
ClaudeBotAnthropicTrainingYour future pages stay out of training.
PerplexityBotPerplexitySearch, not used for trainingYour pages may not appear in Perplexity's search results.
Perplexity-UserPerplexityA person's requestPerplexity says it generally ignores robots.txt.
GooglebotGoogleSearch, including AI OverviewsYou leave Google Search.
BingbotMicrosoftSearchYou leave Bing.
ApplebotAppleSearch in Siri, Spotlight and SafariYou leave Apple's search. Training is a separate name, Applebot-Extended.
DuckAssistBotDuckDuckGoAI answers with sources, not used for trainingYou're not a source for those answers. Your DuckDuckGo rankings don't change.
MistralAI-UserMistralA person's requestMistral's assistant can't open your page for a user.
Meta-ExternalAgentMetaTraining and Meta's productsYour pages stay out of Meta's training.
AmazonbotAmazonAmazon's products; may train its modelsA mixed bot. Amazon now has a separate Amzn-SearchBot for search.
CCBotCommon CrawlAn open web archive that many models train onFuture snapshots skip you. Old ones stay.

Seven search bots, four request bots, five training bots. Only the last five are a real choice.

What my own two sites answer

On October 4 I sent all 16 bots to my two sites, mxkeey.com and plants.place: the home page and one article on each. That’s 64 requests, and I ran the whole set twice. Each request borrows the bot’s name, its user agent, and asks for the page:

$ curl -s -o /dev/null -w "%{http_code} %{size_download}\n" \
    -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot" \
    https://plants.place/wiki/monstera
200 123909
Checkmxkeey.complants.place
robots.txt names any AI botNoNo
16 bots × 2 pages32 of 32 got 20032 of 32 got 200
Same page as a browser getsYes, except a request ID Cloudflare writes into every pageYes, byte for byte
Who answeredThe site, through CloudflareThe site, through Cloudflare
Cloudflare stopped anyoneNoNo

So I let every one of them in, the training bots too. My robots.txt files name no AI bot. On mxkeey.com a comment in the file says the site is open to search engines and AI assistants on purpose, but training isn’t mentioned on either site. Cloudflare let all 128 requests through. For training, the default made the choice for me.

The default depends on when you joined. Since September 15, 2026, domains newly added to Cloudflare get a different one: on pages that show ads, training bots and bots acting for a user are blocked, and search bots are let in. My domains were added earlier and show no ads, so it doesn’t touch them. If your site is new on Cloudflare and runs ads, look at this setting. It may already be turning away bots that fetch pages for your readers.

What this test can’t show. The requests come from my laptop and only borrow the bots’ names. That’s enough for rules that look only at the name, like the server rule at the client. Cloudflare’s AI controls on the free plan also go by the name, according to its documentation. The test doesn’t check robots.txt at all. A polite bot reads robots.txt and stays away by itself, while the server would still answer 200. And it says nothing about whether the real bots come. I don’t know yet how often they visit my sites.

Three places that can say no

1 · robots.txt

A note the bot reads before it comes. A polite bot obeys it. It refuses nothing by itself.

2 · Cloudflare

Answers before your site does, if your site sits behind it. Can block by name, or by its own check of who the bot is.

3 · Your server

Answers last. A rule here can turn a bot away by name, whatever robots.txt says.

A bot gets the page only if all three say yes. Each is set up separately, often by different people in different years.

At the client, robots.txt said yes and the server said no. On my sites, all three say yes, because none of them names an AI bot. On a new Cloudflare domain with ads, the middle one can say no on its own. A check of robots.txt alone would have missed the client’s problem: the file was fine.

A bot’s name is a claim

The same month of logs had plenty of requests named GPTBot, PerplexityBot and Google-Extended. Not one of them came from the addresses those companies publish. 100% were fakes.

Google-Extended gives itself away. Google says it has no user agent of its own: Google crawls with its usual bots, and Google-Extended only exists as a line in robots.txt. A request that calls itself Google-Extended is fake by definition.

So don’t judge AI bots by the name column in your logs. Check the address. OpenAI, Anthropic, Perplexity, Google and Apple publish the addresses their bots come from.

How to check your own site

  1. Open your robots.txt and find every AI bot name. Remember that Google-Extended and Applebot-Extended don’t visit. They only tell Google and Apple how they may use what their search bots collected.
  2. Request your home page and one article with each bot’s user agent, then once more with a normal browser’s. Compare the codes and the sizes. A 403, or a much smaller page, means something refused the bot.
  3. If something refused it, find out what. The error page usually gives it away: Cloudflare’s block page looks nothing like your server’s own error page.
  4. Make the training decision once, write it in robots.txt, and make Cloudflare and the server say the same thing. Careful with Cloudflare: it says that from September 15 its Training block also stops crawlers that do search and training at once, and it names Googlebot, Bingbot and Applebot among them.
  5. In logs, check addresses, not names.

If you want search and readers in and training out, robots.txt needs one group:

# Training bots: keep out. Search and request bots need no lines:
# they follow the rules for everyone below.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: Amazonbot
Disallow: /

User-agent: *
Disallow: /admin

The Google-Extended line also keeps your pages out of answers in the Gemini app. Drop it if you want to be a source there.

One trap: don’t give the search bots a group of their own “to be safe”. A bot that finds its own name in robots.txt follows only that group and ignores the rules for everyone. Your Disallow lines stop applying to it.

On my sites the training bots are in by default, not by decision. That is the one choice left to make on purpose.