Short read · SEO
Which AI bots should a site let in?
Each big AI company sends several bots, and at most one of them collects pages for training. The rule I use, a test of 16 bots on my own sites, and a client's site that kept ChatGPT search out for at least six weeks.
In September we went through a month of server logs for happymonday.ua, a Ukrainian career platform. OpenAI’s search crawler, OAI-SearchBot, had asked for pages again and again. Every one of its requests was refused. We checked the addresses against the list OpenAI publishes: this was the real bot, not someone using its name.
robots.txt allowed it when we checked. A rule on the server turned it away by name. At the same time, pages opened fine when a person in ChatGPT asked about them. That is a different bot, ChatGPT-User. So people could reach the site from ChatGPT, but the crawler that feeds ChatGPT search couldn’t read a single page. Apple’s search crawler, Applebot, was turned away the same way.
The logs covered August 9 to September 8, and live checks showed the same refusal until the fix: at least six weeks. The client’s developer removed the rule on September 23. On September 26 we checked again: both bots get the page. Once the logs showed it, the team treated it as a bug, not a choice. That’s the point of this note: a site can end up with a choice about AI bots that nobody made. Mine had one too.
I can’t show yet what the fix brought. The comparison of visits from ChatGPT is planned for a month after it.
More people seem to be meeting these names. In the US, Ahrefs shows no searches for “what is gptbot?” before January 2026. Since June it gets about 500 a month.
US searches per month: “gptbot” and “what is gptbot?”
“gptbot” stayed between 282 and 690 searches a month for two years. “what is gptbot?” had none before January 2026 and has had 480–566 a month since June.
Three jobs, not three companies
OpenAI and Anthropic send three bots each, Perplexity two. OpenAI also has a fourth, OAI-AdsBot, but it only visits pages submitted as ads in ChatGPT. A rule that blocks “OpenAI” throws out the search bot together with the training one. So I sort bots by what they do with the page.
1 · Search
Builds an index for answers that link back to you. OAI-SearchBot, Claude-SearchBot, PerplexityBot, DuckAssistBot, and the classic search crawlers: Googlebot, Bingbot, Applebot. Let them in. This is how you get cited.
2 · A person's request
Opens your page because someone in a chat asked about it. ChatGPT-User, Claude-User, Perplexity-User, MistralAI-User. Let them in. There is a reader on the other end.
3 · Training
Collects pages to train future models. GPTBot, ClaudeBot, CCBot, Meta-ExternalAgent, Amazonbot. Plus two names that never visit and only live in robots.txt: Google-Extended and Applebot-Extended. Your call. It doesn't decide whether you show up in search.
The third group is the only real decision. Closing it costs nothing in search, by the companies’ own documentation. Google says Google-Extended “does not impact a site’s inclusion in Google Search”. Apple says pages that block Applebot-Extended still appear in Siri, Spotlight and Safari. AI Overviews are part of Google Search: they use pages Googlebot has indexed, not a separate AI crawler.
One catch: Google-Extended covers more than training. Google also uses it to decide whether the Gemini app may pull your pages into its answers. Closing it keeps you in Google Search, but it can keep you out of the Gemini app.
The 16 bots, one by one
What each company says its bot does, and what blocking it means. Checked against their documentation on October 4, 2026.
| Bot | Whose | Job | If you block it |
|---|---|---|---|
| OAI-SearchBot | OpenAI | Search | You don't appear in ChatGPT search answers. |
| ChatGPT-User | OpenAI | A person's request | ChatGPT can't open your page for a user. OpenAI says robots.txt may not apply to it. |
| GPTBot | OpenAI | Training | Your pages stay out of training. Search is not affected. |
| Claude-SearchBot | Anthropic | Search | You may show up less in Claude's search results. |
| Claude-User | Anthropic | A person's request | Claude can't fetch your page when a user asks for it. |
| ClaudeBot | Anthropic | Training | Your future pages stay out of training. |
| PerplexityBot | Perplexity | Search, not used for training | Your pages may not appear in Perplexity's search results. |
| Perplexity-User | Perplexity | A person's request | Perplexity says it generally ignores robots.txt. |
| Googlebot | Search, including AI Overviews | You leave Google Search. | |
| Bingbot | Microsoft | Search | You leave Bing. |
| Applebot | Apple | Search in Siri, Spotlight and Safari | You leave Apple's search. Training is a separate name, Applebot-Extended. |
| DuckAssistBot | DuckDuckGo | AI answers with sources, not used for training | You're not a source for those answers. Your DuckDuckGo rankings don't change. |
| MistralAI-User | Mistral | A person's request | Mistral's assistant can't open your page for a user. |
| Meta-ExternalAgent | Meta | Training and Meta's products | Your pages stay out of Meta's training. |
| Amazonbot | Amazon | Amazon's products; may train its models | A mixed bot. Amazon now has a separate Amzn-SearchBot for search. |
| CCBot | Common Crawl | An open web archive that many models train on | Future snapshots skip you. Old ones stay. |
Seven search bots, four request bots, five training bots. Only the last five are a real choice.
What my own two sites answer
On October 4 I sent all 16 bots to my two sites, mxkeey.com and plants.place: the home page and one article on each. That’s 64 requests, and I ran the whole set twice. Each request borrows the bot’s name, its user agent, and asks for the page:
$ curl -s -o /dev/null -w "%{http_code} %{size_download}\n" \
-A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot" \
https://plants.place/wiki/monstera
200 123909
| Check | mxkeey.com | plants.place |
|---|---|---|
| robots.txt names any AI bot | No | No |
| 16 bots × 2 pages | 32 of 32 got 200 | 32 of 32 got 200 |
| Same page as a browser gets | Yes, except a request ID Cloudflare writes into every page | Yes, byte for byte |
| Who answered | The site, through Cloudflare | The site, through Cloudflare |
| Cloudflare stopped anyone | No | No |
So I let every one of them in, the training bots too. My robots.txt files name no AI bot. On mxkeey.com a comment in the file says the site is open to search engines and AI assistants on purpose, but training isn’t mentioned on either site. Cloudflare let all 128 requests through. For training, the default made the choice for me.
The default depends on when you joined. Since September 15, 2026, domains newly added to Cloudflare get a different one: on pages that show ads, training bots and bots acting for a user are blocked, and search bots are let in. My domains were added earlier and show no ads, so it doesn’t touch them. If your site is new on Cloudflare and runs ads, look at this setting. It may already be turning away bots that fetch pages for your readers.
What this test can’t show. The requests come from my laptop and only borrow the bots’ names. That’s enough for rules that look only at the name, like the server rule at the client. Cloudflare’s AI controls on the free plan also go by the name, according to its documentation. The test doesn’t check robots.txt at all. A polite bot reads robots.txt and stays away by itself, while the server would still answer 200. And it says nothing about whether the real bots come. I don’t know yet how often they visit my sites.
Three places that can say no
1 · robots.txt
A note the bot reads before it comes. A polite bot obeys it. It refuses nothing by itself.
2 · Cloudflare
Answers before your site does, if your site sits behind it. Can block by name, or by its own check of who the bot is.
3 · Your server
Answers last. A rule here can turn a bot away by name, whatever robots.txt says.
At the client, robots.txt said yes and the server said no. On my sites, all three say yes, because none of them names an AI bot. On a new Cloudflare domain with ads, the middle one can say no on its own. A check of robots.txt alone would have missed the client’s problem: the file was fine.
A bot’s name is a claim
The same month of logs had plenty of requests named GPTBot, PerplexityBot and Google-Extended. Not one of them came from the addresses those companies publish. 100% were fakes.
Google-Extended gives itself away. Google says it has no user agent of its own: Google crawls with its usual bots, and Google-Extended only exists as a line in robots.txt. A request that calls itself Google-Extended is fake by definition.
So don’t judge AI bots by the name column in your logs. Check the address. OpenAI, Anthropic, Perplexity, Google and Apple publish the addresses their bots come from.
How to check your own site
- Open your robots.txt and find every AI bot name. Remember that Google-Extended and Applebot-Extended don’t visit. They only tell Google and Apple how they may use what their search bots collected.
- Request your home page and one article with each bot’s user agent, then once more with a normal browser’s. Compare the codes and the sizes. A 403, or a much smaller page, means something refused the bot.
- If something refused it, find out what. The error page usually gives it away: Cloudflare’s block page looks nothing like your server’s own error page.
- Make the training decision once, write it in robots.txt, and make Cloudflare and the server say the same thing. Careful with Cloudflare: it says that from September 15 its Training block also stops crawlers that do search and training at once, and it names Googlebot, Bingbot and Applebot among them.
- In logs, check addresses, not names.
If you want search and readers in and training out, robots.txt needs one group:
# Training bots: keep out. Search and request bots need no lines:
# they follow the rules for everyone below.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: Amazonbot
Disallow: /
User-agent: *
Disallow: /admin
The Google-Extended line also keeps your pages out of answers in the Gemini app. Drop it if you want to be a source there.
One trap: don’t give the search bots a group of their own “to be safe”. A bot that finds its own name in robots.txt follows only that group and ignores the rules for everyone. Your Disallow lines stop applying to it.
On my sites the training bots are in by default, not by decision. That is the one choice left to make on purpose.