Skip to content
Notes

Short read · SEO

What Googlebot told me about my domain's past

Two months after launch I exported the crawl stats of my plant encyclopedia. The domain had a past, Google kept asking for an icon file that didn't exist, and almost a third of its requests went to copies of pages.

Max Kiriienko
Tech Lead SEO & Marketing · 6 min read

Читати українською

plants.place, my houseplant encyclopedia, opened on August 1. In late July my list of candidate names marked the domain as free to register, and I took that to mean it was new. Two months later I exported Search Console’s crawl stats and found out it wasn’t.

Crawl stats is easy to miss. It sits in Settings, not in the main menu, and Google doesn’t offer it through the API. But it is the one place in Search Console that shows Googlebot’s requests themselves, day by day — including requests for addresses you never made.

What I exported

Two files from the report’s breakdown by response: requests answered 200 OK and requests answered 404 Not found. To get them, open Settings → Crawl stats → Open report, click a response type, then Export. Each export has a daily chart for August 1 – October 2 and a table of example requests. I recounted every crawl number below from those files. Redirects and other answers are not in them.

The first three days: 483 requests, then about 8 a day

Googlebot got a normal 200 answer 1,120 times in 63 days. 483 of those came in the first three days: 108, 225 and 150. After that the median was 8 a day, and on August 17 there were none. The server’s average response was about 0.7 seconds.

Googlebot requests to plants.place per day, answered 200 OK

1,120 requests in 63 days. 483 of them in the first three days (108, 225, 150), then a median of 8 a day.

Only requests that got a normal 200 answer. 404s, redirects and other answers are not in this line. Source: Search Console crawl stats for plants.place, Aug 1 – Oct 2, 2026

There are two smaller bursts: 65 requests on August 11 and 77 on August 31. On August 9, thirty new species went live and the catalogue grew from 103 to 133, which may explain the first one. I can’t tie either to anything for sure.

Google remembered a WordPress site I never had

47 requests got a 404, and Search Console listed 39 of them as examples. 14 were WordPress addresses: the login page wp-login.php, a category feed and an author’s feed. plants.place has never run WordPress. Googlebot came back for them on eight different days, from August 3 to September 19, always over plain http.

What Googlebot got a 404 for

Of 39 sampled 404s: the icon file /favicon.ico 16, WordPress addresses 14, everything else 9.

39 examples Search Console listed out of 47 requests that got a 404 in two months. Source: Search Console crawl stats for plants.place, Aug 1 – Oct 2, 2026

Free to register doesn’t mean new. Someone must have run a WordPress site on this domain before me, and Google still had its addresses on file. I didn’t go looking for who it was, because it doesn’t change what to do. It changes what to check before a launch. More on that at the end.

The other nine were six requests for an Apple app-links file the site never had, one listing that no longer existed and two service addresses.

The icon Google kept asking for

The biggest group of 404s was the icon. Googlebot asked for /favicon.ico 16 times in the sample and got a 404 every time. The site did have an icon: a small SVG written straight into the HTML of every page, as a data URI. Browsers show that fine. Google’s favicon guide describes something else — a file that Googlebot-Image must be able to crawl. It says nothing about icons embedded in the page, and I wasn’t going to bet on it.

The embarrassing part: I had learned this a week earlier. On September 26, the crawl stats of another small site of mine led to the same fix — its icon also lived only inside the HTML. I fixed it on every site of that project and didn’t think of plants.place. A lesson learned in one project doesn’t travel to the next one by itself.

Almost a third of the crawl went to copies

In the table of 874 example requests answered 200, 269 (31%) went to addresses with parameters:

  • 216 went to the “similar plants” tabs on species pages: ?rel=care, ?rel=family, ?rel=all. Each tab was a link, so each was a new address for Google.
  • 42 went to the form for posting a listing, filled in with a species: /market/new?p=aspidistra.
  • 11 went to other parameters: catalogue picks and login links.

Where 874 sampled Googlebot requests went

605 clean addresses, 216 copies of species pages with ?rel=, 42 copies of the listing form with ?p=, 11 other addresses with parameters.

Search Console’s examples of requests answered 200 OK over the same two months. The “similar plants” copies had a canonical pointing at the clean page, and Google fetched them anyway. Source: Search Console crawl stats for plants.place

Counted by distinct addresses, it looks worse. In the sample, Google fetched 163 different “similar plants” copies and 121 different species pages.

Every “similar plants” copy had a canonical tag pointing at the clean species page, and the form was marked noindex. That probably kept the copies out of the index. It didn’t stop Google from fetching them. In the sample’s first three days, 183 of its 400 requests (46%) went to the “similar plants” copies. Later Google mostly left them alone — 33 more in two months. But in September the form copies grew: 27 of 230 sampled requests, more than the “similar plants” copies got that month.

Is that a real problem for a site this size? By Google’s own crawl budget guide, probably not: it is mainly meant for sites with at least 10,000 pages that change every day, or a million that change weekly. It also names sites where a large share of pages sits in “Discovered – currently not indexed”, and I haven’t measured that share for plants.place. I can’t prove the copies delayed a single real page. I fixed it anyway. The fix was cheap, and the site had just doubled: on October 3, the day of the fix, the English version of all 133 species went live.

The site existed twice

13 of the 874 sampled requests got a full page over plain http, with no redirect — the home page 9 times. Nothing sent http to https: not Cloudflare, not my code. The canonical pointed at https, so again it probably protected the index, but not the crawl.

What I changed on October 3

What Googlebot foundWhat changedAnswer now
/favicon.ico → 404Icon files: favicon.ico at 16, 32 and 48 px, an SVG, a 96 px PNG, an Apple touch icon and a web manifest, linked from every page200
WordPress addresses → 404410 Gone straight away, over http too410
"Similar plants" tabs in ?rel=The tab now lives after # in the address, which Google ignores. Old ?rel= addresses redirect to the clean page301
Form copies in ?p=The same move to #p=. The form with any parameter is closed in robots.txt301
http answered 200One redirect to https, and browsers are told to use https only301

I checked each answer on the live site on October 4. Two parameters from the sample are still there. The catalogue picks, such as “Fine if you forget to water”, were 7 of the 874 requests, and their canonical points at the catalogue. The login links were 4, and the login page is marked noindex.

About the 410: Google says it treats every 4xx answer except 429 the same way, so I don’t expect it to do more than the 404 did. It is simply the honest answer — gone for good — and it cost nothing.

What comes next

This is the “before”. The “after” needs a second export, and crawl stats can’t be pulled through the API, so it’s a manual step. I’ll do it around October 19 and add the results here. I’ll look at four things: whether /favicon.ico now gets a 200, whether the WordPress requests fade, how much of the crawl still goes to parameters, and whether plants.place shows its icon in Google’s results.

One caveat in advance. The same day, I also moved the whole site to new addresses: English at the root, Ukrainian under /ua/. Every old address now answers with a redirect, so the next export will be full of 301s that have nothing to do with these fixes. I’ll separate the two before drawing any conclusion.

Check a domain’s past before you launch

In the plants.place case study I wrote that I’d do this differently. Here is what I now check:

  • Open crawl stats in the first week, not the third month, and read the 404s. Addresses you never made are the domain’s history.
  • Look at Sitemaps in Search Console. On another site of mine, on an expired domain, it still listed a sitemap someone had submitted a year before I registered it.
  • Put an icon file at /favicon.ico from day one, even if the page already has an icon.
  • Keep tabs and filters you don’t want in search out of the query string. If a state has to live in the address, put it after #.