Skip to content
GetProfit

← All terms

Googlebot

Googlebot is the generic name for Google Search's web crawlers, the programs that fetch a site's pages so that Google can consider them for its index.

How it works

Googlebot has two types, Smartphone and Desktop. Most requests come from the smartphone crawler, because for most sites Google mainly indexes the mobile version. Googlebot finds new addresses mostly through links on pages it has already crawled.

Other Google crawlers visit a store too. Storebot-Google crawls product, cart and checkout pages to collect price, shipping and payment details for Google Shopping. AdsBot checks ad landing pages. Google-InspectionTool fetches pages for URL Inspection in Google Search Console.

The user agent, the crawler’s name in each request, is easy to fake, so Google’s own check is the IP. A reverse DNS lookup must return a name on googlebot.com, google.com or googleusercontent.com. A forward lookup of that name must return the same IP. Google’s published IP lists work too.

If bot protection answers the real crawler with a 403, Google treats the page as missing and removes it from the index. A 429 or 5xx slows crawling; indexed pages stay for a while but are eventually dropped. Our guide shows how to verify Googlebot behind bot protection on a CDN.

Example

Example store, not client data.

The tableware shop’s CDN log shows two requests with the Googlebot user agent. For 66.249.66.1, the reverse lookup gives crawl-66-249-66-1.googlebot.com, which resolves back to the same IP: this is Googlebot. The second IP resolves to a hosting company’s domain: an impostor, safe to block.

Not to be confused with

  • Indexing — Googlebot fetches pages, and indexing decides whether Google keeps them. Blocking the crawler in robots.txt does not keep a URL out of results.

Right and wrong readings

  • Wrong: “We sent a request with Googlebot’s user agent, got a 403, so the CDN blocks Google.” Right: a request from your own IP shows how the protection treats an impostor, not Googlebot.
  • Wrong: “robots.txt allows Googlebot, so Google Shopping’s crawler gets in.” Right: Storebot-Google has its own user-agent name in robots.txt, and rules written for Googlebot do not apply to it.

Sources