Search engine crawlers such as Googlebot and Bingbot visit your pages to index them. When a page runs a search as it loads (a search results page, or a category page that uses InstantSearch), every crawler visit triggers search requests.
This is often a large source of unexpected usage. A common cause is filters and sorting: if every refinement and sort option produces its own crawlable URL, a crawler can visit a huge number of combinations. For example, a page with dozens of filters can produce more crawlable URLs than a crawler will ever finish, and each one triggers a search.
Crawlers like these identify themselves and follow your instructions, so the fix is to control what they crawl rather than to block them. Blocking them outright can hurt your visibility in search engines.
Control what crawlers visit
-
Use
robots.txtto tell crawlers which paths they shouldn't crawl, such as search results pages or URLs with refinement and sort parameters. - Use sitemaps to point crawlers to the pages you do want indexed, such as product pages, so they don't have to discover them through search.
- Use meta tags to indicate which pages and links crawlers should or shouldn't follow.
Done properly, crawlers reach your product pages directly instead of exploring your site through searches. See Google's introduction to robots.txt and blocking indexing for more detail.
Example: excluding refinement and sort URLs
If refining and sorting your results produces URLs like this:
https://example.com/?color=blue&manufacturer=acme&price=150-234&product_list_order=new
you can exclude those parameters in robots.txt:
User-agent: * Disallow: /*?*color Disallow: /*?*manufacturer Disallow: /*?*price Disallow: /*?*product_list_order
With InstantSearch's default URL format, you can do the same with:
Disallow: /*?*refinementList Disallow: /*?*sortBy
Adjust these to match the parameters your site actually uses.
Check that it really is the crawler
User agents are easy to fake. Bots that scrape sites often claim to be Googlebot or Bingbot. Before treating traffic as a genuine crawler, check the IP addresses against the crawler operator's published verification method. If the traffic isn't genuine, robots.txt won't stop it. See How can I mitigate bots impacting my usage of Algolia?