What actually happens

A search engine crawls politely and predictably: it has spent twenty years learning not to hurt the sites it indexes. The current generation of AI crawlers has not, and there are a lot of them — seventeen named agents in this census alone, each fetching on its own schedule, and behind them the on-demand ones that fetch a page because somebody just asked a question about it.

The result is not a spike you can see on a graph and forget. It is a floor that rises: more requests, on paths nobody optimised, at hours nobody staffed, from addresses no cache warms.

  • Pages nobody visits get fetched anyway, so your cache hit rate falls and your database does work no reader asked for.
  • Search and filter URLs — the expensive ones — get crawled combinatorially, because a crawler has no reason not to.
  • A crawler that gets a slow response does not back off the way a person does. It retries.
  • Your bill goes up before your traffic does, because bandwidth and compute are metered and attention is not.

Why blocking everything is the wrong answer

It is the first thing everybody tries, and it costs more than it saves. The agents that train models send you nothing; the ones that answer questions send you readers, with a link. Blocking all of them with one rule is how a site disappears from the place people are starting to ask questions, in exchange for a load problem that was solvable another way.

The census exists partly to make that distinction visible: it measures each agent separately, because what blocking costs you depends entirely on which one you block.

What we actually do

It starts with looking, not with buying anything. Your logs say which agents are hitting you, how often, and which URLs they are burning your server on — and that is a different list for every site.

  • Separate the agents that pay you back from the ones that do not, and write a robots.txt that says so per agent instead of per site.
  • Put the expensive paths behind rules a crawler respects: rate limits per agent, and a crawl-delay where it is honoured.
  • Cache what is being fetched repeatedly, which is usually not what your cache was designed for.
  • Serve the cheap thing where the expensive thing was not needed: a static answer for a page whose HTML never changes per visitor.
  • And measure it afterwards, so the invoice and the response times say whether it worked.

How to tell whether this is you

Two signs, and neither needs an audit to spot. The first is a server that is busier than your analytics say it should be — because analytics count people and agents run no JavaScript. The second is slow responses at hours with no readers in them.

If either sounds familiar, the useful first message is not a request for a quote. It is your logs and a sentence about what you are seeing.

Send us your logs, not a brief

The first reply is a question or two, not a quote: what a site needs depends on what it is running and who is hitting it. A page and an address is enough to start, and if you already know which agents are hurting you, say so.