Bots Bots and 429s
How we got here

I remember when Google made it possible to actually find what we wanted. The internet was full of junk, the old tried and trusted web rings failed, and search engines were a little bit of a joke. Then along came Google, with its idea of ranking pages based on how well respected they are.
This was great, but like everything, people worked out how to game the system (I'm writing this blog for a similar reason) and the SEO consultant was born. We now had to make our site optimised for the all-powerful Google algorithm, and if you fell below expectations, well, down the rankings you go.
To help us achieve this and also to make your time on the website enjoyable (insert something about users leaving if the site takes longer than....), we all heavily relied on caching. We did object caching internally to speed up database calls and then edge caching to speed up delivery of content to the user, and then we even started mashing our CSS and JS to squeeze out even more ms of performance.
This works really well: the most popular pages on a site are usually fast and efficient for users. Sadly you might be the one who warms the cache, but most users see a fast page.
Then came the bots
Bots have always been a part of this. Google and other services have bots that crawl pages and index them. This is how things like PageRank work. It was an annoyance, but the pros outweighed the cons. It was worth that little uptick in server time for those juicy rankings.
Then in the past few years, more and more services started scraping the internet. This started off as the now powerhouse AI companies scraping for training data, and then later on the agents themselves, as users started using LLMs to ask questions.
We are now in a situation where more and more traffic is being generated by bots and increasingly less by humans. This is a problem: we are still gearing websites for the days of SEO and human users, so we are still tolerating complex setups and letting cache paper over the cracks. However, those 3s uncached and 0.3s cached pages are an issue.
Whereas before you had users looking at popular pages, cached and fast, we now have bots hitting every single possible page on your site, and as these pages are rarely visited, every page adds 3s of server time. They grab the content and hit another, and another, and before you know it, your server is running out of workers and real users are getting hit with 429 or 503 errors.
Why blocking them is not the answer

The knee-jerk reaction is to block bots, but increasingly, being cited by ChatGPT is more important than being on the first page of Google. So that's not an option: just as blocking Google's spiders was an issue before, so too is this.
This is before we start talking about the bots and agents that people use to scrape content, probe sites for vulnerabilities and generally cause issues. We have tools we can use, such as robots.txt, llms.txt, Cloudflare, rate limiting and more, but these just plaster over the problem.
The Problems
Inner page links
I have recently been looking into this issue for a client. They are a big tech site with a lot of users. This means we have a lot of content, but also a crazy amount of comments (we are talking around 0.5m comments). As this site is popular with commenters, they have a collection of tools for liking, sharing, promoting and other operations. These, however, all have links, links that bots love to grab hold of.
When the site was made, having a nice landing page to log in/sign up was great; it drove users. However, when a bot hits that unique URL, it loads it cold with no caching, all the assets are loaded (more calls are made to the server), and the bot sees a signup/login page, turns off, and you have just paid 3s of server time. Then it hits the next button in that comment and we do it again.
You have 500k comments, 5 actions on each, and that's 2.5m URLs that a bot can hit, each one taking 3s of server time. This is a problem: we have to make sure that the bots are not hitting these pages and that they are not being indexed by search engines. We can do this by using robots.txt and llms.txt to block these pages from being crawled, but we also need to make sure that the links are not being shared on social media or other platforms.

Cache warming
As I alluded to in the section above, caching is great, but when mixed with bots trapped in a maze of inner links, it doesn't help. A few caching services out there are not as simple as they seem. Pressable, for example, will cache a page, but not the first time you access it. If someone accesses a page for the first time, a timer is started; then if anyone else accesses it again within 5 mins, the cached copy of the page is saved and served. So you can set this to many days for older, less popular pages. But bots might only hit this page 3 times in 2 days, so it's not helping.
We had an issue on another site, a collection of poems, some 35,000 posts. Unlike the tech site mentioned above, we had no issue with comments and honestly not that many links on the pages. A very clean and minimalist site, but still the usual bot issues. The site is pretty popular; bots seem to be both scraping and actually just accessing posts frequently (we suspect it is referenced in some training data).
We have content on there going back many, many years, and when we reviewed the actual traffic, humans mostly read new content and bots the older. So we tried to set some caching rules: if a post is older than 6 months, we cache it for 7 days. The likelihood of the nav/footer changing was low, so it felt like an easy win.
However, the staggering of this meant the cache didn't really hold up due to the 2 visits in 5 minutes rule, so we had to make changes to the caching rules to ensure the first visit was cached. This saw our server overhead drop substantially, and we now comfortably sit under the worker limit.
Do you need that?
One issue that we have seen has been around how much a webpage does. Stick Query Monitor on a site or look in dev tools, and you often see it's the server that takes the most time to load. I know the attitude in the world of WordPress is "there's a plugin for that", and developers are getting much better at optimising. But honestly, do you need all that?
If your site is throwing a 404 as someone tried to access a term that never existed, do you need a fancy, fully featured page with a nav, footer and all your tracking cookies? Or can you render a simple 404 message? If a user is not logged in, do you even need to show login links all over the site? In the world of bots, the less your site can be triggered to do, either on load or on click, the better.

What you can do
We can't stop bots, and honestly we don't want to. If your website sells a service or products, or promotes a cause or events, you need ChatGPT and Gemini to cite you. However, we can do our best to make sure the bots can access the data they need faster, and reduce the number of links they can follow and how much load they put on the server.
If you would like to discuss how we can help you with your site, please get in touch. We can help you with caching, rate limiting and other techniques to make sure your site is fast and efficient for both humans and bots.
Get in touch