Web owner battles bot invasion: 99% of traffic are automated crawlers
A website owner running a database of philanthropists found that 99% of traffic comes from bots. After battling scrapers for a year, including a massive wave from China that generated 3.6 million requests in a single day, he documented his defense strategies against both general web scrapers and AI crawlers. AI bots from companies like Anthropic and Amazon consumed tens of thousands of pages per day without sending any actual visitors.
Cloudflare recently reported that bots have exceeded human activity online, and AI is accelerating this trend further. If someone runs a website where traffic is predominantly generated by entities—organic or synthetic—that consume resources without providing value, how long can such a site realistically operate? The internet is fundamentally changing, so let's hope it evolves in a direction we actually want.
How can website administrators distinguish between bot and human traffic?
Bots typically lack referrer information, visit only one page with high bounce rates, and don't execute JavaScript. Administrators should monitor raw server logs since JavaScript-based analytics only capture human visitors. Cloudflare provides bot-specific traffic insights and enables edge-level filtering rules.
What are typical page crawl-to-visitor ratios for different crawlers?
Quality search engines like Googlebot maintain ratios around 46 pages crawled per visitor referred. AI training crawlers are much worse: Claude's crawler achieved 35,000:1 and Amazon's crawler sent no visitors at all. Bing sits at 406:1, worse than Google but far better than AI-focused crawlers.
Are geographic blocks an effective solution against bot traffic?
Geographic blocks can help against mass waves from specific countries like China, but they're not foolproof. Scrapers adapt and can route through other regions, such as AWS datacenters in America. Combining geographic blocks with other measures like CAPTCHA, referrer checking, and user-agent filtering is more effective.
Should website owners block AI training crawlers entirely?
Polite AI companies respect HTTP 403 responses, making simple firewall rules effective. Since many AI crawlers consume significant bandwidth while sending zero traffic or attribution, blocking them is reasonable. The key metric is pages crawled per visitor referred—AI crawlers score terribly on this metric.
- Kitesurf: Browser engine for agents running in V8 isolates on Cloudflare Workers — blog.cloudflare.com 74 % match
- Artificial intelligence is consuming the web, the internet's collective memory is disappearing — thewalrus.ca 73 % match
- GitHub hit by major web and API service outage — news.ycombinator.com 72 % match
- PatronView
- Cloudflare5
- Anthropic16
- Claude8
- Amazon3
- Google37
- Bing
- Plausible