Although it’s unclear why, European publishers are being targeted more extensively by AI scraping bots, according to a recent report from TollBit.
And the situation is getting worse. AI-powered scrapes were up 20% in June 2026 (compared to January of the same year), with local news and sports sites experiencing the greatest rise.
How Do AI Bot Scrapers Work?
AI bot scrapers typically use headless browsers (which operate automatically) and residential proxy networks to simulate human behavior and mask bot origins with real IPs to evade detection.
Bots steal cookies to bypass paywalls, effectively hijacking sessions, and extract article text even when it’s hidden behind scripts. What about robots.text directives? Many AI bots simply ignore them.
The Stats on the Ground
TollBit’s report identified AI scraping bots from 40 vendors operating across nearly 4,000 publishers, with median AI scrapes per site four times higher on European sites compared to North American sites.
Alarmingly, for every 179 AI bot visits received by European publishers, there was just one human referral visit, a rate which is almost three times worse than for North American sites. Plus, just 0.05% of external referrals to European sites in the first half of 2026 came from an AI app, compared to 0.16% in North America.
To make matters worse, European sites’ instructions not to scrape (via robots.txt) were ignored almost three times more often than North American sites. So why the disparity?
Why are European Sites Being Hit Harder by AI Bot Scrapers?
Olivia Joslin, TollBit’s cofounder, suggests that European sites are being disproportionately targeted simply because there are more languages used on the continent compared to North America.
With LLMs trying to learn many different languages and pulling content from a variety of sources, it makes sense that European publishers are getting hit harder by bot scrapes.
Researcher-in-residence at INMA Grzegorz Piechota agrees this could be the reason for the disparity, although he also points to usage data. In the US, Claude users make up around 21% of total Claude users worldwide, with English-speaking countries making up under a third of the global total viewers.
According to Piechota, this means most Claude usage, and therefore most scraping activity, is probably connected to non-English content, which explains the heavier hit European publishers are experiencing from AI bots.
There are other possible reasons that European publishers are bearing the brunt of AI bot scraping. These include GDPR constraints, which may limit behavioral tracking, thus making bot detection harder, and a pretty fragmented regulatory landscape which doesn’t yet provide full protection regarding scraped content.
On top of this, the presence of high-value European journalism markets is producing the sort of premium content perfect for model training.
Demand for Non-English Language Information
For Piechota, growing consumer demand for non-English language information is actively influencing the strategies of both AI companies and their data suppliers.
This is evidenced by the fact that Microsoft, OpenAI, and Google have all recently announced initiatives and research to boost AI’s relevance to multi-language audiences – especially those in Europe.
The result? Suppliers of data for AI are collecting more non-English content, with European country domains making up just 4.75% of the pages captured by data supplier Common Crawl in 2009. Fast forward to 2026, and they’re making up nearly 29.98%, according to Piechota’s studies.
The mix of publishers in a company’s customer base, and therefore its dataset, could also account for TollBit’s findings. For example, TollBit’s analysis of nearly 4,000 publishers included 456 European publishers, which included the companies Styria, the Telegraph, and Ringier.
In contrast, data from Cloudflare does not indicate that European publishers are getting hit harder by bots than their North American counterparts.
To summarize Cloudflare’s report, European sites may be attracting a higher level of bot traffic, but North American sites are experiencing more bot activity as a whole.
DataDome’s Position
The popular cybersecurity and bot-blocking solution DataDome also doesn’t recognize a consistent gap between North American and European publishers as suggested by TollBit’s analysis.
For Jérôme Segura, DataDome’s VP of threat research, there’s a huge variance in AI bot scraping among individual publishers, probably based on the size and prominence of each publisher and the amount of scraping coming from bots that self-identify, making them easier to track.
Segura points to the fact that, according to DataDome’s own research, some individual publishers, regardless of region, evidenced AI traffic as high as around one AI visit for every ten human visits. This shows how large a share of overall site traffic AI bots can represent when regional averages are further interrogated.
Bot Protection Strategies
The threat is real, with AI bot scrapers causing significant damage to publishers of all sizes and in all locations. It’s essential that companies take steps to protect themselves, to safeguard their revenue, brand, and reputation. Here’s how to bolster your defenses:
- Use a comprehensive bot-protection tool such as DataDome, which analyzes things like mouse movements, fingerprint consistency, anomaly patterns, and scroll depth to find and block bots before they can cause mischief.
- Deploy rate limiting to block high-frequency requests that suggest bot activity.
- Use dynamic paywalls – by rotating your paywall logic every 24 hours, you’ll break scraper scripts.
- Consider API licensing to offer structured content feeds with legal protections, and block everything else.
- Leverage the legalities and regulations – publishers can demand, if necessary, opt-out mechanisms, licensing agreements, and disclosure of scraped content.
Real-World Case Studies
The ramifications of AI scraping can be massive. Recently, several UK local news outlets were overwhelmed with 200 – 400% traffic spikes during breaking news events.
These spikes were predominantly made up of AI bots attempting to scrape council announcements, live updates, weather alerts, and local crime reports. The scraping caused significant site outages, with scrapers targeting this content for ‘hyperlocal’ model training.
Meanwhile, some Spanish digital publishers frequently find their reporting is being replicated by AI systems mere minutes after publication and, to add insult to injury, outranking the original publication on Google.
This is a direct commercial threat, and represents the full bot exploitation cycle, from scraping, to training, to generating, to outranking, to, finally, monetizing content.
Recent years have also seen a number of ‘shadow crawler’ incidents in France, with multiple French publishers reporting these crawlers hitting their sites at up to fifty times the normal rates.
These bots deploy advanced tactics, rotating residential proxies, and mimicking Chrome user agents to scrape full article bodies, rather than just summaries.
Getting Ahead of the AI Bot Scraping Game
AI scraper bots aren’t a theoretical risk, but a real-world threat, already impacting the European (and worldwide) publishing ecosystem. And one thing has become especially clear: robots.txt is no longer a safeguard against the agentic menace.
For publishers, ignoring the commercial threat is no longer an option to avoid content scraping on a potentially industrial scale. The answer is deploying a robust, multi-pronged bot mitigation strategy that includes a reliable, automated solution able to adapt to even newly emerging bot tactics and techniques in real time.
The publishers that survive and thrive in today’s world will likely be those that ensure their content remains valuable, rather than freely available to any bot to scrape at will.
