An Update on the scraper situation
Websites are facing an increasing volume of automated traffic from 'residential proxies' used to scrape data for AI training models. This practice, which often hijacks ordinary devices, makes it difficult for site administrators to distinguish between human users and malicious bots.
Why it matters
The proliferation of AI-driven scraping threatens the sustainability of the open web and forces site owners to implement increasingly complex defensive measures.
Welcome to LWN.net The following subscription-only content has been made available to you by an LWN subscriber. Thousands of subscribers depend on LWN for the best news from the Linux and free software communities. If you enjoy this article, please consider subscribing to LWN . Thank you for visiting LWN.net! By Jonathan Corbet July 10, 2026 Our article " Fighting the AI scraper bot scourge ", published in early 2025, discussed the problem of widespread scraping of web sites in search of training data for large language models and related projects. This activity overwhelms sites with traffic. Over a year after that article is published, the problem is still growing. The hammering of sites by shadowy actors has reached new heights, and the open web is becoming increasingly difficult to maintain. Where is this traffic coming from, and what can be done about it?
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in