Article may be outdated

This article is 88 days old. Some details may have changed since publication.

Hacker News·6 min read·hard

An Update on the scraper situation

C
chmaynard
✦AI Summary

Websites are facing an increasing volume of automated traffic from 'residential proxies' used to scrape data for AI training models. This practice, which often hijacks ordinary devices, makes it difficult for site administrators to distinguish between human users and malicious bots.

Why it matters

The proliferation of AI-driven scraping threatens the sustainability of the open web and forces site owners to implement increasingly complex defensive measures.

✦Dive DeeperCreate a free account to unlock

Welcome to LWN.net The following subscription-only content has been made available to you by an LWN subscriber. Thousands of subscribers depend on LWN for the best news from the Linux and free software communities. If you enjoy this article, please consider subscribing to LWN . Thank you for visiting LWN.net! By Jonathan Corbet July 10, 2026 Our article " Fighting the AI scraper bot scourge ", published in early 2025, discussed the problem of widespread scraping of web sites in search of training data for large language models and related projects. This activity overwhelms sites with traffic. Over a year after that article is published, the problem is still growing. The hammering of sites by shadowy actors has reached new heights, and the open web is becoming increasingly difficult to maintain. Where is this traffic coming from, and what can be done about it?

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyai
✦

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in