Article may be outdated

This article is 49 days old. Some details may have changed since publication.

Hacker News·4 min read·medium

Sites that block AI training crawlers mostly ignore the answer time bots

Z
zeppelin_7
Sites that block AI training crawlers mostly ignore the answer time bots
AI Summary

A study of top websites shows that while many implemented blocks against AI training crawlers like GPTBot in 2023, they have largely failed to address real-time 'answer' bots. New payment models are emerging to allow site owners to charge AI companies for live data access.

Why it matters

The shift from blocking AI crawlers to monetizing real-time data access marks a significant evolution in the economic relationship between web publishers and AI developers.

Dive DeeperCreate a free account to unlock

We read the robots.txt of the top 10,000 sites. 38% of the GPTBot rules we could date were written in a single quarter of 2023 — right after GPTBot launched and the Times sued. Then every AI company quietly added a second bot, the one that answers questions in real time, and almost nobody wrote a rule for it.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologybusiness
Political Bias
Center
LeftLean LCenterLean RRight
Confidence: 85%

The article presents data-driven analysis of web traffic and robots.txt files without taking a side in the AI copyright debate.

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in