Hacker News·4 min read·medium

DoGBench: The first user-facing docs generation benchmark. No model scores >50%

P
prithvi2206
DoGBench: The first user-facing docs generation benchmark. No model scores >50%
✦AI Summary

DoGBench is a new benchmark designed to evaluate how well AI agents can write and maintain user-facing software documentation. Results indicate that current AI models struggle to meet expert standards, with no model scoring above 50% on the benchmark tasks.

Why it matters

As companies increasingly rely on AI to automate technical writing, this benchmark highlights significant gaps in AI reasoning and the ability to produce documentation that is actually useful for end-users.

✦Dive DeeperCreate a free account to unlock

Can agents meet expert standards for user-facing documentation?

DoGBench evaluates agents that write and maintain user-facing software documentation in response to real repository events and reported documentation gaps. Each patch is scored with rubrics validated with project maintainers. The scores measure progress toward that standard, not performance relative to a human expert.

Promptless: This was the same production agent available to all Promptless customers. We made no specific agent improvements based on the DogBench work.

Cloud agent caveat: Cloud-based agents were instructed not to use the internet. We verified that Promptless did not browse the internet to find contaminating information, but we could not verify this for other cloud agents.

Submit your agent’s DoGBench results for review. We review submissions manually and will follow up about verification and evaluation requirements.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyaibusiness
✦

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in