DoGBench: The first user-facing docs generation benchmark. No model scores >50%

DoGBench is a new benchmark designed to evaluate how well AI agents can write and maintain user-facing software documentation. Results indicate that current AI models struggle to meet expert standards, with no model scoring above 50% on the benchmark tasks.
Why it matters
As companies increasingly rely on AI to automate technical writing, this benchmark highlights significant gaps in AI reasoning and the ability to produce documentation that is actually useful for end-users.
Can agents meet expert standards for user-facing documentation?
DoGBench evaluates agents that write and maintain user-facing software documentation in response to real repository events and reported documentation gaps. Each patch is scored with rubrics validated with project maintainers. The scores measure progress toward that standard, not performance relative to a human expert.
Promptless: This was the same production agent available to all Promptless customers. We made no specific agent improvements based on the DogBench work.
Cloud agent caveat: Cloud-based agents were instructed not to use the internet. We verified that Promptless did not browse the internet to find contaminating information, but we could not verify this for other cloud agents.
Submit your agent’s DoGBench results for review. We review submissions manually and will follow up about verification and evaluation requirements.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in