You can't solve computer use by ignoring the interface

The article argues that current AI agents struggle with real-world computer tasks because they are trained on simplified, static benchmarks rather than messy, real-world interfaces. It suggests that the industry's focus on larger models is a dead end and that better interface interaction is the true bottleneck.
Why it matters
As businesses increasingly rely on AI agents for automation, understanding why these tools fail in practical, non-simulated environments is critical for future development.
Right now, agentic computer use is one of the biggest levers for real-world AI impact. LLM-based agents are transforming software development, but most intellectual work is gated behind using software. When coding agents are so good, it is natural to ask: can they file my taxes in a government portal, fix a text document, test a website?
We are not there yet. On OSWorld-V2, a leading benchmark of long-horizon computer tasks, the best model achieves only 20.6% completion rate. On Agents' Last Exam the best result is 26.2% . Users accustomed to the impressive performance of LLMs in chat interactions expect similar results from computer use. But they are met with frustration: agents are unreliable, slow and expensive. At the moment, it's easier to just do the work yourself.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in