Hacker News·4 min read·hard

Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out

S
screm
Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
AI Summary

Researchers conducted a large-scale experiment using 75 simulated software repositories to analyze how AI coding agents like Claude, Codex, and Cursor select tools and frameworks. The study utilized a 'simulated human' orchestrator to evaluate agent performance across various real-world coding tasks.

Why it matters

Understanding the decision-making patterns of AI coding agents is critical for developers and enterprises looking to integrate automation into their software development lifecycles.

Dive DeeperCreate a free account to unlock

We started by running an analysis over thousands of public GitHub repositories from which we extracted statistics about programming languages & frameworks, third-party services, deployment platform, team sizes, and codebase age. Since Tech startups are more likely to have open-source repositories than large enterprises, and stacks are likely very different we then unbiased our statistics based on publicly available data and reached our ideal panel distribution.

We then staffed various coding agents to create real-world repositories to match these exact requirements. Finally, we generated variants in which we removed parts of the codebases and with them, entire third-party service implementations so we could run proper unbiased experiments.

We landed on 75 repositories, in 10 languages, all using fake company names, fake git histories, fake API keys and real lockfiles checked against package manager registries like npm.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologybusinessai

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in