Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out

Researchers conducted a large-scale experiment using 75 simulated software repositories to analyze how AI coding agents like Claude, Codex, and Cursor select tools and frameworks. The study utilized a 'simulated human' orchestrator to evaluate agent performance across various real-world coding tasks.
Why it matters
Understanding the decision-making patterns of AI coding agents is critical for developers and enterprises looking to integrate automation into their software development lifecycles.
We started by running an analysis over thousands of public GitHub repositories from which we extracted statistics about programming languages & frameworks, third-party services, deployment platform, team sizes, and codebase age. Since Tech startups are more likely to have open-source repositories than large enterprises, and stacks are likely very different we then unbiased our statistics based on publicly available data and reached our ideal panel distribution.
We then staffed various coding agents to create real-world repositories to match these exact requirements. Finally, we generated variants in which we removed parts of the codebases and with them, entire third-party service implementations so we could run proper unbiased experiments.
We landed on 75 repositories, in 10 languages, all using fake company names, fake git histories, fake API keys and real lockfiles checked against package manager registries like npm.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in