Google updates Android Bench with new LLMs, but Gemini still lags behind

Google has updated its Android Bench tool to evaluate LLM performance in app development, adding new models and metrics. The results show that Google's Gemini models currently trail behind competitors like Claude and GPT in coding accuracy.
Why it matters
As AI becomes central to software development, benchmarking tools help developers identify the most efficient models for coding tasks.
Home-field disadvantage Google updates Android Bench with new LLMs, but Gemini still lags behind Android Bench is evolving, and developers can help guide that process.
The report objectively analyzes benchmark data and acknowledges the competitive disadvantage of the publisher's own ecosystem.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in