GPT-6 Astra in code review: Gains, privacy, and cost

Early evaluations of OpenAI's GPT-6 Astra show improved performance in identifying bugs during code reviews compared to previous models. The model demonstrates particular strength in complex, cross-file analysis, though researchers note these are early, directional results.
Why it matters
Advancements in AI-driven code review tools could significantly impact software development productivity and the reliability of large-scale codebases.
Some of the hardest work in code review happens outside the changed lines. A change can look correct in isolation and still break code elsewhere in the system.
That is what makes our early results for OpenAI's GPT-6 Astra most interesting. In our evaluation, Astra caught approximately 4% more labeled bugs through actionable findings than GPT-5.6 Sol , and 22% more than Opus 5 .
The biggest jump comes on harder cross-file reviews, where Astra's gains reach 20% over Sol and 33% over Opus 5 . Using that capability at customer scale also means protecting customer data and assessing the model’s public API pricing.
Our measure here is actionable bug coverage, meaning how many labeled bugs a model catches through findings a developer can act on.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in