Claude’s new model is more ‘honest’ when it messes up

Anthropic is releasing Claude Opus 4.8, a new AI model designed to be more 'honest' by flagging uncertainties and avoiding unsupported claims. The company claims internal evaluations show the model is significantly less likely to hallucinate than its predecessor.
Why it matters
As AI adoption grows, improving model reliability and transparency is essential for building trust in automated decision-making and content generation.
Anthropic is releasing Claude Opus 4.8 on Thursday, and the company is touting the model's "honesty." According to Anthropic , it trains "all [its] models to be honest - for instance, to avoid making claims that they can't support." But it notes that "a general problem with AI models is that they sometimes jump to conclusions, confidently presenting their work as making progress despite thin evidence." The AI lab claims that early testers have found that Opus 4.8 "is more likely to flag uncertainties about its work and less likely to make unsupported claims." In the company's evaluations, Opus 4.8 is "around 4x less likely than its predeces … Read the full story at The Verge.
The report neutrally summarizes the company's claims regarding its new product, attributing the information clearly to Anthropic.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in