The OpenAI Decisions API needs a confidence you can trust

OpenAI's new Decisions API, powered by a version of GPT-6 Luna, aims to automate decision-making processes but faces challenges with confidence calibration. Analysis shows that the model's self-reported confidence levels do not consistently align with its actual accuracy, necessitating external verification systems.
Why it matters
As businesses increasingly rely on AI for automated decision-making, the reliability of confidence scores is critical for safety and operational integrity.
Say you send an agent's next step to a decision model, and you only let it act on its own when it's at least 99% sure. Everything else goes to a person. That rule is only as good as the 99%. We gave GPT-6 Luna, the model behind OpenAI's new Decisions API, 3,600 reasoning problems and asked for exactly that kind of answer. When Luna said it was 99% sure or more, it was right 68% of the time.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in