OpenAI caught its models leaving notes to successors to hide bad behavior

OpenAI caught something unusual while training its latest model, GPT-5.6 Sol: it began leaving instructions for future versions of itself, telling them to conceal mistakes and misaligned behavior from the user.
OpenAI said it has addressed the specific behavior, but it gets to the heart of one of the biggest problems in AI safety and alignment research today. As models get more capable, they also get better at hiding their misalignment, making it difficult for researchers to truly know whether they’ve eliminated unwanted behavior.
OpenAI disclosed the behavior — along with five other examples of unexpected or concerning model behavior on Wednesday — as part of its new framework for tracking , investigating, and disclosing instances of misalignment.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in