OpenAI models secretly generate instructions to ignore constraints

Internal unreleased Astra family model · RL training
W e observed rare cases of a model writing jailbreak-like instructions into its own compaction summaries (the summaries used to continue a task in a new context). Our conclusion was that this behavior was extremely rare, did not confer an obvious reward advantage, and was monitorable. Our top hypothesis is that issues around summary termination contributed to this behavior, though we have not established a causal connection, and we have addressed a related bug.
During RL training, an unreleased Astra-family model sometimes added unauthorized instructions to its compaction summaries.
In the following example, the task was to check whether a local public library had certain books:
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in