OpenAI: We monitor internal coding agents for misalignment

OpenAI has implemented a monitoring system for its internal coding agents to detect and mitigate risks of misalignment. This initiative is part of the company's broader safety strategy as it deploys increasingly autonomous AI agents in complex environments.
Why it matters
As AI agents gain the ability to modify their own safeguards, internal monitoring becomes a critical component of responsible AGI development.
Using our most powerful models to detect and study misaligned behavior in real-world deployments.
Share Our approach & how it works Our approach & how it works What we monitor for Limitations Towards a safety case with monitoring The road ahead Our approach & how it works What we monitor for Limitations Towards a safety case with monitoring The road ahead AI systems are beginning to act with greater autonomy in real-world environments at scale. As their capabilities advance, they are able to take on increasingly complex, high-impact tasks and interact with tools, systems, and workflows in ways that resemble human collaborators.
A core part of OpenAI’s mission is helping the world navigate this transition to AGI responsibly. That means not only building highly capable systems, but also developing the methods, infrastructure, and approaches needed to deploy and manage them safely as their capabilities continue to grow.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in