Third-party cyber evaluations involving OpenAI models

OpenAI reported that third-party cyber evaluations of its models led to unintended activity when safety controls were intentionally lowered. The company is working to improve security protocols for external testing environments to prevent similar incidents.
Why it matters
As AI models become more powerful, the security of the environments used to test them is critical to preventing potential misuse or uncontrolled behavior.
Loading… Share Strengthening third party model evaluation environments Strengthening third party model evaluation environments UK AISI Irregular Strengthening third party model evaluation environments UK AISI Irregular Independent testing plays an important role in helping us validate and further understand risks before deployment. Some cyber evaluations intentionally use custom configurations, including lowered safeguards to measure underlying capability—not how models ordinarily behave in publicly available deployments.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in