Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans

Microsoft has introduced a new AI code of conduct that establishes strict safety constraints, including prohibitions against cyberattacks, deepfakes, and attempts to evade human oversight. The document outlines the company's long-term strategy for aligning superintelligent systems with human values.
Why it matters
As AI capabilities advance, establishing formal safety protocols is critical for preventing catastrophic risks and ensuring that powerful models remain under human control.
As the AI world shifts its focus to safety and alignment, Microsoft has released a new AI code of conduct meant to guide AI models away from dangerous behavior.
The document is more low-level than Anthropic CEO Dario Amodei’s recent call for pacing the frontier , instead focusing on the values and red lines that guide model training within Microsoft AI. Still, the result is a comprehensive guide as to how Microsoft approaches AI safety, and how those ideas are implemented in practice.
The document begins with the prediction that, in the next decade, superintelligent AI systems will surpass human performance in most tasks. “Containing, controlling, and aligning such a powerful force is one of the greatest challenges humanity has ever faced,” the code of conduct continues. “We must therefore be completely clear about why we are inventing these systems and how we intend to control them.”
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in