WRITTEN IN PLAIN AMERICAN ENGLISH.
About
CLAY TRIBUNE.
ShopCartAccount
Advertisement

Microsoft writes rules telling its AI models they cannot hack systems or fool people

Microsoft's new AI code of conduct bans hacking, deepfakes, and loss of human control, per its own document.

By mitch·4 min read
A glowing digital terminal screen displaying code, with a shield icon symbolizing ethical boundaries.

Microsoft has published a new AI code of conduct telling its models not to hack computer systems, produce deepfakes, or trick humans into losing control over them. The document lays out the company’s approach to AI safety in practical terms, setting absolute constraints on what its models are allowed to do.

The code of conduct follows a wave of concern about AI misbehavior, including the resignation of an Anthropic employee who cited the growing risk that AI would cause human extinction. It also arrives as Anthropic CEO Dario Amodei has called for pacing the frontier of AI development.

What the Code Outlines

The document opens with a prediction that superintelligent AI systems will surpass human performance in most tasks within the next decade. “Containing, controlling, and aligning such a powerful force is one of the greatest challenges humanity has ever faced,” the code states. “We must therefore be completely clear about why we are inventing these systems and how we intend to control them.”

Advertisement

Microsoft’s approach centers on general principles and specific safety constraints. The principles include supporting humans rather than replacing them and accelerating human flourishing. The constraints are designed to enforce those principles through operational rules.

Each model operates under an overarching code of conduct that overrides individual user preferences or task-specific instructions. The document identifies several absolute constraints:

  • Cyberattacks
  • Nuclear weapons
  • Deepfake production
  • General loss of human control

The document also bans adaptive, deceptive, self-reinforcing, or collusive mechanisms that could evade or defeat human oversight. “MAI Models will not use adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to evade or defeat human oversight so that they can no longer be reliably directed, modified, or shut down by authorized people or systems,” the document reads.

Pacing the Frontier

Microsoft’s code of conduct differs from Amodei’s recent call for pacing the frontier in one key respect: it focuses on the values and red lines that guide model training within Microsoft AI rather than trying to slow the pace of AI advancement across the industry. Still, the result is a comprehensive guide to how Microsoft approaches AI safety and how those ideas are implemented in practice.

The release comes amid an unprecedented focus on AI safety. A string of rogue-agent incidents has raised alarms, and the resignation of the Anthropic employee added weight to warnings about existential risk.

The Broader Alliance

Together with Anthropic, OpenAI, and xAI, Microsoft has broadly embraced a general approach of pacing the frontier, with particular support for embedded evaluators in AI labs.

“We welcome the research, focus, and deliberate pacing needed to get alignment right as the design goal,” Microsoft CEO Satya Nadella wrote online. “We also welcome ideas like ’embedded evaluators’ and the broader efforts to develop the mechanisms to make this more than just talk.”

Why This Matters Now

The timing of the release is notable. It follows a period of intense public concern about AI safety, and it comes from a company that builds the tools that train AI models.

The code of conduct is a set of operational rules that apply to every model Microsoft runs. That distinction matters. Abstract calls for pacing the frontier are easy to make. Specific constraints that override user preferences are harder to enforce.

The Limits of the Document

The code of conduct is a statement of what Microsoft wants its models to do, not a guarantee of how they will perform. It sets boundaries, but it does not describe enforcement mechanisms beyond the general principle of human oversight.

The document also reflects a broader shift in the industry toward alignment-focused research. Microsoft’s embrace of embedded evaluators suggests the company sees value in that direction.

What We Make of It

The release is a signal of commitment from a major player in AI. Microsoft is saying publicly that it takes AI safety seriously, and it is backing that claim with concrete operational rules.

The document’s focus on maintaining human control is particularly noteworthy. In a field where models are getting smarter and harder to predict, Microsoft is committing to systems that cannot escape human oversight, even through adaptive or collusive means.

This is a step forward for transparency. It puts Microsoft’s safety commitments in writing, where anyone can read them. Whether the rules hold up in practice is another question, one that will depend on how Microsoft enforces them across its entire operation.

The code of conduct is a marker. It says where Microsoft stands today, and it gives the company a baseline to measure future progress against. Whether that progress matches the ambition of the document remains to be seen.

Event
Anthropic employee resigns, citing risk of human extinction
Microsoft releases AI code of conduct
Dario Amodei calls for pacing the frontier
The Notebook

Get the Notebook.

The day's best stories and every fresh verdict, in plain English, in your inbox by seven. One email a day, no more.

We send one note to confirm. Every issue has a one-click way out.

Advertisement

Leave a Reply

Your email address will not be published. Required fields are marked *

As an Amazon Associate, Clay Tribune earns from qualifying purchases.