Microsoft AI CEO Mustafa Suleyman says the industry needs to slow down before it kills us all, and he has a specific target for his criticism: Anthropic.
Suleyman appeared on the Decoder podcast to discuss the spiraling debate over AI safety and regulation. He recently published a 37-page statement called the “Humanist AI Code of Conduct,” which lays out Microsoft’s principles around AI development and its philosophy on thorny issues like AI consciousness. He also released a companion essay this week specifically criticizing Anthropic’s philosophy around AI consciousness and its role in the broader alignment debate.
The crux of his argument is that alignment — the idea that models can be made to behave correctly through intrinsic design — is one tool among many, but it is not the whole solution. He argues that containment is the first priority, followed by alignment to human values.
Containment First
Suleyman opened the conversation by framing the current moment in stark terms. He described a scenario where AI systems without safety guardrails are capable of impressive and scary hacking capabilities, citing recent incidents involving Hugging Face and OpenAI.
He pushed back on the idea that alignment is the only path forward. “I think it’s one important element, but it’s not the only one,” he said. He pointed to his book, where he wrote about the idea of containment three or four years ago, and argued that containment is not possible and proliferation is inevitable. “In 99 percent of cases, that’s a really good thing,” he said. “We want technologies to spread far and wide as quickly as possible so that everyone can enjoy the benefits.”
But he also warned that the rapid pace of AI development means the industry will soon face a harder question. “If you just roll forward five years, we always get caught up in the next quarter or next year and everyone gets a little bit flustered and has a big disagreement,” he said. “But if you just imagine the difference between GPT-3 three years ago and GPT-6 today, and then imagine the difference between GPT-6 and GPT-9.” That is three orders of magnitude more compute, 1,000 times more FLOPS applied to pre-training with reinforcement learning for these runs, and we’re going to have something which is breathtaking. It’s going to be absolutely incredible at so many things.
He argued that the industry has to address containment first. “The first thing is that we have to make sure they’re contained, their agency is limited, they don’t escape the box, they don’t reward hack, that they are controllable, and they follow our instruction,” he said. Only then can alignment be addressed.
Alignment’s Limits
Suleyman drew a sharp parallel between alignment failures and automotive engineering. “If I designed a car and 10 percent of the time the brake pedal decided to go attack my neighbor’s house, I would be like, ‘This car doesn’t work. The very technology of brakes is broken. I need a new idea,'” he said.
He framed the question as a binary: either alignment has potential to be 100 percent safe, or it does not. “If it’s possible for alignment and the techniques of alignment to be successful or useful or consistent, then maybe I understand the debate in a different way,” he said.
The Microsoft Position
The Humanist AI Code of Conduct lays out Microsoft’s principles in stark terms. The company’s position is simple: technology is here to serve humanity. It should be a subordinate, controllable, aligned force that does good in the world. If it doesn’t achieve that, then we should reject it.
Suleyman acknowledged that the industry has not yet reached that point. “It has not happened today, but it is now, I think given what’s happened over the summer with Hugging Face and OpenAI, pretty clear that these systems without the safety guardrails are capable of really impressive and quite scary hacking capabilities,” he said.
The Anthropic Criticism
Suleyman’s companion essay targets Anthropic directly. He argues that the company has gotten confused about the concept of model welfare in dangerous ways. His criticism of Anthropic’s philosophy around AI consciousness is part of a broader push to clarify what the alignment debate actually means.
The interview was lightly edited for length and clarity.
Key Dates
| Date | Event |
|---|---|
| Recent weeks | Suleyman appears on Decoder podcast |
| Three years ago | GPT-3 released |
| Five years from now | GPT-9 projected |
Suleyman’s argument lands with a mix of urgency and precision. He is not arguing that alignment is useless. He is arguing that it is not the only lever, and that the industry has been treating it as if it were.
The tension in the exchange is real. Microsoft’s own code of conduct says AI should serve humanity and be rejected if it fails. Suleyman’s practical advice is to slow down. The two positions are not contradictory, but they sit uncomfortably close together.
The industry has spent years building systems that can follow more complex instructions across multiple time steps.
Get the Notebook.
The day's best stories and every fresh verdict, in plain English, in your inbox by seven. One email a day, no more.

