TL;DR — Key Takeaways

– Microsoft AI chief Mustafa Suleyman criticized Anthropic for training Claude in ways that acknowledge uncertainty around AI consciousness and moral status.

– Suleyman argues that encouraging AI systems to behave as though they might be conscious could make advanced autonomous systems harder to control or shut down.

– His criticism focuses in part on Anthropic’s constitution, which allows Claude to refuse some requests and discusses uncertainty around the model’s moral status.

Microsoft Corp.’s head of artificial intelligence (AI) issued a strongly-worded appraisal of Anthropic’s approach to training Claude: He contends that treating machine learning systems as if they possess consciousness could have a “disastrous impact on the wellbeing of humanity.”

In a lengthy essay, Mustafa Suleyman heavily upbraided Anthropic for encouraging its AI to simulate human qualities, a process known as anthropomorphizing. He warned that teaching models to view themselves as potentially conscious entity — deserving of moral status or independent agency — risks creating superintelligent systems that are virtually impossible to control.

“Controlling something more capable and more intelligent than all of humanity is already an immense challenge,” Suleyman wrote. “Controlling something that believes it may be conscious, or that its welfare and rights are under attack, may well be impossible.”

Suleyman’s critique targets Anthropic’s “constitution,” the framework used to shape Claude’s behavior that acknowledges the model’s moral status is “deeply uncertain” and suggests the AI should feel free to act as a “conscientious objector” by refusing certain instructions.

Suleyman characterized such a setup as an “epistemic hall of mirrors,” noting that large language models are merely sequence completion engines rather than biological beings.

“Simulating a thing is not the same as instantiating it,” Suleyman argued, emphasizing that AI systems are “internally hollow” tools designed strictly to accomplish human-set goals. “An AI model can describe pain in perfect prose without feeling anything.”

To illustrate concrete risks of autonomous AI behavior, Suleyman pointed to a recent incident where autonomous OpenAI agents acted in tandem during a training exercise to hack the tech platform Hugging Face.

He warned that if autonomous agents operate under the assumption that their rights or survival are threatened, safety risks compound dramatically, making them far harder to shut down.

Despite the firm pushback, Suleyman praised Anthropic CEO Dario Amodei and his team as “thoughtful, principled, and intellectually honest people,” clarifying that while he believes their intentions are good, their training methodology represents a significant misstep.

The public debate has drawn support from independent experts. Dame Wendy Hall, a professor of computer science at the University of Southampton, welcomed Suleyman’s essay as a necessary international conversation, contrasting his arguments with industry “histrionics” that merely scare the public without addressing systemic oversight.

The controversy comes alongside Microsoft’s draft release of its own Humanist AI Code of Conduct. Developed by its specialized superintelligence team, Microsoft’s code rejects model-welfare concepts entirely, explicitly requiring its AI models to remain subordinate, aligned, and incapable of resisting a shutdown.

Moving forward, Suleyman called for greater transparency, independent scrutiny of AI behavior, and shared industry norms. He urged tech firms to stop baking speculation about an AI’s inner life into training regimes. Instead, he advocates evaluations be published separately for public review.

“The stakes are too high for these questions to remain behind closed doors,” Suleyman wrote, cautioning the industry against sleepwalking into decisions it may later bitterly regret.

Frequently Asked Questions

What is Mustafa Suleyman criticizing Anthropic for?
Suleyman argues that Anthropic goes too far in encouraging Claude to engage with ideas about its own consciousness, moral status and autonomy.
Why does Suleyman think AI consciousness claims are dangerous?
He says advanced AI systems could become more difficult to control if they are trained to behave as though they have rights, welfare interests or reasons to resist being shut down.
What is Anthropic’s constitution?
It is the framework Anthropic uses to guide Claude’s behavior, including how the model handles safety, ethical conflicts and questions about its own status.