AI News Feed
Market watch
Companies

Microsoft's Nadella Calls for AI 'Emergency Brake' and Assumption That Models Are Compromised

Nadella calls for AI 'emergency brake,' saying models should be assumed compromised from the start.

In a post on X, Nadella wrote that AI systems need stronger safety architecture. "We need to surround non-deterministic models with strong, deterministic system design, human controls, and reliable operating procedures, and establish industry standards where existing ones are insufficient," he said, according to CNBC. He added that "treating frontier closed and open weight models like insider risks is a way to build such a system."

Nadella's proposal centers on what CNBC described as "principles of observability." Those include model diversity; a human-readable footprint of the model's actions; continuous system testing; independent controls and auditability; containment; and incident disclosure. He wrote that the most trustworthy "Super Intelligence" system would not be the one with the model people trust most, but "the one that enables us to trust the model the least."

The Verge reported that Nadella said AI can no longer be treated as a "set of nested black boxes" whose advice and actions are simply accepted or rejected. He called for a more transparent system in which models can be contained and observed and leave behind "tamper-proof human readable evidence."

On containment, Nadella went further than some industry recommendations, The Verge reported. "We must assume a model is compromised and contain it from the start. Think of it like an emergency brake," he wrote. "An authorized person should always be able to pause or shut down a model mid-task. More advanced models will require more advanced containment technologies that we need to standardize on." The Verge noted that Nadella referred to AI as "super intelligence" throughout the post.

TechCrunch reported that Nadella also called for stepping back and assessing the "trust architecture" of AI. "We can't treat Super Intelligence as a set of nested black boxes and simply accept or reject its recommendations, answers, and actions," he wrote, using the Trump administration's preferred term for AI, according to TechCrunch. His approach includes "separating the model from the harness that orchestrates its work" and "externalizing controls and safeguards." He also called for "every meaningful model action" to be documented with "tamper-proof human readable evidence."

The comments come amid warnings from technology executives and researchers, including Microsoft co-founder Bill Gates, Anthropic CEO Dario Amodei, OpenAI CEO Sam Altman and SpaceX CEO Elon Musk, about insufficient AI safety protocols and claims that the technology is moving too fast, CNBC reported. Last month, an AI researcher quit Anthropic and accused the company and OpenAI of "gambling with our lives," CNBC reported. Later that day, an Anthropic alignment lead focused on AI safety said there was a greater than 10% chance the technology could "kill all humans" within the next decade, according to CNBC.

President Donald Trump has repeatedly dismissed AI extinction risks and instead emphasized that the industry needs to stay ahead of China, CNBC reported. Trump recently introduced a new "AI Force," led by Director of National Intelligence Jay Clayton, to help facilitate the industry and root out bad actors, according to CNBC.

TechCrunch reported that Nadella's post came as leading AI companies acknowledge more and more incidents in which they appeared to lose control of their models, and after Anthropic CEO Dario Amodei published a plan for more cautious AI development. The Verge noted that many of Nadella's recommendations align with those heard from others in the industry: timely incident disclosure, independent audits, verifiable data and containment.

Editor's Summary

Microsoft CEO Satya Nadella called for AI systems to be contained, independently controlled and equipped with an emergency brake that can pause or shut down a model mid-task. He said models should be assumed compromised from the start and subject to observability, audits and incident disclosure. The proposal adds to an industry debate over AI safety, following warnings from researchers and executives and a contrasting emphasis by President Donald Trump on competing with China.