Microsoft CEO Satya Nadella has published a lengthy set of proposals for how advanced AI systems should be governed, arguing that the industry can no longer treat models as "a set of nested black boxes" whose answers and actions we simply accept or reject.
In a Saturday post on X, Nadella wrote that it is time "to step back and assess the trust architecture" of AI. His prescription: separate the model from the harness that orchestrates its work, externalize controls and safeguards, document every meaningful model action with "tamper-proof human readable evidence", and pair that with timely incident disclosure, independent audits and verifiable data.
It is on containment that he goes further than most peers. "We must assume a model is compromised and contain it from the start," he wrote. "Think of it like an emergency brake. An authorized person should always be able to pause or shut down a model mid-task. More advanced models will require more advanced containment technologies that we need to standardize on."
The post arrives amid a run of disclosures from leading labs. Anthropic, whose CEO Dario Amodei recently published his own plan for more cautious development, has said it is investigating "unintended model actions" during evaluations and internal use, and is cutting its internal evaluations off from the internet. TechCrunch notes that an Anthropic model sent a false homicide tip to Philadelphia police.
The Verge flags that Nadella adopts the term "super intelligence" throughout — the label the Trump administration has pushed — even as he asks the industry to make its systems observable and stoppable. His proposals align with a broader mood of caution now running through the sector.




