AI News Feed
Market watch
Companies

Anthropic and OpenAI's proposed AI risk evaluators lack power to halt models, experts say

Anthropic and OpenAI plan embedded AI safety evaluators with deep access but no power to halt model training or release.

Amodei proposed the arrangement as part of a plan to slow the advance of increasingly capable models, after former Anthropic researcher Jacob Coxon resigned and warned that frontier labs were racing toward systems they might not control. In an essay published last weekend, Amodei committed Anthropic to giving third-party evaluators access comparable to internal risk teams and the right to publish findings without the company's editorial control, subject to limited redactions. The essay has led to a regulatory split between factions within the AI community and government, including President Trump, according to CNBC. He cited embedded bank supervisors as a precedent, writing that his proposal “has precedent in the banking industry, which sometimes involves regulatory ‘supervisors’ embedded along with employees.”

Julie Andersen Hill, dean of the University of Wyoming College of Law and an expert on banking regulation, said the comparison is not accurate without the power to flip the “kill switch” on the whole operation. “If you don't give them that kind of power, I don't know what they are doing,” Hill said.

At the largest banks, government examiners have offices inside the institution, access to internal systems and employees, and a continuous presence. They can direct a bank to stop a practice, restrict its growth, force management changes and, in extreme cases, close it, Hill said. Anthropic's proposed evaluators would be able to investigate and report, but Amodei's plan does not give them comparable enforcement power or legal authority to prevent a model from being trained or released. “That's fundamentally different, because bank regulators have a lot more power than that,” Hill said.

Amodei wrote that evaluators are needed to provide “a neutral third party who can actually see the details” but acknowledged that Anthropic still determines what it includes and omits from its existing public disclosures. Details released to date suggest the evaluators would have extraordinary access to frontier AI systems but limited formal authority over the companies developing them. Neither Anthropic's proposal nor OpenAI's existing third-party evaluation framework gives outside evaluators independent authority to halt the development or deployment of a model.

OpenAI has also committed to a similar safety arrangement but has not yet released details of how its new embedded-evaluator commitment will work. Anthropic and OpenAI did not respond to requests for comment, CNBC reported.

Albert Ziegler, head of AI at cybersecurity company XBOW, who leads a team that evaluates model capabilities, said his firm has received early access to unreleased models from Anthropic, OpenAI and other major developers, and that he was personally involved in those evaluations. XBOW conducts the work in its own environment rather than as commissioned testing, he said, and generally shares its findings with the model providers.

The everyday reality, at least to date, is less cinematic than the existential framing, Ziegler said. Amodei wrote that the new regulatory embeds are necessary because within six to 12 months, a misaligned swarm of agents could potentially seize large parts of the internet and inflict hundreds of billions of dollars in damage. Ziegler said his team may find that a model produces nonsense under an unusual formatting request or that an external safety checker needs to intervene more frequently. “But the kind of insidious, catastrophic consequences produced by subterfuge combined with unprecedented abilities that people are afraid of — that's not something we've seen ourselves,” he said.

Black-box testing can reveal whether a model can perform a dangerous task, while determining whether the larger system is dangerous may require access to the instructions surrounding the model, its tools and permissions, its safety controls and logs of attempted actions, Ziegler said. Even then, he added, a serious risk may emerge only under a combination of circumstances the evaluator never triggers. “It's true that we don't have any veto power,” he said. An evaluator can uncover and document risks a developer missed and “compel an informed decision before release,” he said. But the ultimate decision remains with the company.

Ties to top AI labs remain a concern. Even with its broader authority, the bank model is imperfect, according to Hill. Supervisors have failed to prevent major collapses and are regularly accused after crises of becoming too close to the institutions they oversee.