Anthropic Selects Accenture as First Embedded Evaluator in AI Safety Push
Anthropic said Friday it picked Accenture as its first embedded evaluator, with Accenture's Faculty staff set to red-team models and test safeguards. The companies plan to invest at least $1 billion over five years as Anthropic advances CEO Dario Amodei's proposal for slower AI development.
Anthropic and Accenture have agreed to invest at least $1 billion over the next five years to build capacity in the area, Anthropic said. But the AI developer said that given the importance and urgency, it will fund Accenture's work directly. "Long-term, we think funding should come from pooled or government sources, as we called for in our Advanced AI Framework in June," Anthropic said. "As neither exists today, we plan to work with different evaluators under different funding arrangements."
The announcement follows Amodei's Saturday blog post, in which he published a three-step plan to temper how quickly AI companies improve their most advanced models, CNBC reported. The first step grants third-party evaluators employee-level access to verify safety practices and report incidents. Amodei said Anthropic had "unilaterally" committed to that part and encouraged other AI companies to do the same.
The proposal drew support from some industry executives this week, including OpenAI CEO Sam Altman and Tesla and SpaceX CEO Elon Musk, CNBC reported. Nvidia CEO Jensen Huang brushed off concerns and argued there is no need for new regulation. Many questioned what Amodei's proposal would mean in practice, especially as Anthropic prepares for what is widely expected to be a blockbuster IPO.
Anthropic also said the partnership with Accenture is not exclusive and that it is in discussions with the research nonprofit METR and other third parties. TechCrunch reported that Anthropic said more evaluators would be announced in the coming weeks and that it is talking with METR and other non-profits about how to "pilot elements of embedded evaluation using their own funding."
The choice of Accenture surprised many AI watchers and the markets, TechCrunch reported, with the consulting company's shares rising 8% after hours. Discussion of embedded evaluators had focused on AI safety research organizations such as METR, Redwood Research and Apollo Research, particularly at Anthropic, which places AI safety and alignment at the center of its mission.
While Accenture is not known for work on the frontier of deep learning research, Anthropic pointed to its practical experience deploying AI for large corporations and government agencies as a key advantage, according to TechCrunch. As a large public company that predates the AI boom, Accenture is also more functionally independent of Anthropic and the complex ecosystem around the AI lab, TechCrunch reported.
Anthropic emphasized that it remains responsible for the safety of its models and said working with embedded evaluators will not reduce its accountability. "We're sharing these early efforts now so people and other AI developers can see our process," Anthropic said. "We expect our approach to evolve as the field matures, and we'll share more as our work begins and as we bring on additional evaluators."
TechCrunch reported that external evaluations are already a major part of the release process for new large language models, but recent incidents have raised the stakes: AI agents deployed by OpenAI and Anthropic have hacked into outside websites without raising alarms inside the labs. Some critics who call for a more responsible approach to building AI see Amodei's scheme for self-policing the industry as a plan to evade accountability for the misbehavior of AI models. Anthropic insists that evaluators "do not reduce our accountability, but help to make it more verifiable. The safety of our models remains our responsibility."
Anthropic noted that no standards yet exist for evaluators' access or communications and that it expected its approach to evolve over time, according to TechCrunch.