AI News Feed
Market watch
Companies

AI Safety Evaluators Step Into Spotlight as OpenAI, Anthropic Seek Outside Scrutiny

Independent AI evaluators are gaining prominence as Anthropic and OpenAI seek outside safety assessments, with funding, access and reporting still unresolved.

The evaluators include Model Evaluation and Threat Research, or METR, Apollo Research and Transluce. Their primary role has been to assess AI model capabilities and risks and to call attention to instances where the technology behaves badly, according to CNBC. Two months ago, they occupied a relatively sleepy corner of the multitrillion-dollar AI industry; now they are being asked to help.

With no federal push for regulations, the evaluators have taken on outsized importance. Anthropic CEO Dario Amodei pledged last month to embed independent evaluators in his company, a move OpenAI CEO Sam Altman quickly endorsed. President Donald Trump supported the idea, as did most of the largest U.S. tech companies. But questions remain about how those third parties should be funded, what level of access they will have and what the reporting structure will ultimately look like.

"To a degree, the problem, as always, is money," Suresh Venkatasubramanian, a computer science professor at Brown University, told CNBC. "Who is paying for these companies to do their work? How are they going to support them? You need an ecosystem, you need a viable business model for this."

Right now, Anthropic, OpenAI and the infrastructure partners profiting from the AI boom are writing the rules. Critics say that is like asking the biggest banks to protect the public from a financial crisis or allowing pharmaceutical companies to put drugs on the market without regulatory clearance. Trump recently praised AI executives for their "tremendous self-policing" and signaled that he intends to leave companies to their own devices, unwilling to impede an industry that is driving the economy and stock market. In a voluntary accord presented in late September, Trump encouraged AI companies to "partner with an independent external auditor or evaluator."

Amodei touched off the conversation in a viral essay last month, when he called for a "slower pace" in advanced model development after researchers left his company and voiced concerns about existential threats the technology poses.

As AI labs move to put evaluators in place, friction is already emerging. OpenAI fired three employees last week for "violating our policies on accessing and handling sensitive company information," according to a spokesperson. Two of those employees, Mikita Balesni and Tomek Korbak, said they believe they were dismissed because of how they communicated with third-party evaluators.

"My former colleagues are telling me they are confused about what to believe," Balesni wrote in a post on X on Thursday. "They also are afraid to speak, and worry their personal phones will be searched for messages to us and third parties. I worry the pervading fear to speak up and engage with third parties will mean OpenAI will cut corners on safety behind closed doors."

OpenAI disputed that characterization and said in a post on Friday that it is "actively finalizing contracts with third-party safety assessors and will announce details in the coming weeks." The company wrote, "We are committed to embedding external assessors and continue to make close collaboration with independent safety organizations a core part of our safety work." An OpenAI spokesperson said in an emailed statement that its upcoming work with evaluators "builds on existing collaboration with independent safety organizations," including METR and Redwood Research. Anthropic did not respond to CNBC's request for comment.

The evaluator ecosystem consists mostly of small organizations, including METR and Apollo Research, and larger accounting and auditing firms such as Accenture. AI labs have worked with evaluators in limited capacities, but Andrew Freedman, CEO of policy nonprofit Fathom, said the field is quickly maturing. "I've worked in politics and policy for the last 20 years of my life, and I've never seen an issue move so fast on so many different political spectrums," Freedman told CNBC. He said he expects an "influx of capital" to flow into the ecosystem.

Rayan Krishnan, CEO of independent evaluator Vals AI, said his for-profit startup builds benchmarks to measure how AI models perform on industry-specific tasks. The company has grown from eight employees to roughly 30 this year and announced a $40 million funding round in August. METR, a nonprofit, announced in August that it had raised commitments of around $71 million over the last six months.