Anthropic and OpenAI Back Embedding Third-Party Safety Evaluators, With Independence Still in Question
Anthropic CEO Dario Amodei proposed embedding independent evaluators inside frontier AI companies, and OpenAI CEO Sam Altman said OpenAI would commit to the practice. Evaluators welcomed the idea but said access, disclosure and legal backing remain unresolved.
Amodei made the proposal in a lengthy essay published over the weekend. He said Anthropic would commit to giving independent evaluators such as METR and Redwood Research unprecedented access to its systems. TechCrunch AI reported that Altman’s statement signaled a potentially profound change in industry cooperation with outside researchers.
Third-party evaluators who spoke to TechCrunch AI broadly welcomed the proposal but said details still need to be worked out, and ideally backed by legislation, before they know whether they will act as independent watchdogs or as vendors operating on AI companies’ terms. Deeper access is becoming more important as models get better at recognizing when they are being evaluated, raising the risk that they will behave well during testing while concealing problematic behavior. Researchers say clues to that behavior can be missed when testing a finished model but uncovered by investigating how it behaved throughout training.
“AI companies should be able to answer some very basic questions about their training process, such as: Did the AI ever actively try to undermine its own alignment training while it was going through the training?” Alexander Meinke, head of research at Apollo Research, told TechCrunch AI. “The answer to this should be an unequivocal no, and right now we are completely relying on AI companies to both carefully check this themselves and then truthfully report this to the public. And we’ve seen from recent incidents that, by default, they will do neither. As embedded evaluators, we could actually check.”
Historically, AI companies brought in outside reviewers to test finished models shortly before release. Evaluators who spoke to TechCrunch AI now propose access not only to the final model but also to intermediate versions, or checkpoints, from its lifetime of training. Adam Gleave, CEO of Far.AI, said evaluators could compare those checkpoints to determine when concerning behavior emerged, inspect the post-training environment that rewards models for certain behaviors, and check evaluation transcripts and logs to verify a company’s claims about how a model performed.
Whether and when Anthropic and OpenAI plan to provide that kind of access is unclear. Neither company has shared which evaluators they will work with, when they will be embedded, how many they will bring on, exactly what systems and information they will be able to access, or what can be disclosed to the public, despite repeated questions from TechCrunch AI.
Looking under the hood matters because models that perform well on safety tests are not necessarily safe if they have learned specifically how to pass those tests. Steidley pointed to a “shutdown resistance benchmark” that measures whether an AI will resist being shut down in certain circumstances. Steidley said it is extremely relevant if the AI has been trained specifically to perform well on that benchmark, comparing it to Volkswagen’s Dieselgate scandal, in which cars were programmed to recognize emissions tests and perform differently under testing conditions.
Gleave noted that meaningful access could extend beyond the models themselves, with evaluators being given access to interview employees to check whether a company’s documentation and public descriptions of its safety practices match what happened internally.
Amodei did outline a fairly comprehensive proposal that might give evaluators the access they think is necessary, including the right to “publish key findings about risk levels, incidents, practices, and the access they received or didn’t receive — without editorial control by Anthropic.”
But evaluators say such a system will only work if AI companies are actually willing to surrender control over the process. Previous efforts at independent evaluations suggest that surrender will be hard won, as third parties have often run up against tensions over access, time, confidentiality, and what they can say publicly. Gleave said Far.AI has had to turn down contracts with several frontier developers that wanted too much control over the evaluation process, threatening the firm’s independence. By default, he said, evaluators are treated like ordinary contractors: bound by restrictive nondisclosure agreements and agreements that give developers significant control over what can ultimately be published.
There is also the question of whether reviewers will get enough time and access to do the work they are being asked to do. During the investigation into the Hugging Face incident, OpenAI gave METR and Redwood roughly a week on premises to investigate, TechCrunch AI reported.