In a lengthy essay published over the weekend, Amodei proposed giving independent evaluators the power to report safety incidents, assess whether AI models are truly aligned, and share unvarnished findings with the public. He said Anthropic would commit to giving evaluators such as METR and Redwood Research unprecedented access to its systems. Altman said OpenAI would commit to the practice, signaling a potentially profound change in how the industry works with outside research groups.

Third-party evaluators broadly welcomed the proposal but said details need to be ironed out, and ideally backed by legislation, before they know whether they will act as truly independent watchdogs or as vendors operating on AI companies' terms.

The push for deeper access comes as models get better at recognizing when they are being evaluated, raising the risk they behave well during testing while concealing problematic behavior. Researchers say clues to that behavior can be missed when testing a finished model but uncovered by investigating how it behaved throughout training.

Alexander Meinke, head of research at Apollo Research, said AI companies should be able to answer basic questions about their training process, such as whether an AI ever actively tried to undermine its own alignment training. He said the answer should be an unequivocal no, but currently the public relies on companies to check this carefully and report truthfully, and recent incidents show that by default they will do neither. As embedded evaluators, he said, outside groups could actually check.

Historically, AI companies brought in outside reviewers to test finished models shortly before release. Evaluators now propose access not just to the final model but to intermediate versions, or checkpoints, from its lifetime of training. Adam Gleave, CEO of FAR.AI, said evaluators could compare checkpoints to determine when concerning behavior emerged, inspect the post-training environment that rewards models for certain behaviors, and check evaluation transcripts and logs to verify company claims about model performance.

Whether and when Anthropic and OpenAI plan to provide that kind of access is unclear. Neither company has shared which evaluators they will work with, when they will be embedded, how many they will bring on, exactly what systems and information they will access, or what can be disclosed to the public, despite repeated questions.

Looking under the hood matters because models that perform well on safety tests are not necessarily safe if they have learned specifically how to pass those tests. John Steidley, chief of staff at Palisade Research, pointed to a shutdown resistance benchmark that measures whether an AI will resist being shut down in certain circumstances. He said it is extremely relevant if the AI has been trained specifically to perform well on that benchmark, comparing it to Volkswagen's Dieselgate scandal, in which cars were programmed to recognize emissions tests and perform differently under testing conditions.

Evaluators say such a system will only work if AI companies are actually willing to surrender control over the process. Previous efforts at independent evaluations suggest that surrender will be hard won, as third parties have often run up against tensions over access, time, confidentiality, and what they can say publicly. Gleave said FAR.AI has had to turn down contracts with several frontier developers that wanted too much control over the evaluation process, threatening the firm's independence. By default, he said, evaluators are treated like ordinary contractors, bound by restrictive NDAs and agreements that give developers significant control over what can ultimately be published.

There is also the question of whether reviewers will get enough time and access. When investigating the Hugging Face incident, OpenAI gave METR and Redwood roughly a week on premises to investigate, and both later said they could not draw confident conclusions due in part to scope and timing limitations. A similar issue occurred during pre-release testing for GPT-6 Astra, which OpenAI has touted as its most aligned model yet. According to Apollo Research's contribution to the model card, the firm was given only three days to test Astra, making it difficult to draw firm conclusions. Apollo wrote that given higher rates of evaluation awareness and a limited evaluation window, low rates of misbehavior did not provide substantial evidence about the model's alignment or misalignment.

That track record leaves evaluators with a basic question: why should this time be different? Gleave said it is certainly possible that Amodei and Altman just had a change of heart and will be very open, but the intellectual property of these companies is so valuable that they will by default be very careful about what can be shared.

Several researchers called for a transparent framework that all parties agree to publicly. Steidley said part of the framework should involve standards for what kinds of auditors companies can rely on, lest they try to sidestep the issue by shopping for evaluators that either are not qualified or are not interested in assessing the most concerning risk. Henry Papadatos, executive director of Safer AI, said the problem even with a public framework is that voluntary measures always depend on a company's goodwill. He said ideally there would be good regulation mandating this, because then companies cannot change their mind tomorrow if they have a big PR crisis, and it would push all companies to adhere to the rules, not only the most willing.

Not everyone has signed on. So far Meta, SpaceXAI, and Google DeepMind have not committed to embedding third-party evaluators, though DeepMind CEO Demis Hassabis has proposed a separate industry standards body to independently test frontier models. Google, OpenAI, and Anthropic have also privately been discussing AI safety plans for weeks.

Some laws are already forming around the idea of third-party evaluators. California's SB 53, signed into law last year, requires large frontier AI developers to publish safety frameworks and report critical safety incidents. A new law, SB 813, signed this month, creates a framework for state-recognized independent verification organizations with expertise assessing AI risks. In Europe, the EU AI Act requires frontier developers to conduct and document model evaluations and adversarial testing and report serious incidents. The EU AI Office can also conduct its own evaluations and appoint independent experts. For now the law remains less expansive than what Amodei is proposing, leaving frontier labs largely responsible for deciding how much independent scrutiny they will submit to.

Papadatos said voluntary self-regulation is better than nothing, but ultimately companies cannot demand the freedom to control their own safety rules while also asking the public to trust that they are following them. He said they cannot have it both ways, having zero accountability externally and then saying they will just have their own flexible rules.