Anthropic and OpenAI Propose Embedded AI Safety Evaluators: What Changed and What It Costs
Leading AI labs Anthropic and OpenAI have proposed embedding independent safety evaluators into their model development pipelines, but experts question whether these assessors can truly be impartial. Over 100 AI researchers have signed a letter demanding more rigorous oversight.
The Proposal: Embedding Safety Evaluators Directly in AI Development
In September 2026, Anthropic and OpenAI announced plans to embed third-party safety evaluators directly within their model development processes [7]. This move aims to improve the alignment of advanced AI systems—ensuring they behave as intended under real-world conditions. The idea is that having independent assessors review models during training could catch harmful behaviors earlier than post-deployment audits.
This shift marks a change from traditional safety evaluations, which typically occur after a model is released or in controlled testing environments [1]. By integrating evaluators earlier, the companies argue they can build safer systems from the ground up. However, critics remain skeptical about whether these evaluators will maintain independence when embedded within the same organizations that develop the models.
The Skepticism: Can Evaluators Be Truly Independent?
More than 100 AI experts have signed a public letter urging Anthropic and OpenAI to adopt safety evaluators that are not only independent but also externally accountable [2][4]. They argue that embedding evaluators inside the same labs that build the models creates a conflict of interest. If the evaluators are funded or managed by the companies, their findings may be influenced by internal pressures to avoid delays or negative publicity.
“An evaluator embedded within a company cannot credibly claim independence,” said one signatory in an interview cited by TechCrunch [6]. “It’s like asking a bank to audit its own loan practices.”
The concern is not just theoretical. Experts point to past instances where AI safety concerns were downplayed or ignored during product launches, even when internal warnings existed [3]. The proposed embedding model does little to address this structural issue.
Industry and Government Involvement: A Fragmented Landscape
While Anthropic and OpenAI have taken the lead on embedding evaluators, other major players remain outside the initiative. Google DeepMind has not formally joined the effort, though it has participated in broader AI safety discussions with both companies [8]. Demis Hassabis, CEO of DeepMind, has instead proposed establishing a separate industry-wide standards body to oversee model evaluations—a move seen by some as an alternative path to centralized oversight [7][8].
Meta has also not signed on to the embedding proposal, despite its own growing role in AI safety research. Meanwhile, claims about SpaceX AI being a key player in this space appear unfounded; there is no publicly recognized entity named “SpaceX AI” actively competing with OpenAI or Anthropic in model development [5].
What the Experts Are Asking For
The letter from over 100 experts calls for safety evaluations to be conducted by fully independent organizations—free from financial ties to any AI lab—and to publish their findings openly. They emphasize that evaluation timelines must be long enough to detect subtle, emergent behaviors that may not surface in short testing windows.
Critics have pointed out that some existing evaluation frameworks operate on extremely tight schedules. For example, Apollo Risk reportedly completed its assessment of OpenAI’s GPT-6 Astra—though this model has not been officially released—in just three days [6]. Experts argue that such brief evaluations are insufficient for detecting complex alignment failures or deceptive behaviors that may only emerge under prolonged use.
The Bigger Picture: Regulation and Public Trust
The debate over embedded evaluators comes amid increasing pressure from governments to regulate AI development. While political headwinds have included opposition from former figures like Donald Trump and David Sacks—neither of whom held office in September 2026—the current administration under President Joe Biden has shown growing interest in AI governance [1][3].
The push for independent oversight reflects a broader demand for transparency. As AI systems grow more capable, the public and policymakers alike are demanding proof that these technologies are safe before widespread deployment. The embedded evaluator model may represent a step forward—but only if it includes real independence, long-term testing, and public accountability.
What Changed and What It Costs
The core change is not in technology, but in process: moving safety evaluations earlier in the development cycle. However, the cost of this shift—both financial and reputational—is borne by the companies themselves. Embedding external evaluators likely increases R&D timelines and operational costs. Yet, without trust from users, regulators, and investors, these added expenses may be unavoidable.
For independent researchers and watchdog groups, the real cost is credibility. If embedded evaluators are perceived as rubber stamps, their role will erode. The challenge ahead is not just technical, but institutional: building systems where safety is not just checked, but genuinely prioritized.
Sources
- Anthropic and OpenAI Leaders Commit to Independent Evaluators for Powerful AI Models - Americans for Responsible Innovation (ari.us) — 2026-09-12
- Anthropic and OpenAI want to embed safety evaluators. Will they really be independent? | TechCrunch (techcrunch.com) — 2026-09-16
- More Than 100 AI Experts Sign a Letter Saying that OpenAI & Anthropic Need Independent AI Safety Evaluators (www.ibtimes.com) — 2026-09-18
- Anthropic and OpenAI need truly independent safety evaluators, experts say in public letter (www.cnbc.com) — 2026-09-18
- Anthropic and OpenAI want to embed safety evaluators. Will they really be independent? | daily.dev — 2026-09-16
- Anthropic and OpenAI want to embed safety evaluators. Will they really be independent? | TechCrunch (techcrunch.com) — 2026-09-16
- OpenAI, Anthropic, Google have been in talks on AI safety for weeks | TechCrunch (techcrunch.com) — 2026-09-15
- Anthropic and OpenAI want to embed safety evaluators. Will they really be independent? (www.industryevents.com) — 2026-09-16