What AI labs' safety pledges still don't solve

Sep 15, 2026

11:42pm UTC

Copy link
Share on X
Share on LinkedIn
Share on Instagram
Share via Facebook

As discussions of an AI slowdown escalate, leaders of AI's top labs may be aligned on where to start: third-party accountability.

Leaders from Anthropic, Google and OpenAI are in discussion about creating an AI industry standards body to test advanced AI models before deployment, CNN reported. However, these conversations were underway before the chaos of the past week incited new fervor in the debates around AI safety, and were instead spurred by Google DeepMind CEO Demis Hassabis' July essay that pitched a US-led standards body similar to the Financial Industry Regulatory Authority, according to CNN.

"The rapid progress we’re seeing in AI requires a new approach to testing frontier AI model capabilities that is dynamic, adaptable, and rigorous," Hassabis wrote in the essay.

But this isn't the only sign that the industry is looking for new ways to be held accountable:

  • OpenAI is backing the FRONTIER act, a bipartisan House proposal that would make it required for frontier AI labs to embed outside evaluators into their development processes to ensure model safety, according to a Tuesday Politico report.
  • Embedded evaluators were also part of Anthropic CEO Dario Amodei's pitch for pacing frontier development, calling for AI companies to give "employee-like access" to teams that can "verify adherence to safety practices and commitments."
  • And SpaceXAI CEO Elon Musk this week called for AI companies to work together to test each other's models before they are released to the public, specifically calling for OpenAI, Anthropic, Google, Meta and "three or four of the leading Chinese companies” to let rivals evaluate their models for safety.

However, third-party evaluators may only be one piece of the puzzle of a much larger framework necessary for keeping models in line. Miranda Bogen, director of the governance lab at the Center for Democracy and Technology, told The Deep View that while outside testers can spot the pitfalls in models before they go out to the public, they "can't change companies' behavior without a complementary suite of tools."

For instance, Bogen said, other necessary pieces include measurement standards, channels to communicate failed safety checks to relevant external stakeholders, making it mandatory to fix deficiencies, as well as a "clear allocation of responsibilities" to keep these evaluations from "ending up as a checkbox exercise."

"We've seen examples of the limitations of third-party assessments time and again across contexts, from the financial industry to aviation," Bogen told The Deep View. "The fresh energy around external evaluation is exciting and third-party evaluations are a critical piece of the puzzle, but we need to be realistic about what it will take for them to effectively reduce the many risks that AI systems pose."

Our Deeper View

As Bogen said, third-party evaluations are a great first step in spotting the flaws in powerful AI models before they reach the hands of users. But in order for this to actually be effective, these evaluators have to have leverage over these powerful companies. For instance, if a lab fails its safety standards evaluations, mandatory requirements should force that lab to either adjust its model to make it safe, or not release the model at all. Without that leverage, there is no consequence for a company not meeting these standards, or forgoing them entirely. Rather, evaluations would become a symbolic, good faith measure that doesn't actually do much to mitigate risk. The problem is that organizing this kind of effort generally takes public-private collaboration, and in the US, the Trump Administration has made it clear that it doesn't believe that AI presents the kind of risks that the industry is warning about. Additionally, given that this would require a global effort, getting Chinese labs to cooperate may be similarly difficult.