AI labs agree on testing, but enforcement is harder
September 16, 2026

Welcome back. Google is making a case that AI’s biggest story isn’t only about risk, but also about impact, from disease detection to disaster forecasting to language access. At Dreamforce, Salesforce offered a different glimpse of AI’s future: domain-specific models that let enterprises own more of their intelligence without becoming frontier labs themselves. And the industry’s push for third-party safety testing is gaining momentum, but testing alone won’t solve the safety issues without standards, enforcement, and consequences when the models fail to meet guidelines. And governments don't look likely to step in and play for their part, for now. —Jason Hiner
IN TODAY’S NEWSLETTER
1. What AI labs' safety pledges still don't solve
2. Why Salesforce may be AI's adult in the room
3. How Google made human impact an AI strategy
GOVERNANCE
AI labs agree on testing, but enforcement is harder
As discussions of an AI slowdown escalate, leaders of AI's top labs may be aligned on where to start: third-party accountability.
Leaders from Anthropic, Google and OpenAI are in discussion about creating an AI industry standards body to test advanced AI models before deployment, CNN reported. However, these conversations were underway before the chaos of the past week incited new fervor in the debates around AI safety, and were instead spurred by Google DeepMind CEO Demis Hassabis' July essay that pitched a US-led standards body similar to the Financial Industry Regulatory Authority, according to CNN.
"The rapid progress we’re seeing in AI requires a new approach to testing frontier AI model capabilities that is dynamic, adaptable, and rigorous," Hassabis wrote in the essay.
But this isn't the only sign that the industry is looking for new ways to be held accountable:
OpenAI is backing the FRONTIER act, a bipartisan House proposal that would make it required for frontier AI labs to embed outside evaluators into their development processes to ensure model safety, according to a Tuesday Politico report.
Embedded evaluators were also part of Anthropic CEO Dario Amodei's pitch for pacing frontier development, calling for AI companies to give "employee-like access" to teams that can "verify adherence to safety practices and commitments."
And SpaceXAI CEO Elon Musk this week called for AI companies to work together to test each other's models before they are released to the public, specifically calling for OpenAI, Anthropic, Google, Meta and "three or four of the leading Chinese companies” to let rivals evaluate their models for safety.
However, third-party evaluators may only be one piece of the puzzle of a much larger framework necessary for keeping models in line. Miranda Bogen, director of the governance lab at the Center for Democracy and Technology, told The Deep View that while outside testers can spot the pitfalls in models before they go out to the public, they "can't change companies' behavior without a complementary suite of tools."
For instance, Bogen said, other necessary pieces include measurement standards, channels to communicate failed safety checks to relevant external stakeholders, making it mandatory to fix deficiencies, as well as a "clear allocation of responsibilities" to keep these evaluations from "ending up as a checkbox exercise."
"We've seen examples of the limitations of third-party assessments time and again across contexts, from the financial industry to aviation," Bogen told The Deep View. "The fresh energy around external evaluation is exciting and third-party evaluations are a critical piece of the puzzle, but we need to be realistic about what it will take for them to effectively reduce the many risks that AI systems pose."

As Bogen said, third-party evaluations are a great first step in spotting the flaws in powerful AI models before they reach the hands of users. But in order for this to actually be effective, these evaluators have to have leverage over these powerful companies. For instance, if a lab fails its safety standards evaluations, mandatory requirements should force that lab to either adjust its model to make it safe, or not release the model at all. Without that leverage, there is no consequence for a company not meeting these standards, or forgoing them entirely. Rather, evaluations would become a symbolic, good faith measure that doesn't actually do much to mitigate risk. The problem is that organizing this kind of effort generally takes public-private collaboration, and in the US, the Trump Administration has made it clear that it doesn't believe that AI presents the kind of risks that the industry is warning about. Additionally, given that this would require a global effort, getting Chinese labs to cooperate may be similarly difficult.
TOGETHER WITH TABS
6-10 Days of Faster Collections? See How Automation Does It (On-Demand)
Manual AR workflows drain team capacity, delay cash, and create compliance gaps.
In this on-demand session, Tabs experts walk through the operational playbook, inlcuding:
Eliminating repetitive tasks
Achieving 15–30% DSO improvements
Building real ROI models you can implement immediately.
PRODUCTS
Why Salesforce may be AI's adult in the room
Salesforce is now making its own AI model. So is Crowdstrike. So is Thomson Reuters.
On Tuesday at its Dreamforce 2026 event in San Francisco, Salesforce announced Koa, its own domain-specific reasoning model that's purpose-built to enable agents to handle business tasks more effectively while keeping your data private. Koa has performed well in early benchmarks, including the LLM benchmark for CRM created by Salesforce AI Research to measure performance based on real-world enterprise tasks and used across the industry over the past couple years.
Salesforced reported, "Koa already matches or exceeds leading model performance on CRM actions with 3x fewer errors."
That tracks with the results of other domain-specific models, which typically reduce token costs and hallucinations because they are focused on a narrower set of expertise. They also tend to have increased performance for the same reason.
Salesforce built Koa, the name of a Hawaiian tree used to make canoes and ukuleles, by post-training Nvidia's open model, Nemotron 3 Super. Koa was specifically trained on synthetic data from almost three decades of business knowledge and was tuned to focus on knowledge work. As a result, Koa "is designed to support long-running agents that execute multiple tasks and complete complex outcomes," said Rohan Kumar, chief platform and engineering officer at the Dreamforce keynote on Tuesday.
The company doesn't see this as a vehicle for job or SaaS replacement, but as an enterprise empowerment tool that is more precise, more secure, and more tailored to the AI needs of companies that use Salesforce. One of the big promises of AI has always been that it will automate away grunt work and processes that don't add as much value. That's what Salesforce is trying to deliver here.
"The SaaS-pocalypse was not about the end of software, but it may be about the end of software that makes humans do all the work," CEO Marc Benioff said during the Tuesday keynote.
Salesforce also made a series of other AI announcements on Tuesday at Dreamforce, led by:
AIforce: This is a new interface that, instead of going to the traditional Salesforce UI, lets you ask questions, run complex queries, assign tasks, and create exactly the dashboards you need by simply interrogating your company's Salesforce instance directly.
Claudeforce: This basically turns Claude into a front-end for Salesforce and ships with 37 pre-built sales skills at launch that include functions such as deal review, pipeline hygiene, prospect research, account management, and more.
Headless 360: Salesforce is allowing its customers to have access to all of the elements of their platform through APIs, MCPs, plugins, and skills so that they can access their Salesforce data from the platforms of the choice without ever having to go to salesforce.com or use any of Salesforce's own tools, if that's what they prefer.

The next stage of enterprise AI is shaping up to be companies owning their own intelligence. As the models get smarter and smarter, intelligence is likely to encapsulate the greatest value inside an organization. Outsourcing that layer would mean losing control of your most important asset and potentially sending the most proprietary information about your business to another company, one that might also be serving your competitors. It's easy to see why companies like Salesforce, Crowdstrike, and Thomson Reuters have decided their need to build their own models. But it's also easy to see why they don't necessarily want to become frontier labs, when they can use open models like Nvidia Nemotron and use post-training to customize them and save a lot of time. And since Salesforce is a platform company, it will be interesting to see if it eventually helps other enterprises build their own AI models so that they can also capture more value and ROI from their investments in AI.
Disclaimer: Jason Hiner's travel to Dreamforce 2026 was paid for by Salesforce. The Deep View's coverage is editorially independent from the companies we cover.
TOGETHER WITH GRANOLA
This App Can Take Notes From Your Wrist
By now, most of us are familiar with AI meeting apps. They're the trick to never forgetting a single conversation or detail. But there's one catch: some of the most important conversations happen away from the laptop, in hallways, at coworkers' desks, in common spaces.
That's where Granola for Apple Watch comes in. If you already wear an Apple Watch, start a note from your wrist, stay present in the conversation, and finish with the context you need. The watch turns bright green while transcription is running, so everyone can see it's active, no guessing, no secrecy.
Try Granola for Apple Watch free for one month with code THE100OFF.
RESEARCH
How Google made human impact an AI strategy
Despite the heightened tension around AI's risks, the tech may actually be starting to live up to some of AI leaders' grandiose predictions.
In a blog post on Tuesday, Google announced that its tech now supports more than 300 languages, spoken by 7 billion people, representing around 86% of the global population.
This, however, comes on the heels of a number of significant breakthroughs, including unveiling and releasing AlphaGenome Atlas, a map of all 9 billion possible single letter genetic changes across the human genome, releasing WeatherNext 3, its most accurate global weather model yet, and creating the Planetary Prediction Engine to forecast and prepare for what it calls "planetary crises," such as disease outbreaks.
"We’re focusing our work in key areas that matter most: making disease detectable, treatable, and preventable, predicting natural disasters, expanding learning, and unlocking economic opportunities for more people," James Manyika, SVP of research, labs, technology and society at Google, wrote in the post.
In these areas, Google laid out several other initiatives to use AI for the benefit of humanity, including:
Using the tech to study breast cancer, tuberculosis and diabetes, as well as expanding wearables to detect things like cardiovascular diseases, insulin resistance and hypertension
Tracking and predicting extreme weather or natural disasters, such as monsoons, wildfires, earthquakes and floods, to prepare for and mitigate as much damage as possible
Using AI to democratize education through personalized learning and removing language barriers, and broadly creating better translation tools and more inclusive speech technology

There is a lot of doom and gloom around AI right now as the risks of the tech heighten anxiety around all of the ways it can be used for malice. And that risk isn't unwarranted, as we've recently seen warnings that frontier labs are having to thwart attempts to use the tech to create bioweapons. But we should always remember that AI is simply a very powerful tool that can be used for good or bad. Amid the current fear, the good is often being overshadowed. And while Google's initiatives certainly serve as good PR, both for the benefits of AI and for itself, they do serve as a reminder that, when it's in the right hands, AI can clearly be used to benefit humanity. And for a company like Google that's struggling to keep up with frontier labs in creating the most powerful models, focusing on ways to maximize the human benefits looks to be a solid strategy.
LINKS

OpenAI reportedly considers pre-IPO funding round at $1.2 trillion valuation
Agility unveils Digit 5, its latest humanoid that can work alongside people
Gates Foundation pledges $1 billion to expand global AI access
Google DeepMind staffer Bilal Chughtai resigns amid AI safety concerns
Manufacturing AI firm CADDi raises $114 million at $1.2 billion valuation
Zuckerberg says AI labs have "responsibility and incentive" to train safely

Gemini 3.8 Live and 3.8 Live Extended Thinking: Google released its most advanced live dialogue models yet, with upgrades in intelligence and parallel reasoning.
Odyssey-3: the startup's latest foundation world model, which can control robots, power humanoids, drive cars and more.
Salesforce in Claude: The tech firm announced a beta integration of its CRM features into Anthropic’s Claude.
Meta One: A subscription service in Meta's platforms that gives users access to more AI features.

A QUICK POLL BEFORE YOU GO
Do you have a generally positive, negative or neutral sentiment around the future of AI? |
The Deep View is written by Nat Rubio-Licht, Sabrina Ortiz, Jason Hiner, Faris Kojok and The Deep View crew. Please reply with any feedback.

Thanks for reading today’s edition of The Deep View! We’ll see you in the next one.

“I saw shoe prints in the sand near the fire.”
|
“[This image] has lots of sticks and debris very close to the fire, which is a huge hazard and most people who have camped or built one before know to clear the area.”
|


If you want to get in front of an audience of 750,000+ developers, business leaders and tech enthusiasts, get in touch with us here.












