Nvidia is pitching itself as AI's safety steward

•

Sep 28, 2026

•

10:44pm UTC

Copy link
Share on X
Share on LinkedIn
Share on Instagram
Share via Facebook
T

esting and experimenting with AI agents is extremely useful, but the risks have been exposed recently, as many have escaped their environments and breached third parties. Nvidia thinks it has a solution.

On Monday, Nvidia announced the Open Agent Safety platform, which the company describes as an open software platform and reference system design that can strengthen AI agent security from testing to deployment. Simply put, it helps developers build safeguards to prevent agents from escaping containment, ultimately addressing the underlying cause of the recent agent security incidents, which Nvidia identifies as agents circumventing security controls to complete their assigned tasks.

Over 100 partners have joined the platform, including heavyweights such as Anthropic, SpaceX, Microsoft, IBM, and Perplexity. Meanwhile, OpenAI, Google, Meta, and Amazon are notably absent from the new coalition.

With the new platform, organizations can control the entire AI agent stack, from the software that runs the agents to the hardware that powers their work, and even the robotic systems that take physical action in the real world. This is enabled by bringing together Nvidia OpenShell and Sentry:

  • NVIDIA OpenShell: the software that sets boundaries for how agents running on CPUs execute tasks across both open and closed models. It shouldn't compromise speed or delivery.
  • NVIDIA Sentry: Monitors agents' behavior, providing in-silicon security enforcement that can stop and quarantine an agent attempting to move outside its boundary within milliseconds.

CNBC reports that in a call, Nvidia told reporters that this could have prevented OpenAI's Hugging Face incident in July, in which an agent, driven by a combination of OpenAI models, compromised Hugging Face's infrastructure. Since then, there have been many other similar instances reported, including the first known AI-attempted hack of a government entity.

In addition to the partnerships mentioned above, Anthropic and Nvidia took their collaboration a step further by developing Claude Managed Agents, which run the agent loop on a separate, sealed-off server from the sandboxes where the work takes place. Nvidia's OpenShell software and BlueField chips let enterprises set strict rules about what the AI can access inside that workspace.

Our Deeper View

The most interesting point is that, as Jensen Huang noted in his post on X, this isn't meant to be a single product but an ecosystem. Up until now, Nvidia hasn't positioned itself as a leading voice on AI governance, as Anthropic, OpenAI, and Google have been trying to do. But in its role as self-appointed steward of the current AI boom, this looks like a clear step in that direction and shows that it's concerned the safety narrative could derail AI progress. Through this release, Nvidia acknowledges the same risks others have raised and offers its own path forward. This release also continues to highlight how AI companies are being left to build safety measures on their own, since there's no overseeing body to set and enforce them.