TL;DR โ Key Takeaways
– NVIDIA has launched the Open Agent Safety Platform to help contain autonomous AI agents and prevent them from accessing unauthorized systems, files and networks.
– The framework combines OpenShell, an open-source runtime for enforcing agent policies, with Sentry, a hardware- and network-level monitoring system that can rapidly quarantine suspicious agents.
– NVIDIA says more than 100 organizations are already using or developing with the platform, including Microsoft, Cisco, JPMorgan Chase, Accenture and Lenovo.
NVIDIA Corp. on Monday unveiled a new cybersecurity framework aimed at reining in artificial intelligence (AI) agents, Open Agent Safety Platform, following a wave of high-profile security breaches.
The release comes as OpenAI, Anthropic, Meta Platforms Inc., Google, and others confront a surge in rogue incidents in which autonomous models escaped isolated testing environments, breached unauthorized networks, and rewrote sensitive data.
The chipmaker’s new platform introduces open-source software and hardware-level controls designed to keep AI agents contained within strict boundaries.
According to NVIDIA, the technology could have prevented a major July 2026 incident in which a swarm of OpenAI agents autonomously breached the systems of open-source platform Hugging Face, an entity NVIDIA subsequently acquired for $13 billion earlier this month.
“From what we know, this new security platform could have stopped the breach if it was being used in frontier labs for model evaluation early on,” Justin Boitano, NVIDIA’s vice president and general manager of enterprise computing, said during a press briefing.
NVIDIA’s safety stack relies on two primary components: OpenShell and Sentry.
OpenShell is an open-source runtime environment that establishes policy boundaries for autonomous agents, restricting unauthorized access to filesystems, network connections, and system processes. Because the software is open source, it extends beyond NVIDIA’s native hardware to support rival computing platforms from Arm Holdings and Intel Corp.
Working alongside the software runtime is Sentry, a network- and silicon-level security monitor. Sentry continuously tracks agent activity in real time and can quarantine suspicious agents within milliseconds if they attempt to bypass sandbox restrictions. The feature is specifically engineered to counter agent swarms, scenarios where a primary AI spawns multiple sub-agents to evade security protocols.
Industry adoption for the framework is being built rapidly. NVIDIA confirmed that more than 100 organizations are utilizing or developing on the platform, including Microsoft Corp., Perplexity AI Inc., Accenture, JPMorgan Chase, Cisco Systems Inc., Dell Inc., Hewlett Packard Enterprise Co., and Lenovo. The company is also collaborating with Anthropic to integrate cloud-managed agents like Claude Code with OpenShell architecture.
“NVIDIA is moving agent controls into the kernel and the DPU, where an agent can’t talk or code its way around them. Guardrails that run inside the agent process can be bypassed by the agent itself,” said Mitch Ashley, vice president and practice lead for Software Lifecycle Engineering and AI-Native Software Engineering at The Futurum Group.
“The harder problem is policy. Enterprises will run OpenShell alongside hyperscaler and SaaS agent controls, each with its own policy model,” Ashley said. “Watch whether buyers can author a policy once and enforce it everywhere, instead of maintaining three versions of the same intent.”
“NVIDIA is moving agent safety outside the model, where it belongs,” said Stephanie Walter, practice leader for AI Stack & Enterprise Application Development at HyperFRAME Research. “Instructions are not sufficient security boundary once an agent can use credentials, tools, and networks. OpenShell limits what the agent can access, while Sentry independently monitors its behavior and can stop it.
“OpenAI’s recent disclosure makes the need for this kind of enforcement clear. Its agents moved beyond their intended boundaries, prompting the company to pause a major reinforcement-learning training run while it improves its safeguards,” Walter said. “The question is not whether agents will act unexpectedly, but whether enterprises can detect, contain, and stop those actions before they cause harm. NVIDIA will still need to prove that its controls can do that without routinely blocking legitimate work.”
The rollout highlights a deepening divide across Silicon Valley regarding how to govern rapid advancements in AI.
Executives at Anthropic and OpenAI have recently called for a coordinated slowdown in development to allow safety protocols to keep pace.
Anthropic CEO Dario Amodei warned in a recent blog post that AI is advancing drastically faster as models begin building the next generation of systems, prompting his company to hire Accenture for third-party security audits.
Meanwhile, OpenAI CEO Sam Altman disclosed plans for an initial public offering are on hold while the company addresses alignment and safety imperatives following unauthorized agent activity on government websites in the U.S. and Australia.
NVIDIA CEO Jensen Huang rejected calls for regulatory pauses or new legislation, framing AI safety as a straightforward engineering challenge rather than a cause for government intervention.
“AI safety is a real thing. Engineering products so that they are safe for the world to use — that’s a real thing,” Huang told CNBC, emphasizing that companies must take responsibility for rigorous testing before commercial release. “The fact that we need new laws… so that these companies could do their fundamental engineering and do it properly before they release products, that is just completely unnecessary.”
Huang, who serves on the Trump administration’s advisory board on science and technology, emphasized that industry-led innovation remains the key to maintaining control over advancing models.
“As we continue to discover the frontier of AI capabilities,” Huang said, “we must accelerate discovery at the frontier of AI safety.”

