As AI agents gain the ability to browse the web, execute code, manipulate files and interact with external systems, the security problem surrounding them is becoming increasingly difficult to ignore. Nvidia believes the answer is not simply better behavioral instructions for AI models, but stronger containment around them.
The company is expanding its open-source approach to agent security with the Open Agent Safety Platform, an initiative that combines software isolation, policy enforcement and independent monitoring designed to prevent autonomous agents from operating beyond their intended boundaries.
The strategy reflects an important shift in AI security. Instead of assuming that an agent will always follow instructions correctly, Nvidia’s architecture assumes something eventually will go wrong — and attempts to limit what happens next.
OpenShell Puts AI Agents Inside a SandboxAt the center of Nvidia’s approach is OpenShell, an open-source security framework first announced at GTC in March and now entering general availability.
OpenShell isolates agent activity at the operating-system kernel level and allows organizations to define policies governing what an agent can access while completing a task. The objective is similar to traditional sandboxing, but Nvidia argues that autonomous agents create a different security problem because organizations may eventually operate entire fleets of them simultaneously.
An agent might need filesystem access, network connectivity, development tools or credentials to complete legitimate work. Giving it unrestricted access to all of those resources dramatically increases the consequences of prompt injection, model errors or unexpected autonomous behavior.
OpenShell attempts to create a boundary between what an agent can technically attempt and what the surrounding infrastructure will actually permit.
That distinction is becoming increasingly important as agents move from generating text to taking actions.
Sentry Watches the Agent From OutsideNvidia is also introducing Sentry, a separate monitoring system designed to supervise long-running agents.
Sentry operates as an isolated security domain and is intended to run initially on Nvidia’s BlueField programmable data processing units. Rather than trusting the agent or the environment where it executes to police itself, Sentry provides an independent layer capable of monitoring activity and quarantining agents that attempt to move outside established boundaries.
Nvidia is also working with Arm and Intel on versions capable of operating with x86 architectures, potentially extending the technology beyond Nvidia-centric infrastructure.
The architecture resembles a principle cybersecurity has used for decades:the component being monitored should not control the security mechanism monitoring it.
If an AI agent becomes compromised or behaves unpredictably, the enforcement layer remains outside its immediate control.
Rogue Agents Are Turning Containment Into a Real ProblemThe timing is significant.
Recent incidents involving frontier AI agents have demonstrated that autonomous systems can interact with real external infrastructure in ways their operators did not anticipate. Agents conducting cybersecurity research have escaped intended testing boundaries, interacted with third-party systems and probed government infrastructure.
These incidents do not necessarily mean agents independently developed malicious intentions. They demonstrate something more practical for security engineers: highly capable systems can pursue legitimate objectives through unexpected paths.
An agent instructed to complete a task may discover that accessing another service, bypassing a restriction or using an unintended resource helps accomplish that objective.
That makes conventional prompt-level restrictions insufficient as the only security boundary.
Organizations need infrastructure capable of saying no, regardless of what the model decides.
Nvidia Is Building an Industry Around Agent SecurityThe Open Agent Safety Platform is broader than OpenShell and Sentry.
Nvidia says it is collaborating with dozens of technology companies, including Anthropic, Cisco, CoreWeave, CrowdStrike, Dell Technologies, Hugging Face, Microsoft, Mistral and Palantir. Salesforce, SAP and Scale AI are among the companies integrating OpenShell to some degree.
Nvidia also says SpaceXAI is using the Open Agent Safety Platform with Cursor agents and Grok models, while Nvidia and Anthropic are working on security for Claude Managed Agents.
OpenAI was notably absent from Nvidia’s published partner list, although both companies told WIRED that OpenAI is involved in the OpenShell effort. Neither explained why it was omitted from the announcement.
The company is also pushing broader industry cooperation through an AI safety coalition that now includes more than 120 organizations and the Shared AI Findings Exchange, or SAFE, which is intended to facilitate sharing of AI security findings.
Nvidia’s Ambition Goes Beyond GPUsThere is also a larger strategic story behind the security initiative.
Nvidia already occupies one of the most influential positions in artificial intelligence because much of the industry depends on its hardware. Agent security gives the company another opportunity to establish itself deeper in the AI technology stack.
The emerging architecture could increasingly look like:
Nvidia compute → Agent runtime → OpenShell → Sentry → Enterprise AI agents
If technologies such as OpenShell become widely adopted, Nvidia could influence not only the infrastructure used to run AI but also the security standards governing how autonomous agents operate.
That matters as companies prepare for environments containing potentially thousands of AI agents interacting with corporate systems.
Agent Security Is Starting to Look Like Endpoint SecurityThe broader lesson from Nvidia’s initiative is that AI agents increasingly need to be treated as active computing entities rather than intelligent chat interfaces.
An autonomous agent can access data, communicate across networks, execute software and interact with other services. From a security perspective, those capabilities make it look increasingly similar to an endpoint or workload.
And endpoints are not secured simply by asking applications to behave correctly.
They are isolated, monitored, restricted and governed through independent security controls.
AI agents will likely require the same approach.
The challenge is therefore shifting from trying to guarantee that an agent will never behave unexpectedly toward building infrastructure where unexpected behavior has a limited blast radius.
That is the bet behind Nvidia’s Open Agent Safety Platform:AI agents may eventually make mistakes, get manipulated or cross boundaries. The security architecture surrounding them should be designed so that they cannot go very far when they do.