On September 28, 2026, NVIDIA released the Open Agent Safety Platform, pitching an open software stack and reference design to govern AI agents from runtime up through hardware and robotics. According to CyberScoop, more than 100 organizations, including Anthropic, Arm, Microsoft, SpaceXAI, Palantir, and JPMorgan Chase, have committed to adopt pieces of the platform to harden products that rely on autonomous or semi-autonomous agents.
The company’s message is blunt: model alignment alone is not enough. One expert told CyberScoop the launch reflects a wider industry recognition that after years of training models to behave, teams now need stronger outside controls. NVIDIA’s approach centers on tougher sandboxes and deeper monitoring to curb “rogue” agent actions. The platform debuts a pair of open tools, including OpenShell, which the company says provides sandboxed, policy-driven execution for agents in test environments. CyberScoop reports OpenShell is built on Apache 2.0–licensed software and lets operators specify which files, networks, and system resources an agent can touch.
What the Open Agent Safety Platform actually offers
NVIDIA describes the Open Agent Safety Platform as a full-stack governance layer across software and hardware, compute, and even robotics systems that run agents, CyberScoop notes. The emphasis is on two pillars: keeping agents boxed in during sensitive tasks and watching what they do with enough fidelity to catch misuse or drift before it spreads into production.
OpenShell sits at the center of that plan. By constraining runtime access and requiring explicit policies, the tool turns agent permissions into code rather than trust. That gives security teams a practical way to enforce guardrails that do not depend on a single model’s internal alignment. CyberScoop says the platform also introduces a second tool focused on containment and restriction; taken together, the pair is meant to shrink the blast radius of agentic failures and simplify post-incident forensics.
The company’s framing matters because more enterprises are wiring agents to take real actions: file operations, API calls, code execution, and control of physical devices. Sandboxing and monitoring sound like old ideas, but applied to agents they address fresh risks: prompt-induced tool abuse, indirect injection via data sources, and lateral movement across integrated services. These are attack paths cataloged by efforts such as MITRE ATLAS, which has been mapping adversary techniques against AI-enabled systems.
Why agent security is moving outside the model
Security controls are following the architectural center of gravity. Enterprises are shifting from static assistants to agent-driven execution that plans, acts, and learns across workflows. Forrester analysts Charlie Dai and Meng Liu argue that the next wave of competition will hinge on who can operationalize agents at scale, with dedicated runtimes, governance, and observability, not just bigger models. In their analysis of Alibaba’s Apsara Conference on September 28, 2026, they describe an “Agent-Native Cloud” push with components like Agent Sandbox, AgentCore, and an Agent Security Center (Forrester).
NVIDIA’s release slots cleanly into that direction. If agents are the new computing abstraction, then guardrails need to exist where actions occur: at the agent runtime and the system boundaries it touches. That means hardened shells, execution policies, identity-scoped permissions, and traceable logs that can be inspected and replayed. It also means vendors must publish reference designs that customers and partners can extend, since every enterprise stitches agents into different stacks.
There’s a regulatory current as well. Governance frameworks like the NIST AI Risk Management Framework call for measurable controls tied to context and impact. Agents bring new context and higher impact. Outside-the-model safeguards—sandboxes, circuit breakers, signed tool integrations, and audit trails—map to requirements that regulators and risk teams already understand from software security. NVIDIA’s open approach signals that these controls should be community-tested, not hidden behind proprietary walls.
How enterprises should evaluate NVIDIA’s agent security tools
For CISOs and platform teams, the question is less “Does it align?” and more “Does it constrain?” A practical evaluation plan for NVIDIA’s agent security stack should start with where agents already act today and expand outward:
- Inventory agent tool use. List system calls, APIs, file paths, and network targets an agent can reach. Aim to turn each into a policy in OpenShell or equivalent.
- Adopt least-privilege for actions. Bind tool access to the smallest necessary scopes, and rotate credentials as part of the agent’s lifecycle.
- Red-team the runtime. Use known agentic attack patterns—prompt injection via retrieved content, tool call escalation, data exfiltration—to validate that sandbox policies hold under stress. MITRE ATLAS can inform scenarios.
- Treat monitoring as evidence. Ensure logs capture inputs, tool selections, and outputs with timestamps and identities. If an incident occurs, you need replayable traces for root cause and regulatory response.
Buyers should also test how easily the Open Agent Safety Platform integrates with existing compute, container, and CI/CD stacks. If the controls require a parallel universe, teams will bypass them. If they drop into current pipelines with minimal friction, they will get used. The fact that more than 100 firms have signaled support, as reported by CyberScoop, suggests vendors see customer demand for controls that work across heterogeneous environments.
What this means for policy and open-source ecosystems
Open tooling can change the tempo of safety work. Because OpenShell is Apache 2.0–licensed, per CyberScoop, enterprises and researchers can inspect, modify, and redistribute changes under permissive terms (Apache License 2.0). That encourages red teams to publish test harnesses, and encourages cloud providers to wire policies into their managed runtimes. It also raises the floor: if a common, vetted sandbox exists, customers can expect baseline controls across vendors.
Policy debates are circling the same point. Lawmakers want clearer guardrails without freezing innovation. Open, auditable controls at the agent runtime give both sides a path forward. They make it plausible to certify that a system can only perform approved actions in defined contexts, and they let independent researchers validate those claims.
There are trade-offs. Overly tight sandboxes can hinder agent usefulness. Under-specified policies invite abuse. That is where reference designs help. By publishing how governance spans software, hardware, and robotics, NVIDIA is telling integrators where the seams are, and where to place tripwires. It also puts pressure on model providers and app builders to meet the same bar.
The cadence of agent incidents will test this strategy. If open sandboxes and stronger monitoring reduce the number and scope of mishaps—accidental deletions, unwanted purchases, code injection, or unauthorized device control—demand will grow. If they overpromise, teams will retreat to narrower assistants until better controls arrive.
NVIDIA’s bet is that the industry is ready to standardize outside-the-model safeguards. The Open Agent Safety Platform gives security teams something concrete to test, extend, and hold vendors to—and that is what agent-native operations will require. For more on this, see reuters.com and bloomberg.com and nytimes.com.
Related reading: Hugging Face • Fine-Tuning • Open Source AI
