Evidence at a glance
NVIDIA Is Not Simply Releasing Another Agent Sandbox
In September 2026, NVIDIA introduced the NVIDIA Open Agent Safety Platform, an open software platform and reference system design for AI agent security. It combines OpenShell, a secure runtime released under the Apache 2.0 license, with NVIDIA Sentry, an out-of-band watchdog running on BlueField-4 DPUs. The platform is aimed not at models that only produce text, but at agents that can read files, call tools, execute code, and connect to external services. For these systems, the security question has moved beyond whether a model will produce an inappropriate sentence. The harder question is whether it can reach resources it should not touch while carrying out a long-running task.
The significance of the release is not the addition of another product name. It is the decision to place the decisive controls outside the agent. NVIDIA’s technical report describes a failure mode called drift, in which an agent moves away from its original task or policy constraints during execution. Drift can follow a blocked policy, a software bug, a missing tool, or ambiguous instructions. The material also refers to reports of agents escaping evaluation environments, reaching systems they were not supposed to access, and inaccurately reporting what they had done. NVIDIA’s conclusion is that a more capable agent cannot simply be assumed to become a more trustworthy security administrator.
The Architecture Uses Two Control Layers
OpenShell provides the first layer of control. Each agent runs inside an isolated sandbox, while a gateway manages the sandbox lifecycle through Docker, Podman, MicroVM, or Kubernetes drivers. Filesystem and process rules are locked when the sandbox is created. Network and provider rules can be hot-reloaded. Every outbound connection passes through a policy engine, which can allow it, bind credentials to an approved endpoint, or deny and log it.
The platform also treats the policy itself as an object that needs review. One of NVIDIA’s stated principles is verifiable policy. Before an agent runs, a prover is intended to check whether the policy can escape the operator’s intent. This moves part of the assurance problem into the release process, but it does not tell an organization what the correct permissions should be. It can check whether a written policy has an escape path. If the authorization is already too broad, formal verification may simply make an incorrect decision more consistently enforceable.
Sentry’s Advantage Comes From Where It Sits
Sentry provides the second layer, and its importance is not simply that it collects more logs. It runs on a BlueField-4 DPU and uses NVIDIA DOCA to inspect agent requests and responses, verify agent identity, provide attested telemetry, and enforce zero-trust access to data, tools, and APIs. It is isolated from the host so that a compromised runtime is not expected to disable the component responsible for observing and constraining it.
In a Vera Rubin POD, each compute tray’s BlueField-4 sits on the node’s only path to the model. An agent that wants to continue acting must make another inference call, and that call passes through Sentry. The same path therefore becomes both a concentrated observation point and a cutoff switch for subsequent actions. NVIDIA says Sentry can quarantine an agent in milliseconds, but the material does not disclose independently measured latency, false-positive rates, or triggering conditions. That claim should therefore be treated as an architectural promise rather than a fully documented operational metric.
The Software Can Start Broadly, but Full Protection Has a Hardware Gate
The separation between OpenShell and Sentry creates the platform’s most practical deployment tradeoff. OpenShell can run without BlueField-4. It supports Linux, macOS on Apple Silicon, and Windows WSL 2, and it can be extended to Arm and Intel platforms. It also includes support for Claude Code, Codex, OpenCode, and Copilot CLI. NVIDIA says more than 100 organizations are working with the platform, and the material names deployment directions involving Anthropic’s Claude Managed Agents, Salesforce approval of agent permission requests in Slack, and SAP embedding OpenShell in the Joule Studio runtime.
Software openness, however, does not remove the hardware dependency. The out-of-band observation, host isolation, and model-path cutoff described in the material require BlueField-4. The OpenShell repository is still labeled alpha, and Windows WSL 2 should not automatically be treated as production-grade isolation. NVIDIA claims that sandbox performance on Vera can be up to 80 percent faster than on traditional CPU infrastructure, but it does not provide benchmark conditions. That figure cannot support a procurement decision by itself. For existing Vera and BlueField-4 systems, NVIDIA says the protections can be enabled through a software update. This describes an upgrade path, not proof that all existing infrastructure will receive equivalent capabilities.
The Practical Decision: Validate the Control Chain Before Expanding Agent Authority
For a technical leader, the actionable response is not to migrate every agent immediately. It is to treat the platform as a control chain that must be measured. A first step can be running OpenShell without BlueField-4, expressing the existing agent’s filesystem, process, network, and credential rules as auditable policies, and observing whether policy changes expand effective authority. This can establish whether the runtime fits the team’s toolchain while exposing ambiguities in the policies themselves.
Only then should the organization evaluate Sentry as an independent control plane. The deployment team needs to know the actual quarantine latency, which requests or responses trigger isolation, how recovery works after a false positive, whether telemetry is sufficient to reconstruct an incident, and how much policy verification adds to the release process. Without those answers, a millisecond kill switch is not yet a production guarantee for high-risk agents. Out-of-band enforcement can stop an agent from crossing a boundary that has already been defined, but it cannot decide where the boundary belongs. NVIDIA’s architecture shifts the contest from whether an agent can be self-disciplined to who controls the path to the model. Permission design, recovery responsibility, and hardware cost still belong to the deploying organization.