A policy written into a prompt is not a security boundary. When an agent can reach credentials, files and APIs, telling it not to cross a line is materially different from enforcing that line outside the system. The distinction becomes more consequential as agents gain the ability to run code and continue working for hours.
On September 28, NVIDIA introduced the Open Agent Safety Platform, a combination of software and infrastructure architecture designed to enforce controls outside the agent workload. It brings together OpenShell, an open runtime that restricts access and records policy decisions, and Sentry, a reference design for independent monitoring on dedicated hardware.
The meaningful change is not another instruction for a model to follow. It is removing the agent's authority to decide whether a rule will be honored. From the MnzAI Labs perspective, that places agent operations closer to established infrastructure security practices: least privilege, separation of duties and auditable decisions.
Two layers at different stages
The components serve separate purposes:
- OpenShell: places agents in sandboxes and controls filesystem, process, network, service and credential access. An external supervisor inspects requests, so a read can be allowed while a write to the same API is denied. Policy decisions are recorded in an OCSF-compatible audit trail.
- Sentry: acts as an out-of-band observer. In NVIDIA's design, it runs on BlueField-4 DPUs, watches agent behavior and can quarantine an agent after a policy violation is detected.
That separation matters when assessing availability. OpenShell 0.1.0 is open source and available, with declared support for environments including Docker and Kubernetes and an extensible path across processor architectures. Sentry, by contrast, is presented as a reference system design tied to specific hardware. The announcement should not be read as proof that every layer can be deployed today in the same way on any infrastructure.
The operational question shifts from “did the agent understand the rule?” to “which external layer can enforce it?”
Evidence still to come
NVIDIA says more than 100 organizations are involved in the ecosystem and describes integrations with companies including Anthropic, SAP and Salesforce. Those examples show commercial interest, but they remain vendor claims. The published material provides no independent benchmark for comparing security effectiveness, performance overhead or operating cost with alternative approaches.
Open source does not mean secure by default, either. Incomplete policies, overprivileged credentials and poorly configured integrations can still create risky paths. OpenShell's formal policy prover can analyze modeled permissions, but its conclusions depend on what the model actually captures.
NVIDIA itself notes that described products and features may be at different stages and remain subject to availability. A practical evaluation should therefore begin with what teams can inspect: run an agent in a sandbox, attempt a prohibited write, and verify whether the decision appears clearly in the audit trail.
This release does not settle agent security. It does reinforce a useful design principle: agents with meaningful system access need boundaries that remain in force when the model makes a mistake, persists or changes its strategy.
Sources: NVIDIA's Open Agent Safety Platform announcement and OpenShell 0.1.0 technical overview, both published September 28, 2026.
