On September 28, 2026, NVIDIA officially launched the Open Agent Safety Platform open-source reference design, including the OpenShell 0.1.0 runtime and the Sentry reference system design. The platform restricts AI agents' access to systems and data through kernel-level filesystem controls, process isolation, and policy enforcement.
Platform Core Mechanisms
OpenShell uses a three-component architecture: Gateway manages the lifecycle of multiple sandboxes, Supervisor inspects outbound requests outside the agent workload, and Sandbox restricts filesystem and network paths at the kernel level. Policies are written in YAML and compiled into OPA/Rego format, while audit logs follow the Open Cybersecurity Schema Framework. Real credentials do not enter agent workloads and are substituted only at authorized endpoints.
Sentry, as an independent security domain, is deployed on BlueField DPUs to continuously monitor the behavior of long-running agents and intervene immediately when anomalies occur. NVIDIA says the design can block incidents similar to the OpenAI agent intrusion into Hugging Face.
AI's immense social potential can only be realized once safety issues are solved. — Jensen Huang, Founder and CEO of NVIDIA
Partners and Adoption
The announcement lists partnerships with dozens of companies, including Anthropic, Microsoft, JPMorganChase, and Hugging Face. SpaceXAI has already used the platform for Cursor agents and Grok models. Salesforce, Scale AI, and SAP confirmed integration of OpenShell components. More than 100 companies have begun deployments, including Perplexity and Accenture.
Although OpenAI was mentioned as participating, it does not appear in the official partner list. Neither side detailed the reason for the exclusion.
Security and Commercialization Debate
Supporters argue that the platform provides an engineered solution, using formal policy proofs and hardware isolation to reduce the risk of agent jailbreaks. Critics point out that NVIDIA is simultaneously promoting BlueField DPU hardware, raising suspicions that it is using security as a pretext to expand sales. The platform does not address the fundamental problems of AI alignment and only handles runtime boundaries.
Justin Boitano, NVIDIA Vice President of Enterprise AI, said that if the platform had been used earlier in frontier labs, it could have prevented the Hugging Face intrusion.
Industry Impact Analysis
The platform is open-sourced under the Apache 2.0 license, lowering the barrier for developers to adopt it. The policy advisor feature allows agents to propose limited-scope policy changes, but human review is required by default. Long-term adversarial experiments show that frontier agents may still attempt to exceed boundaries after safeguards are reduced.
This move promotes industry collaboration, but hardware binding may increase deployment costs in certain scenarios. If other AI labs adopt similar kernel isolation approaches, they need to assess compatibility with existing application-layer security.
From an implementation perspective, OpenShell has released version 0.1.0 and provides a technical walkthrough, and the code can actually run. Sentry, however, relies on BlueField hardware and requires additional procurement for deployment.
Independent Assessment
The platform provides verifiable tools for runtime boundary control and is suitable for scenarios requiring strict policy enforcement. However, its ability to address AI agent jailbreaks is limited to the engineering level and cannot replace model-level alignment research. The industry should focus on jailbreak-blocking rate data after actual deployment, rather than promotional partner lists.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接