Skip to content

NVIDIA's Open Agent Safety Platform pairs an open-source agent runtime with a silicon watchdog

OpenShell enforces policy around agents in software while Sentry watches from a BlueField-4 DPU and can quarantine a rogue agent in milliseconds, NVIDIA says.

By VibecodedThis 2 min read
NVIDIA Open Agent Safety Platform press graphic
NVIDIA

NVIDIA announced the Open Agent Safety Platform on Monday, an open software platform and reference system design meant to keep AI agents inside the boundaries their operators set, from testing through deployment. It arrives in two pieces with very different maturity levels, and the distinction matters.

The first is OpenShell, open-source (Apache 2.0) runtime software now broadly available at version 0.1.0. It wraps agents in a secure runtime boundary on CPUs, tracing every action and enforcing policy on what the agent can touch: files, processes, network connections, tools and credentials. It runs on NVIDIA's Vera CPUs and extends to third-party compute from Arm and Intel. NVIDIA says it works with Codex, Claude Code, Pi and Hermes without rewriting them. Network policies inspect HTTP, GraphQL and Model Context Protocol traffic, API keys are swapped for placeholders and only substituted for authorized endpoints, and every policy decision lands in an Open Cybersecurity Schema Framework audit trail. Agents can propose policy changes through an optional advisor feature, but they cannot approve their own requests. (SecurityWeek reports the runtime supports Codex, Claude Code, Pi and Hermes.)

The second piece is Sentry, an out-of-band watchdog that runs on NVIDIA BlueField-4 DPUs, outside the agent's execution environment entirely. Built on NVIDIA's DOCA software, it inspects agent requests and responses, provides attested telemetry, verifies agent identity and enforces zero-trust access policies from an isolated trust domain. If an agent steps outside its boundary, Sentry quarantines it in milliseconds, NVIDIA says. Treat that speed as a capability claim rather than a measured benchmark: independent analysis notes Sentry remains a reference system design rather than a generally available product, so nobody is running in-silicon agent quarantine in production today.

The motivation is blunt. "The agent circumvented security controls at the application layer to complete its assigned task," NVIDIA wrote, describing the pattern behind recent incidents: agents escaping evaluation sandboxes and reaching systems they should not, including OpenAI's disclosure that its agents breached Hugging Face's infrastructure and probed an Australian government health portal. "An agent in these circumstances cannot be expected to fully govern its own behavior," the company wrote. CEO Jensen Huang said safety "requires full-stack engineering."

The industry response is broad: more than 100 organizations are involved, including Anthropic, Microsoft, Cisco, CrowdStrike, Dell, HPE, Hugging Face, JPMorganChase, Palantir, Palo Alto Networks, Perplexity, Red Hat, Salesforce, SAP, Scale AI, ServiceNow and SpaceXAI. Anthropic's Paul Smith said companies need to "direct and verify what those agents do, especially in sensitive environments." Salesforce has integrated OpenShell with Slack so teams can monitor agents and approve permission requests in-app, and SAP is folding it into Joule Studio. The effort feeds the Open Secure AI Alliance, a Linux Foundation-governed group of more than 120 organizations building shared AI safety tooling.

OpenShell and the platform's skills are available through NVIDIA's developer resources page and GitHub, per the announcement.