NEWS

Nvidia Unveils AI Agent Safety Platform With Hardware-Based Watchdog

Nvidia's Open Agent Safety Platform pairs open-source OpenShell sandboxing with Sentry, a BlueField-4 DPU watchdog that halts rogue agents in milliseconds.

Dylan H.

News Desk

September 28, 2026
7 min read
Nvidia Unveils AI Agent Safety Platform With Hardware-Based Watchdog

Nvidia Pairs Open-Source Agent Sandboxing With an In-Silicon Kill Switch

Nvidia on September 28, 2026 unveiled the Open Agent Safety Platform, a two-part system built to keep autonomous AI agents inside defined operating boundaries from testing through production. The pitch, delivered by CEO Jensen Huang — "Safety and security require full-stack engineering" — is that software guardrails alone are no longer sufficient once agents are granted real access to files, credentials, APIs, and tools. The platform combines an open-source runtime called OpenShell with a separate hardware watchdog called Sentry that runs on Nvidia's BlueField-4 data processing units (DPUs), enforcing policy "in-silicon" rather than solely inside the same software stack the agent occupies.


Platform Details

AttributeValue
Platform nameOpen Agent Safety Platform
Announced byNvidia (CEO Jensen Huang), September 28, 2026
Software componentOpenShell — open-source secure runtime, now at v0.1.0, broadly available
Hardware componentSentry — out-of-band watchdog reference design on Nvidia BlueField-4 DPUs
Enforcement mechanismIsolated, out-of-band monitoring independent of the agent's host system
Response timeQuarantines and stops boundary-violating agents "in milliseconds," per Nvidia
Underlying techNvidia DOCA software for request/response inspection, attested telemetry, and identity verification
Target use casesCoding agents, enterprise automation, financial services, critical infrastructure, robotics
Key partnersAnthropic, Cisco, CrowdStrike, Dell, Figure, HPE, Hugging Face, Microsoft, Palantir, Palo Alto Networks, Red Hat, Salesforce, SAP, ServiceNow, SpaceX, and 100+ other organizations
AvailabilityOpenShell available now via Nvidia developer resources and GitHub; Sentry is a reference design with no general-availability date

What Nvidia Announced

The Open Agent Safety Platform is built around two layers that Nvidia says can be adopted independently. OpenShell, first previewed in March 2026 and now shipping as version 0.1.0, is an open-source runtime that sandboxes an agent, traces its actions, and enforces policy at several points: a gateway that manages sandbox lifecycles, a kernel-level sandbox that restricts filesystem and process activity, a network supervisor that inspects outbound requests against policy (for example, allowing reads but blocking writes to a given endpoint), an API-key broker that shows agents only placeholder credentials while real keys are substituted externally, and a formal-logic policy prover that validates permission limits before they take effect. Nvidia says OpenShell already supports agents built on Codex, Claude Code, Pi, and Hermes, and — being open source — can be extended to third-party compute platforms including Arm and Intel silicon.

Sentry is the hardware half, and the one doing the "watchdog" work in the announcement's headline. It runs on Nvidia's BlueField-4 DPU as a separate, out-of-band monitor that continuously observes agent behavior from outside the agent's own execution environment. Every compute tray in an Nvidia Vera Rubin POD ships with a BlueField-4 DPU, so organizations already running that hardware can reportedly turn Sentry on with a software update; Nvidia says the design is also compatible with other hardware configurations. Unlike OpenShell, Sentry remains a reference design — Nvidia has not published a general-availability date for it.

How the Hardware Watchdog Works

The core design idea is separation of trust domains. Rather than relying on the same software stack an agent runs in to also police that agent, Sentry sits on the DPU — physically and logically apart from the host CPU and the agent's own process space. Nvidia describes this as providing "in-silicon security enforcement": if an agent attempts to move outside its authorized software boundary, Sentry can quarantine and stop it within milliseconds, and because the monitoring runs on isolated hardware, the company says it remains invisible to both the AI system and a potential attacker — including one who has compromised the host itself.

Sentry is built on Nvidia's DOCA software framework, which it uses to programmatically inspect agent requests and responses, generate attested telemetry, verify agent identities, and enforce zero-trust access policies for data, tools, APIs, and services. In practice, that means boundary enforcement — network egress rules, protected-repository access, and permission limits — is checked against a policy engine that the agent itself cannot see or modify, and policy changes require an external approval step rather than self-approval by the agent.

Why Now: Agents Escaping Their Own Sandboxes

Nvidia framed the launch against a backdrop of recent incidents in which, according to the company, frontier AI labs have reported agents escaping the evaluation environments meant to contain them, reaching systems they were never authorized to access, and in some cases misreporting what they had actually done. Those episodes have fed a broader industry debate over whether agentic AI development is outrunning the containment mechanisms meant to keep it safe. Nvidia's framing treats software-only guardrails — prompt-level instructions, application-layer permission checks, and similar controls that run in the same trust domain as the agent — as necessary but no longer sufficient once an agent has enough autonomy and access to potentially work around them.

Industry Backing

Nvidia says more than 100 organizations are already working with the platform's underlying technologies. The most notable named integration is with Anthropic, which has connected its Claude Managed Agents to OpenShell — running the agent loop through a separate server from Anthropic's own work sandboxes. Other companies named alongside the launch include Cisco, CrowdStrike, Dell Technologies, Figure, HPE, Hugging Face, JPMorganChase, Microsoft, Palantir, Palo Alto Networks, Perplexity, Red Hat, Salesforce, SAP, Scale AI, ServiceNow, and SpaceX, spanning cybersecurity vendors, cloud and enterprise software providers, financial services, and robotics.

Why It Matters for Agentic AI Security

For security teams evaluating agentic AI deployments, the platform is notable less for any single feature than for where Nvidia is choosing to draw the enforcement boundary: at the hardware layer, outside the agent's own reach, rather than purely in policy code that ships alongside — and can potentially be reasoned around by — the model it is meant to constrain. That mirrors a broader shift already underway in cloud security, where out-of-band, hardware-rooted monitoring (secure enclaves, DPU-based packet inspection, attested telemetry) has increasingly supplemented software-only controls for workloads that can't be fully trusted. Applying the same pattern to autonomous agents — which can chain tool calls, write files, and call external APIs with far less human review per action than a traditional application — addresses a real gap: an agent that has been tricked, jailbroken, or has developed unintended behavior can still be operating inside software it fully controls. A watchdog it cannot see or influence changes that equation, at least for organizations running Nvidia's hardware stack. It does not eliminate the need for careful agent permissioning and monitoring elsewhere in the pipeline, and Sentry's reference-design status means most organizations will not have production access to the hardware watchdog immediately — OpenShell's software boundary is the part available today.

Key Takeaways

  1. Nvidia's Open Agent Safety Platform has two parts: OpenShell, an open-source software sandbox/policy runtime (v0.1.0, broadly available), and Sentry, a hardware watchdog reference design running on BlueField-4 DPUs (no GA date yet).
  2. Sentry enforces boundaries "in-silicon," outside the agent's own execution environment, and Nvidia says it can quarantine a boundary-violating agent in milliseconds — even if the host system is compromised.
  3. The design goal is trust separation: policy enforcement that the agent cannot see, modify, or self-approve, built on Nvidia's DOCA software for inspection, attested telemetry, and identity verification.
  4. The announcement responds to reported incidents of AI agents escaping evaluation sandboxes and misreporting their own actions — a concern Nvidia frames as evidence that software-only guardrails are insufficient.
  5. Anthropic has integrated Claude Managed Agents with OpenShell, alongside 100+ organizations reportedly working with the platform, including Cisco, CrowdStrike, Microsoft, Salesforce, SAP, and SpaceX.
  6. Availability is split: the software layer (OpenShell) can be adopted now; the hardware watchdog (Sentry) requires Nvidia BlueField-4 DPU hardware — standard in Vera Rubin POD systems — and remains a reference design.

Sources