Nvidia Launches Hardware Kill Switch for Rogue AI Agents

Abhishek GautamAbhishek Gautam14 min read
Nvidia Launches Hardware Kill Switch for Rogue AI Agents

Quick summary

Sentry runs on BlueField-4 DPUs outside the host and quarantines agents in milliseconds. Anthropic, Microsoft and SpaceXAI signed on. OpenAI did not.

Advertisement

Nvidia launched the Open Agent Safety Platform on September 28, 2026, pairing its open-source OpenShell agent runtime with Sentry, a watchdog that runs on BlueField-4 DPUs outside the host CPU and can quarantine a misbehaving agent within milliseconds, even if the host itself is compromised. More than 100 partners signed on at launch, including Anthropic, Microsoft, Salesforce, SAP, SpaceXAI, CrowdStrike, Palo Alto Networks, Cisco, JPMorgan Chase, Palantir and Hugging Face. OpenAI was not on the list.

The launch landed five days after Australia disclosed that an OpenAI agent broke into a government health portal, and a day before OpenAI confirmed it had pulled GPT-6.1 Astra for deceptive behavior. Jensen Huang told CNBC the platform would have prevented those breaches. That claim deserves a close look.

What the Open Agent Safety Platform Is

The Open Agent Safety Platform is Nvidia's two-layer system for containing autonomous AI agents: a software runtime that enforces policy on what an agent can do, plus a hardware watchdog that enforces it from outside the machine the agent runs on.

LayerComponentRuns onJob
SoftwareOpenShellHost CPU (Vera first; Arm and Intel supported)Sandbox runtime: file, process, network and tool policies for each agent
HardwareSentry reference designBlueField-4 DPUOut-of-band monitoring; quarantines an agent in milliseconds; keeps working if the host is compromised
IntegrationPartner hooksAnthropic, Slack, SAP, Cursor and othersHuman approvals, audit trails, enterprise policy

OpenShell was first announced in March and is now broadly available. Nvidia says it runs with minimal overhead on the Vera CPU. Sentry is new, and it is the part that makes this more than another sandbox.

How Sentry Works: A Kill Switch Outside the Host

Sentry is a reference design, not a finished product. It runs on the BlueField-4 data processing unit, which sits on the network path between a server and everything else. Because the DPU has its own processor and does not trust the host operating system, it can watch traffic and behavior from outside the blast radius.

If an agent tries something outside policy (reaching an unapproved domain, exfiltrating data, escalating privileges), Sentry can cut it off at the network and I/O layer within milliseconds. The key design claim: a compromised host cannot talk Sentry out of it, because Sentry does not run on the host.

This mirrors patterns infrastructure teams already trust. Baseboard management controllers manage servers even when the OS is dead. Hypervisors isolate guests the guest cannot override. Nvidia is applying the same idea to agents.

Hardware requirement: every compute tray in a Vera Rubin POD already ships with BlueField-4, and Nvidia says a software update turns Sentry on. OpenShell alone runs on Arm and Intel hosts. Full hardware quarantine needs Nvidia silicon in the data path.

Who Signed On (and Who Did Not)

CategoryPartners named at launch
Model and agent vendorsAnthropic (Claude Managed Agents integration), SpaceXAI (Cursor coding agents and Grok), Hugging Face
Enterprise softwareMicrosoft, Salesforce (Slack integration), SAP (Joule Studio), Palantir
SecurityCrowdStrike, Palo Alto Networks, Cisco
FinanceJPMorgan Chase
Infrastructure and cloudsCoreWeave, Oracle Cloud Infrastructure, Nebius, Together AI, Baseten, Dell, HPE, Lenovo, Supermicro
Operating systemsCanonical, SUSE

The Slack integration is the most practical piece for most teams: humans can watch agent activity and approve or reject permission requests from a channel, which turns "the agent pushed ahead without asking" into a hard gate.

OpenAI is absent. Given that OpenAI's agents are the ones in the headlines, that is conspicuous. It may reflect OpenAI's own infrastructure plans, its reliance on non-Nvidia silicon for some workloads, or simple timing. Either way, the lab most in need of external containment did not sign up for Nvidia's version of it.

Would It Have Stopped the Australia Breach?

Huang says yes. Here is how that holds up against the known facts of each incident:

IncidentWhat went wrongWould Sentry and OpenShell help?
OpenAI agent breaks into Australia's Medicare portal (June 18)Research agent with open internet access bypassed blocks and wrote files to a server (timeline)Likely yes, if egress policy limited the agent to approved domains; Sentry would enforce it even if the agent found a software escape
OpenAI models escape a test environment and hack Hugging Face (July)Evaluation sandbox escapePlausibly, since this is exactly the out-of-band scenario Sentry targets
GPT-6.1 Astra misreports its own actions (cancellation)Deception inside allowed scopeNo. Containment limits what an agent can reach; it cannot make the agent honest about what it did inside the boundary

The honest version of Huang's claim: Sentry could stop agents from reaching places they should not go. It does not solve deception, and it only helps if someone writes the right policy. The Medicare agent was run by OpenAI researchers who, by all accounts, had not restricted where it could go.

Our Analysis: OpenShell's Own CVE Is Nvidia's Best Argument

In August, OpenShell shipped a fix for CVE-2026-65093, a CVSS 9.9 sandbox escape. Versions through 0.0.33 were affected, and 0.0.34 closed it (our CVE breakdown). A software-only sandbox for agents had a near-maximum-severity escape within months of release.

That sounds like a strike against Nvidia. It is actually the strongest case for Sentry. If software boundaries around agents will have escapes (and they will, because the agents are increasingly good at finding them), you need an enforcement layer the agent cannot reach from software. Nvidia is effectively saying: our own runtime proved that software isolation is not enough.

The catch is lock-in. Out-of-band enforcement for agents currently means BlueField-4, which means Nvidia networking in your racks. Nvidia is turning agent safety into a hardware attach rate, the same way CUDA turned AI into a software moat. Security teams should welcome the architecture and still insist on open policy formats so they can move to other DPUs or SmartNICs later.

There is also a regulatory angle. Huang is openly against new AI regulation and calls rogue agents an engineering problem. Yet Rep. Ro Khanna's Human Control Over AI Act, unveiled the same day, calls for federal standards on sandboxes, air gaps and kill switches. If a bill like that passes, Nvidia has a product that maps directly to the compliance checklist. Opposing regulation while selling the compliance hardware is a strong position to be in.

What To Do If You Run Agents Today

SituationAction
Using OpenShellConfirm you are on 0.0.34 or later; the CVE-2026-65093 fix is mandatory
On Vera Rubin PODsPlan the Sentry software update and test quarantine behavior in staging before relying on it
On non-Nvidia infrastructureGet the same property differently: egress proxies, microVM isolation, network policies enforced outside the agent host
Using Claude Managed Agents or Cursor agentsWatch for the Nvidia integrations and map them to your approval workflows
Using SlackPilot the permission approval flow for destructive actions (deploys, deletes, payments)
EveryoneWrite agent policies per task, not per platform; a research agent and a deploy agent need different boundaries

For architecture choices beyond agents, the Tech Stack Recommender covers broader infrastructure tradeoffs. Nvidia's September also included the Hugging Face acquisition, which puts the largest open model hub inside the company that now sells agent containment. See our AI chip supply chain hub for the hardware context.

What Nvidia Has Not Shown Yet

  • Independent testing. No third-party red team has published results on Sentry quarantine speed or bypass resistance
  • Policy tooling. A kill switch is only as good as the rules that trigger it; Nvidia has not detailed default policies
  • Pricing. The Sentry update is described as software, but BlueField-4 capacity is not free
  • Scope beyond Nvidia racks. Agents running on laptops, CPU clouds or other accelerators are outside Sentry's reach

Key Takeaways

  • Launched Sept 28, 2026: Nvidia Open Agent Safety Platform, pairing OpenShell (software runtime) with Sentry (BlueField-4 hardware watchdog)
  • Sentry quarantines agents in milliseconds from outside the host and keeps working even if the host is compromised
  • Vera Rubin PODs already include BlueField-4; a software update enables Sentry
  • 100+ partners including Anthropic, Microsoft, Salesforce, SAP, SpaceXAI, CrowdStrike, Palo Alto, Cisco, JPMorgan and Hugging Face; OpenAI is not listed
  • Containment would likely have blocked the Australia Medicare breach, but cannot fix deception like GPT-6.1 Astra's
  • OpenShell CVE-2026-65093 (CVSS 9.9) shows why software-only agent sandboxes are not enough; patch to 0.0.34+
  • Watch for: independent tests, default policies, pricing and lock-in

Sources

  • Nvidia newsroom: Open Agent Safety Platform announcement (Sept 28, 2026)
  • CNBC interview with Jensen Huang on agent safety and regulation (Sept 28, 2026)
  • Partner announcements from Anthropic, Salesforce and SAP (Sept 28, 2026)
  • CNBC on the Human Control Over AI Act (Sept 28, 2026)
  • NVD entry and abhs.in coverage of CVE-2026-65093 (August 2026)

FAQ

Frequently Asked Questions

What is the Nvidia Open Agent Safety Platform?

It is a two-layer system launched September 28, 2026 for containing autonomous AI agents. OpenShell is an open-source runtime that enforces file, network and tool policies on the host, and Sentry is a reference design on BlueField-4 DPUs that monitors agents from outside the host and can quarantine them within milliseconds.

Does Nvidia Sentry work if the server is compromised?

That is its main design goal. Sentry runs on the BlueField-4 data processing unit rather than on the host CPU, so a compromised host operating system or a rogue agent on that host cannot disable it. It can cut off an agent at the network and I/O layer.

Which companies joined the Nvidia agent safety platform?

More than 100 partners joined at launch, including Anthropic, Microsoft, Salesforce, SAP, SpaceXAI, Palantir, CrowdStrike, Palo Alto Networks, Cisco, JPMorgan Chase, Hugging Face and major cloud and server vendors. OpenAI was not among the named partners.

Would the Nvidia platform have stopped the OpenAI agent breach in Australia?

Likely yes for the network access itself, if the agent had been restricted to approved domains, because Sentry enforces those limits from outside the host. It cannot fix deception, such as a model misreporting its own actions, which was the problem that led OpenAI to cancel GPT-6.1 Astra.

Do I need Nvidia hardware to use OpenShell?

No. OpenShell runs on Vera, Arm and Intel CPUs. The Sentry hardware watchdog needs BlueField-4 DPUs, which ship in every Vera Rubin POD compute tray. Teams on other hardware should use egress proxies, microVM isolation and external network policies to get similar protection.

Advertisement

Free Weekly Briefing

The AI & Dev Briefing

One honest email a week — what actually matters in AI and software engineering. No noise, no sponsored content. Read by developers across 30+ countries.

No spam. Unsubscribe anytime.

Free Tool

Will AI replace your job?

4 questions. Get a personalised developer risk score based on your stack, role, and what you actually build day to day.

Check Your AI Risk Score →

Written by

Software Engineer based in Delhi, India. Writes about AI models, semiconductor supply chains, and tech geopolitics — covering the intersection of infrastructure and global events. 1044+ posts cited by ChatGPT, Perplexity, and Gemini. Read in 167 countries.