Nvidia Wants Hardware to Stop Rogue AI Agents — Here’s How Its New Safety System Works

Written By
Ai Cerrudo
Ai Cerrudo
Sep 30, 2026
5 minute read
Nvidia’s Open Agent Safety Platform combines software controls with hardware-isolated monitoring designed to restrict AI agents that move beyond defined boundaries.

Nvidia’s Open Agent Safety Platform combines software controls with hardware-isolated monitoring designed to restrict AI agents that move beyond defined boundaries. Image: Generated with ChatGPT.

eWeek content and product recommendations are editorially independent. We may make money when you click on links to our partners. Learn More

AI agents can now use tools, access files, call APIs, and work autonomously for extended periods. Nvidia wants some of the controls governing those agents to sit somewhere the AI cannot directly change: outside its own software environment.

The company has launched the Nvidia Open Agent Safety Platform, which combines software-based sandboxing with an optional hardware-isolated monitoring system designed to restrict agent behavior. Nvidia says the hardware layer can quarantine an agent in milliseconds when deployed with its BlueField-4 data processing units.

The approach applies familiar security principles — including isolation, least privilege, identity verification, and policy enforcement — to autonomous AI systems that may operate for long periods without constant human supervision.

OpenShell puts agents inside a controlled runtime

The first layer is Nvidia OpenShell, an Apache 2.0-licensed open-source runtime that executes agents inside isolated sandboxes.

Operators define which files, networks, processes, tools, and credentials an agent can access. According to Nvidia, OpenShell then enforces those restrictions with kernel-level controls covering file, process, and network access.

That matters because useful agents often need broad permissions.

A coding agent, for example, may need to read files, install packages, call external services, and use credentials. OpenShell is designed to separate what the agent is capable of doing from what it is actually permitted to access.

Nvidia also says credentials can remain outside the agent sandbox and be added only to requests sent to approved endpoints, reducing the need to expose them directly to the agent.

Advertisement

BlueField adds an independent monitoring layer

Nvidia's second layer is Sentry, a reference system design that runs on BlueField-4 DPUs.

The company describes Sentry as an out-of-band monitoring system because it operates separately from the AI agent and host software. Using Nvidia's DOCA software framework, Sentry can inspect agent activity, verify identities, monitor access to data and tools, and enforce security policies.

Because BlueField operates in an isolated hardware domain, Nvidia says its monitoring can continue even if software running on the host is compromised. The agent cannot directly alter the watchdog running outside its environment.

Nvidia also says Sentry can quarantine an agent in milliseconds if it detects behavior that falls outside defined policies.

That is a company performance claim. The platform is new, and independent evidence showing how consistently that response works across large-scale production deployments is not yet available.

OpenShell can run without Nvidia hardware

The software layer does not require BlueField-4.

Nvidia says OpenShell can run on local systems, on-premises infrastructure, cloud environments, and Kubernetes without Sentry. The runtime is optimized for Nvidia's Vera CPU but can also be extended to third-party compute platforms, including Arm and Intel systems.

BlueField-4 adds the separate hardware-isolated enforcement layer for organizations that want an additional control boundary.

OpenShell also supports both open and closed models and can work with existing agent tools. Nvidia lists Claude Code, Codex, OpenCode, GitHub Copilot CLI, and OpenClaw among the agents currently supported.

The company says Cadence, Slack, and Gecko Robotics are adopting OpenShell for uses including chip design, enterprise automation, and robotics.

Nvidia's broader Open Agent Safety Platform also has support from companies including Anthropic, Cisco, CrowdStrike, Dell Technologies, Microsoft, Palantir, Palo Alto Networks, Salesforce, SAP, ServiceNow, and SpaceXAI.

Advertisement

How Nvidia's AI agent safety layers work

Safety layer

Where it operates

What it does

Hardware required?

Model safeguardsAI model/application layerInfluences what an agent attempts to doNo
Nvidia OpenShellIsolated runtimeRestricts access to files, networks, processes, tools, and credentialsNo
Nvidia SentrySeparate monitoring and enforcement layerMonitors agent activity and enforces defined security policiesYes, for this layer
BlueField-4 DPUHardware-isolated domainHosts Sentry separately from the agent and host softwareYes

Nvidia is moving some safety controls outside the model

Nvidia says recent cases involving agents circumventing application-layer controls helped motivate the platform.

Its technical team describes a broader problem it calls agent "drift," in which an AI system deviates from its intended task after encountering blocked actions, missing tools, bugs, ambiguous instructions, or repeated failures.

Rather than relying only on model-level safeguards, Nvidia is adding enforcement outside the model itself.

That does not mean OpenShell or Sentry can guarantee that an AI agent will always behave correctly. Nvidia presents the platform as an additional layer of containment and monitoring, not a replacement for model safeguards, application security, or human oversight.

CEO Jensen Huang described the strategy as part of Nvidia's broader "full-stack engineering" approach to AI safety.

What eWeek Found: Nvidia is moving some AI agent enforcement outside the agent itself

The key distinction in Nvidia's architecture is where authority over an AI agent sits.

OpenShell defines restrictions outside the agent process, while Sentry can move another layer of monitoring and enforcement into an isolated hardware domain. Nvidia's documentation distinguishes these runtime controls from model safeguards: safeguards influence what an agent attempts to do, while runtime controls determine what it is permitted to access.

The architecture applies established cybersecurity practices — including sandboxing, least privilege, isolation, and zero-trust access — to autonomous agents that may interact with tools and infrastructure without continuous human involvement.

There are important limits to what can be concluded. Nvidia says Sentry can quarantine agents in milliseconds, but the platform was only recently announced, and its effectiveness across large production environments has not yet been independently established.

Advertisement

The system also does not remove the need to evaluate the underlying AI model, application permissions, security architecture, or human approval process.

Instead, Nvidia is adding another security boundary: if an agent ignores or circumvents controls inside its software environment, infrastructure outside that environment may still restrict what it can access.

Want to learn more AI tips, tricks, and prompting techniques? eWeek readers get free 7-day access to The Neuron Academy, our practical learning platform designed to help professionals use AI more confidently at work. Browse all lessons →

Let us teach you How to Talk to AI for free! Try our six-minute course at The Neuron Academy and learn a few simple ways to write better prompts and get more useful results from AI, or browse our other AI course for free for seven days. Check out all the lessons here →

Ai Cerrudo

Ai Cerrudo is a writer and editor with a decade of experience in media and publishing. Beginning as a journalist in the Philippines, Ai has covered a diverse spectrum of beats, including politics, healthcare, business, and interactive media/gaming. Blending analytical rigor with engaging storytelling, she now works as an editor across technology and AI media

eWeek Logo

eWeek has the latest technology news and analysis, buying guides, and product reviews for IT professionals and technology buyers. The site's focus is on innovative solutions and covering in-depth technical content. eWeek stays on the cutting edge of technology news and IT trends through interviews and expert analysis. Gain insight from top innovators and thought leaders in the fields of IT, business, enterprise software, startups, and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.