In Focus
Nvidia’s software solution is known as the Open Agent Safety Platform
The engineering solution is designed to address AI agent safety challenges
The new platform features two tools: Nvidia OpenShell and Nvidia Sentry
Nvidia has developed a software platform that enables AI labs to define safeguards for AI agents and keep them from escaping testing environments. Known as the Open Agent Safety Platform, the new platform includes Nvidia OpenShell, a system that runs on central processors to restrict what AI agents can access and do. The chipmaker also unveiled Nvidia Sentry, a tool that monitors agent activity while operating a network of chips.
Why Nvidia Built the Open Agent Safety Platform
In recent months, Nvidia CEO Jensen Huang has become an important voice in the AI safety discussions. According to Huang, most AI security concerns stem from engineering issues that can be addressed through computer science during product development.
“To deliver an agentic system in a safe way, you have to make sure that the sandbox around it is designed in a way that keeps the agent with minimal rights, whatever rights it has and it needs in order to do its job. No more rights than that. And then it can't do anything that it's not supposed to. Also, you have to monitor it,” Huang said during an interview with CNBC’s Squawk Box.
Nvidia’s AI agent safety platform launched amid calls for tighter AI testing safeguards. Recently, the EU President said the risks posed by self-improving models are increasingly becoming apparent and supported a slowdown in frontier AI.
Nvidia has played an important role in the growth of generative AI since ChatGPT launched about four years ago. Its GPUs provide the computing power needed to train large language models and support the AI applications and services developed by major cloud providers.
Can Nvidia’s Agent Safety Platform Prevent AI Agent Hacks?
Nvidia is offering an engineering solution designed to address the safety challenges posed by AI agents. The company said the new Nvidia AI agent safety platform could have prevented OpenAI’s AI agent hack into Hugging Face systems in July 2026.
“Each security incident is unique, and we have to look at all of them in detail. From what we know, Hugging Face reported over 17,000 agents attacking their infrastructure that went on for days and weeks,” Nvidia VP of enterprise AI, Justin Boitano, said as cited by CNBC.
Since July, other AI developers, including Anthropic, Meta, and Google, have disclosed incidents of AI agent hacking. Last week, Australia’s Prime Minister Anthony Albanese revealed that an OpenAI agent hacked into the country’s universal healthcare system in June 2026, exposing government systems to AI safety threats.
Nvidia’s AI Safety Platform Roll-Out Plan
Nvidia is open-sourcing parts of its AI agent sandboxing platform. The company considers the platform a reference design that partners like Cisco, Intel, ARM, Microsoft, Oracle, CoreWeave, and Dell can use to build and commercialize their own products. The chipmaker is working with Anthropic to integrate cloud-managed agents with OpenShell.
.webp&w=3840&q=75)

.webp&w=750&q=75)
.webp&w=750&q=75)
.webp&w=750&q=75)
