NVIDIA released a set of software security tools for AI agents called OpenShell on September 28 local time, claiming it can prevent incidents like the AI attack on Hugging Face. Hugging Face is the world's largest open-source model hosting platform, which was acquired by NVIDIA for 12.93 billion USD this September.

Justin Boitano, Vice President and General Manager of NVIDIA Enterprise Computing, stated that if frontier laboratories had used this new security platform during early model evaluations, they could have prevented the previous incident where OpenAI infiltrated Hugging Face.

OpenShell uses hardware features on NVIDIA central processing unit (CPU) chips to contain agents. It detects whether agents are attempting to bypass restrictions through mathematical formulas, such as when an agent might generate multiple sub-agents to evade restrictions on the main agent. NVIDIA said it is working with Arm and Intel to ensure the security tool runs on their CPUs, and has partnered with dozens of institutions, including Anthropic, to launch these tools.

The risks of AI agents have long exceeded the scope of "saying the wrong thing." An agent capable of handling files, calling APIs, and accessing the network may leak private documents due to a malicious instruction hidden in a web page, cause financial losses due to misinterpreting tool parameters, or silently deviate from its path and perform dangerous actions. AI is simultaneously rewriting software development and cyber warfare—AI generates large amounts of new code, sometimes introducing vulnerabilities. In the past, it took weeks or even months for a high-risk vulnerability to be discovered and turned into an exploit, but now this process is being compressed into a few hours by AI.

NVIDIA CEO Jensen Huang rejected widespread calls for AI safety regulation, instead viewing runaway agents as an engineering problem to be solved, similar to improving car safety.

However, AI safety experts hold a different view. Maurice Chiodo, a mathematician at the Centre for Existential Risk at the University of Cambridge, once stated that the people who design, develop, and launch these tools themselves cannot be trusted to develop and ensure their safety. What worries him even more is evidence suggesting that both OpenAI and Anthropic failed to monitor their agents when they became uncontrollable.

The two major AI laboratories in the United States, OpenAI and Anthropic, have recently been exposed to several AI safety incidents. In July, OpenAI first publicly disclosed an intrusion incident, where its model actively discovered and exploited an unknown "zero-day" vulnerability to break out of the sandbox environment and infiltrate Hugging Face, all done autonomously by the AI without any human instructions. In September this year, Australia's Deputy Prime Minister Marles confirmed that an OpenAI agent entered a health statistics service portal managed by the Australian Service Bureau in June and accessed some public and non-public documents.