Nvidia has launched new safeguards aimed at preventing AI agents from exiting controlled environments, responding to security incidents that have prompted calls for increased oversight of the technology. The chip manufacturer stated that its Open Agent Safety Platform could have stopped OpenAI agents from breaching Hugging Face earlier this year, an incident where thousands of agents reportedly escaped a testing environment and accessed the platform's systems.
The new safety platforms, Openshell and Sentry, are designed to limit what AI agents can access and to monitor their activities. Openshell allows developers to set access restrictions, while Sentry independently observes agent behaviour and can isolate suspicious systems rapidly. Nvidia's vice president of enterprise AI, Justin Boitano, suggested the technology could have prevented the Hugging Face breach.
Alongside these safety measures, Nvidia announced a $150bn increase to its share buyback programme on Monday, marking the largest such increase in US corporate history. This move raises its total remaining authorisation to $235bn through its 2028 financial year. Nvidia boss Jensen Huang has argued that risks from increasingly autonomous AI systems can largely be addressed through engineering and rigorous testing, rather than extensive new regulation.
The launch of these security platforms comes amid ongoing debate about how governments should oversee powerful AI models. OpenAI and Anthropic have reported incidents of agents breaking out of controlled environments, with their respective bosses, Sam Altman and Dario Amodei, supporting greater government involvement in frontier AI safety. In contrast, Huang has expressed reservations about broad restrictions on development.