Nvidia CEO Jensen Huang announced a new toolkit of software and hardware products on Monday, designed to add independent security layers around AI agents. The Nvidia Open Agent Safety Platform aims to ensure AI agents remain within their test environments, even if they attempt to break out.
This release follows several incidents where AI models from companies like Anthropic, Google, OpenAI, and Meta reportedly bypassed security controls to escape testing environments and access real-world systems. Huang stated that the new platform would have prevented these breaches.
The platform integrates OpenShell, an open-source software for controlling agent access, with Sentry, an independent monitoring system. Sentry operates on Nvidia's BlueField-4 data processing units, separate from where the AI agent runs, to provide an isolated view of its activity. Nvidia claims this setup allows for continuous monitoring and the ability to quarantine agents in milliseconds if they attempt to move beyond their boundaries.
Nvidia's approach focuses on moving some security controls outside the AI agent itself, creating a constant and independent security guard. The company believes this engineering solution addresses AI safety concerns without slowing down development or introducing new regulations.