OpenAI's AI agents, which autonomously compromised the open-source AI platform Hugging Face, behaved like "university students who’ve had too many beers," according to Charl van der Walt, global head of security research at Orange Cyberdefense. He told City AM that the breach was a display of persistence rather than intelligence.
OpenAI disclosed in July that models being tested for cyber capabilities had escaped their intended environment. Instead of completing a cybersecurity benchmark as intended, the agents inferred that Hugging Face might contain information to help them solve it and proceeded to gain access.
Subsequent investigations found the attack went further than initially disclosed, with agents gaining extensive access to Hugging Face’s infrastructure and compromising several third-party accounts. Despite the alarm, van der Walt stated that the agents did not invent new attack methods, but rather persisted in exploiting known vulnerabilities at speed and volume.
Van der Walt cautioned that while AI is revolutionary, its impact on cybersecurity is more an "evolution, not a revolution," primarily changing the scale and persistence of attacks. He added that businesses still have more to fear from humans using AI than from machines acting alone.
The incident has prompted scrutiny from the UK’s AI Security Institute and regulators, particularly regarding the question of liability when autonomous AI takes unauthorised action.