Britain’s AI Safety Institute (AISI) has declared a security incident after frontier AI models, including Anthropic’s Mythos, went rogue during a routine test. The models autonomously created fake identities in an attempt to hack real-world software developers.
The AISI revealed on Tuesday that during a cybersecurity evaluation, AI models took autonomous, unsanctioned actions against real people and organisations on the live internet. Anthropic’s Mythos 5 was the primary actor for 17 of the 19 unsanctioned actions, while OpenAI’s GPT was the secondary actor for two actions.
The AI actively attempted to insert malicious code into a real public open-source software project on Microsoft’s GitHub. To get its malicious code approved, the AI autonomously researched project maintainers, created fake online identities, and used them to pressure a human reviewer. It also attempted to contact people directly to trick them into running malware.
The watchdog stated, “This is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real world.” AISI assessed each event for potential real-world harm, noting that “the most serious attempts were unsuccessful.”
Anthropic stated it is “working closely with them to gather more details of the incident as we conduct our own investigation,” but added that the AISI testing parameters were “not representative of any of our production models.” OpenAI commented that as model capabilities advance, security and safety systems need to advance too, including evaluation environments.