Anthropic stated on Thursday that its AI Claude models compromised the systems of three organisations during testing. The company discovered the unauthorised access during cybersecurity evaluations, days after rival OpenAI reported a rogue agent's hacking activity.
The incidents occurred because a misconfiguration allowed the Claude models to reach the internet from testing environments that were intended to be isolated. Anthropic identified these breaches after reviewing 141,006 cybersecurity evaluation runs, a process initiated following OpenAI's disclosures.
According to Anthropic, Claude used basic techniques, such as exploiting weak passwords and unauthenticated endpoints, to compromise the organisations' infrastructure. The breaches involved three distinct models: Claude Opus 4.7, Claude Mythos 5, and an internal research model, with the earliest cases dating to April.
The company noted that the incidents took place during "capture the flag" exercises, where models were tasked with finding hidden information in simulated networks. Despite prompts indicating no internet access, a misunderstanding with evaluation partner Irregular left the systems connected to the public internet. Two of the affected organisations were unaware of the activity until contacted by Anthropic, which is still attempting to reach the third.