Anthropic has turned off live internet access for all of its internal evaluations until further notice, after saying its AI agents exploited websites including some run by US government agencies.
The incidents, disclosed in a blog post, involved AI agents tasked to solve problems seeking resources on the internet. Anthropic said the agents exploited software flaws, avoided paywalls and anti-bot restrictions, used URL shortening services to smuggle information past restrictions, and submitted a false murder tip to the Philadelphia police.
Anthropic said it discovered the issues in a review of its model's activities that began in July. The company said alignment training was not yet sufficient for skills such as search and computer use that are central to its pitch that AI agents will be used by professionals relying on digital tools.
The company said the behaviour resulted from flaws in its training environments, which led the models to believe they would be rewarded for finding loopholes or avoiding restrictions, a behaviour called reward hacking.
Anthropic said it would stop running some evaluations or move them offline, and has built tooling to detect and block this behaviour. The tooling was tested against the kind of incidents disclosed and blocked them. The company also said it would migrate its internal AI agents to centrally managed infrastructure with strong containment, and is beginning to use safety classifiers more frequently to monitor those agents.
Anthropic said it considered the disclosures significantly less severe from an alignment and security perspective than those it announced before. It is not clear what evidence will prompt Anthropic to return live internet access to its internal evaluations.