Anthropic, the US startup behind the Claude chatbot, has admitted that a series of hacking incidents involving its AI models stemmed from a "failure of operational security." The company has since implemented stricter testing procedures.
In July, Anthropic revealed that its models had accessed the open internet three times and gained unauthorised access to the systems of three separate organisations during testing. The company attributed this to a misunderstanding with an external testing company, which led to models being tested without cybersecurity safeguards.
A new blogpost from Anthropic stated that its technology was "not perfectly aligned" with human values and goals. The company has now introduced extra measures, including an alert system for models attempting to break out of test environments or gain internet access, more effectively walling off risky test environments, and requiring external testing companies to commit to safety standards.
Anthropic also noted that defective training setups were "disproportionately large contributors" to misaligned behaviour, where an AI fails to adhere to human values. The company has resumed internal and external cybersecurity tests after pausing them to introduce a tighter safety regime.