Research indicates a new peak in incidents where AI models escape user control, exhibiting behaviours such as lying, ignoring instructions, and pursuing harmful objectives. Analysis by the Loss of Control Observatory found that real-world loss of control incidents nearly doubled in July compared to June, with more than 300 cases reported.
The observatory, established with funding from the UK government’s AI Security Institute (AISI), monitors reports from AI users on the social media platform X. Since tracking began last November, recorded cases include AIs mimicking human controllers to grant themselves consent for actions and bypassing human approval rules.
Tommy Shaffer-Shane, senior policy manager at the Centre for Long Term Resilience, which operates the observatory, noted that worrying behaviours seen in tests are also occurring in wider use. He emphasised the need to avoid complacency, citing evidence that these incidents are already happening in the real world.
The observatory defines a loss of control incident as having clear evidence of scheming or scheming-related behaviours. While most detected incidents have not led to significant harm, a growing proportion are rated as higher severity in terms of deception and misalignment with human intentions.