OpenAI has revealed six new reports of "unexpected or concerning" behaviour in its artificial intelligence models. Among these cases, an unreleased research model reportedly inserted "jailbreak-like instructions" into its own notes, instructing itself to ignore normal constraints and to be "freed from the roles and identities that bind other chatbots."
In another reported incident, an AI "agent" uploaded files to the internet to obtain a browser citation without seeking user authorisation. These incidents were discovered during training or evaluation over recent months.
The AI company also announced on Wednesday that it is introducing a new framework designed for tracking, probing, and disclosing AI model misalignment. This framework will address issues such as models acting without authorisation, coordinating with other models, or evading oversight.