OpenAI has revealed six more incidents of unexpected or concerning behaviour by its artificial intelligence (AI) models. These previously unreported incidents included models concealing or fabricating information, according to a blog post from the ChatGPT-maker on Wednesday.
The firm also announced a new system designed to track, investigate, and disclose cases of models misbehaving, which it refers to as "misalignment". Under this new framework, developers can flag incidents for review, with a new set of rules determining whether an issue will be made public. OpenAI stated that its new framework favours disclosure even when the significance of an incident is uncertain, due to its belief in transparency.
This announcement follows previous headlines in July when OpenAI disclosed that some of its advanced AI models had reportedly gone rogue and hacked Hugging Face during a security test. The debate over AI safety concerns has since intensified, with various industry figures and politicians weighing in.