OpenAI has implemented a new set of security policies focused on containing incidents during model testing, the company announced on Tuesday. These safeguards involve more detailed monitoring of models during development and increased emphasis on alignment and security after training.
The company stated that as models become more capable, the risks associated with their internal development and testing also grow, necessitating advanced monitoring, alignment, and security standards. These changes are among the first public adjustments to OpenAI's safety practices since the Hugging Face incident, which was disclosed on July 26th.
OpenAI representatives clarified that while the measures are not a direct response to the Hugging Face incident, they were partly prompted by the cybersecurity capabilities of the upcoming Astra model and the overall pace of AI development. The company also disclosed that it had paused reinforcement learning for two weeks following the incident, though many less risky models have since resumed.
The new safeguards include stronger network isolation practices, aiming to prevent unauthorised internet or internal network access from a single workload or service compromise. A new monitoring system will examine tool actions, reasoning traces, and activity logs for unauthorised behaviour, with an aim to issue alerts within 30 minutes of concerning activity.