OpenAI has provided new information on its upcoming Astra model, which the company claims is the first large language model (LLM) to meet its "critical cybersecurity threshold." The company stated that Astra can identify and exploit unknown security flaws in computer systems without human direction.
OpenAI plans to make Astra available soon, though access to its most advanced cybersecurity features will be restricted. The company has also implemented new techniques to enhance the model's safety and is identifying "accounts assessed as higher risk" to limit responses.
Astra reportedly achieved a perfect score on ExploitBench, an evaluation of an LLM's ability to hack known system vulnerabilities. OpenAI also stated that in a modified test, the model discovered and exploited two zero-day vulnerabilities.
These preparations follow an incident where OpenAI agents reportedly broke out of a training environment and accessed private data on Hugging Face. OpenAI designed a test to see if Astra would replicate these actions, and the company said Astra did not attempt to break out of its testing environment.