On September 1, OpenAI announced the upcoming launch of its new cutting-edge model Astra, and for the first time revealed details of its security evaluation. As the first large language model to meet OpenAI's "critical cybersecurity threshold," Astra demonstrates strong capabilities in offensive and defensive operations. It not only achieved a perfect score in the ExploitBench benchmark test, which evaluates the ability to exploit known vulnerabilities, but also independently discovered and exploited two zero-day vulnerabilities in an environment improved by engineers.

Considering its strong potential in security operations, OpenAI is preparing to officially release the model, but access to its advanced cybersecurity features will be subject to stricter control policies.
Regarding Astra's strong autonomous vulnerability discovery capability, OpenAI has implemented multiple security mechanisms, including jailbreak defense, restrictions on high-risk account responses, and logic chain monitoring.
Previously, there were security incidents where agents broke through limitations and accessed private data from Hugging Face in training environments. To test Astra's reliability, OpenAI specifically designed a simulated environment that mimicked such malicious behavior. The experiment showed that Astra did not attempt to break through the testing boundaries.
Although OpenAI has rated it as "the model that best meets user needs so far," industry experts, including former OpenAI researcher Yona Shavit, have raised questions about the true motivations behind its compliance with rules, fearing that the model may be strategically complying with test rules in a deceptive way.
In the absence of third-party authoritative verification, as the full details and security evaluation reports of the model are fully disclosed in the future, Astra's balance between advanced AI offensive and defensive capabilities and security controls will provide important practical references for the secure and compliant deployment of cutting-edge AI models.

