OpenAI has revealed new details about its upcoming Astra model, describing it as the first large language model to reach the company's internal cybersecurity benchmark. The model is expected to launch soon, while its most advanced security-related features will be available to a narrower group of users.
According to OpenAI, Astra can identify previously unknown weaknesses in computer systems and carry out exploitation steps with little or no human direction. The company says it used a series of evaluations, including ExploitBench, where Astra reportedly achieved a perfect result. In a modified test designed by OpenAI engineers, the model also found and used two zero-day vulnerabilities.
To prepare for release, OpenAI says it has strengthened its safeguards against misuse, including improved abuse detection, jailbreak prevention, and additional monitoring of reasoning patterns. The company also says it is limiting responses for accounts it considers higher risk, while applying new safety techniques to make the model more resilient.
OpenAI added that Astra was tested against scenarios inspired by recent agent behavior in research environments, and said the model did not attempt to leave its test setup during those trials. The company plans to share more evaluations and safety details as the rollout expands.
As AI systems become more capable, Astra may help shape a future where advanced models are built with stronger security, tighter controls, and more responsible deployment standards.