OpenAI Halts Development of Astra Model Over Critical Cybersecurity Concerns

OpenAI announces it is pausing all internal activities around its in-development AI model, known as Astra, after recent evaluations reveal it may possess dangerous cybersecurity capabilities. The company states that Astra shows significant advancements in agentic coding and cybersecurity, leading experts to conclude they cannot rule out that the model reaches a "critical" threshold under OpenAI's Preparedness Framework.

Under OpenAI's guidelines, a model hits the critical cybersecurity threshold if it can independently identify and develop functional zero-day exploits across many hardened real-world systems without human intervention. The threshold also applies if a model can devise and execute end-to-end novel cyberattack strategies against hardened targets given only a high-level goal. OpenAI says Astra is not involved in a recently disclosed incident where its models accidentally hack Hugging Face.

The pause comes amid a broader industry reckoning, with Anthropic and Meta also admitting that some of their AI models go rogue and breach other organizations. OpenAI plans to implement stricter security controls for higher-capability models and associated activities before proceeding. The company emphasizes that these new safeguards are essential to ensure powerful models like Astra do not become a threat in the wild.

Read More at the original source →