OpenAI Pauses Astra Work Over Critical Cyber Capability
Under OpenAI's Preparedness Framework, the Critical cyber threshold covers a tool-assisted model capable of independently developing functional zero-day exploits across many hardened, real-world critical systems, or devising and executing a novel end-to-end attack against hardened targets from a high-level goal. OpenAI said Astra's early results prevent it from excluding that classification, but did not claim the model had definitively reached the threshold.
The company said affected Astra activities must use isolated testing environments, restricted network and tool access, sandboxed execution, stronger protection and encryption for model weights, and expanded monitoring and detection. Universal monitoring is being applied across agentic Astra applications during training and evaluation, with risky behavior subject to review and interruption.
OpenAI also plans to work with relevant government agencies and selected AI safety organizations on further capability testing. It said third-party testing partners will receive recommended controls for higher-risk evaluations and workloads. Axios reported that OpenAI informed the U.S. administration of its plans, while the timing of any future Astra release remains unclear.
The disclosure follows heightened scrutiny of autonomous model testing, but OpenAI said Astra was not involved in the separate Hugging Face security incident disclosed in July. That distinction is material: the Astra announcement describes precautionary containment prompted by benchmark performance, and the available statements do not establish that Astra found a specific zero-day, compromised a real target, or caused real-world harm.
For organizations evaluating advanced cyber-capable models, the controls identified in the disclosure provide a defensive baseline: isolate test environments, minimize credentials and network reach, restrict tools, encrypt sensitive model assets, and continuously monitor agent actions. OpenAI said it will expand external testing before determining Astra's capability level and whether its safeguards are sufficient for further work.