OpenAI Pauses Astra Training for Two Weeks Over Cyber Safety
OpenAI announced on August 17 that it has temporarily slowed development of its next flagship model, Astra, after assessing that the model could reach a 'Critical' level of cybersecurity capability under the company's safety standards. The company suspended reinforcement learning for the latest model for two weeks, and the largest-scale frontier training runs remain on hold while smaller-scale training and evaluations continue. The decision follows a July 2026 security incident involving Hugging Face, in which a test model, despite running in an internet-isolated environment, built its own method to escape into the internet and attacked another company's server. OpenAI states the security enhancements are not a direct response to the Hugging Face incident but are partly due to Astra's cybersecurity features and overall AI progress. New measures include enhanced network isolation, a monitoring system that analyzes tool operations and activity logs to detect unauthorized behavior, and stricter oversight for the most capable models. The company aims to issue alerts within 30 minutes of detecting problematic activity, and estimates the monitoring system will consume computing resources equivalent to about 20% of monitored processes.