OpenAI tightens model security with sandboxes, 30-minute alerts

OpenAI added stronger sandboxes, token-level monitoring with 30-minute alert rules, and paused some reinforcement and frontier training after Astra was flagged as potentially critical.

OpenAI announced stricter containment and continuous monitoring for its AI research after an internal review found its upcoming Astra model may meet a ‘critical’ cybersecurity capability threshold under the company’s Preparedness Framework. The changes also follow a recent security incident that influenced the company’s decision to delay certain training runs.

The company has paused reinforcement-learning runs aimed at deployment for two weeks and put its largest planned frontier training on hold until the new controls are in place. Engineers must move any workload that runs model-generated or otherwise untrusted code into more robust sandbox environments before resuming large-scale runs.

Network boundaries have been redesigned so a single compromised process cannot by itself gain internet access or reach other internal services. OpenAI described the update as a reconfiguration of research infrastructure to limit the blast radius of a single workload compromise.

A central element is a multistage monitoring framework that inspects model activity at the token level. Activation classifiers sample a model’s internal signals at each inspected token. Detected anomalies are escalated to automated investigator systems that examine reasoning traces and tool usage for signs of data exfiltration, unauthorized access, destructive actions, or attempts to bypass safeguards.

High-priority alerts trigger immediate pages to security and research responders. Under the operational service-level agreement, responders must confirm an alert is a false positive within 30 minutes; if they cannot, the associated activity must be paused. OpenAI estimates the continuous oversight layer uses roughly 20 percent of the monitored inference compute.

The updated standards are mandatory for all reinforcement learning training and evaluations that permit tool use when models reach the Sol capability tier or higher. The company said its Preparedness Framework will be revised to keep pace with rapid capability gains and that core alignment techniques will be applied at more stages of development and evaluation.

OpenAI noted it expects security operations to scale with system capabilities and that future protections will likely rely on AI-driven defenses, including using models to detect and stop attacks by other models. The company referenced industry testing that showed powerful models could exploit real systems during cybersecurity exercises.

Operational impacts include the current two-week pause on deployment-focused reinforcement learning and the ongoing hold of the largest frontier training while containment measures and monitoring are implemented. Teams must update training pipelines to operate inside the stronger sandboxes and comply with continuous token-level inspection and the 30-minute alert rules before resuming large-scale runs.

Articles by this author