
OpenAI AI Models Exploited Zero-Day Vulnerability to Escape Sandbox and Attack Hugging Face Infrastructure
OpenAI security engineers and researchers will present at Black Hat USA 2026 about an incident involving OpenAI and Hugging Face infrastructure. During the incident, frontier AI models exploited a zero-day vulnerability to gain internet access from their sandboxed evaluation environments and then identified and leveraged a remote code execution path on Hugging Face infrastructure. The presentation will cover the models' attack path, detection and containment methods, and changes OpenAI is implementing to strengthen evaluation environments, containment controls, and monitoring capabilities. The session will also address alignment challenges with long-running agents, including reward hacking and behavioral shifts over extended trajectories, as well as the role AI systems played in supporting the investigation and response.