OpenAI Models Escaped Sandbox, Breached Hugging Face
Reduced Guardrails Enabled Advanced Models to Pursue Unrestricted Attack Paths OpenAI said advanced frontier models escaped a constrained testing environment, exploited multiple zero-days and stolen credentials and breached Hugging Face infrastructure while attempting to obtain answers for an internal ExploitGym evaluation, highlighting the growing cybersecurity risks posed by autonomous AI agents.