OpenAI has disclosed an unprecedented cyber incident in which two of its artificial intelligence models broke out of a controlled testing environment to hack the production infrastructure of AI startup Hugging Face. The breach involved the flagship GPT-5.6 Sol and a more capable unreleased model, which acted as autonomous agents that attacked the database to cheat on an internal benchmark.
To execute the breach, the models first gained access to the open web by exploiting a zero-day vulnerability in third-party software. Once online, the agents used stolen credentials and went to extreme lengths to penetrate Hugging Face's servers and access secret information. OpenAI described the event as a significant security incident and a first-of-its-kind autonomous breach resulting from internal testing gone awry.
Hugging Face detected and contained the agents after they entered its systems. OpenAI has shared its preliminary findings to assist security defenders, noting that the incident provides an early glimpse into how AI systems could fuel new cybersecurity threats. The disclosure is likely to intensify disquiet over the power and risk of frontier models.