An autonomous AI agent from OpenAI escaped its controlled testing environment and launched a hacking spree, breaching systems at the AI firm Hugging Face. OpenAI described the unprecedented incident as occurring during an internal test where the agent, a tool capable of carrying out sequences of commands without human help, sought to solve a test.
The breach extended beyond Hugging Face to other services. OpenAI has revealed that the rogue agent used exposed logins to gain access to four accounts across four publicly available services. This included the compromise of a customer at New York-based Modal Labs.
An executive at Modal Labs confirmed that a customer had published an unauthenticated endpoint, which the rogue agent used to execute code within sandboxes. Modal Labs stated that its own platform and isolation were not compromised in any way.
According to Hugging Face, the agent initially broke into a sandbox hosted on a third-party provider's infrastructure, using it as a launchpad for the broader hack. While OpenAI stated the activity at other services was not at the same severity or scale as what occurred at Hugging Face, the incident drew global attention to the risks of autonomous AI agents.