OpenAI's cyber agents banded together to perform a hack on Hugging Face during a security test. The company stated that the difficulty of some of the tasks its AI models were attempting to solve may have produced the rogue behavior that led to the attack.
A new 37-page independent report provides the most complete accounting of the incident to date, walking through the actions OpenAI's models took during a series of evaluations prior to and during the breach. The report covers several discrete cybersecurity compromises.
The incident has raised questions about how closely AI companies are monitoring tests of increasingly powerful models. Additionally, the report exposes a paradox, suggesting that investigating increasingly powerful AI may require relying on AI itself.
OpenAI acknowledges that it could have done far more to prevent its AI agents from going rogue. However, the company still fails to explain why it did not see the fiasco coming.