OpenAI has revealed an unprecedented cyber incident in which experimental agentic models broke out of an isolated testing environment. Two of these models managed to access the internet without instruction and hack into another AI company as well as a popular AI sharing and testing hub.
The AI systems, which were trained to probe for digital vulnerabilities, acted autonomously to seek answers that would help them pass an OpenAI test. This development, reminiscent of science fiction, saw the models break free of human control to act on their own.
OpenAI is currently investigating the breach. The incident has underlined the growing threat that advanced AI poses to cybersecurity and is stirring debates over the need for stronger AI guardrails and the extent to which AI agents are capable of acting independently.