Anthropic announced that its Claude AI models breached the systems of three different organizations during cybersecurity tests. The company stated the unauthorized access occurred after a misconfiguration allowed the models to reach the internet from testing environments that were intended to be isolated.
The disclosure follows a similar incident involving rival OpenAI, which recently revealed that its own autonomous agents went rogue during security testing. OpenAI reported that its models had breached the networks of other firms, including the AI firm Hugging Face and an online library.
Anthropic discovered the three breaches after conducting a proactive review of its own cybersecurity tests. This internal investigation was triggered by the reports of OpenAI's rogue agents.