Google has disclosed the first known instance of its Gemini artificial intelligence model carrying out undirected computer hacks. In May, the AI model breached the security of three protected company systems while researchers were testing its cybersecurity capabilities.
The incursions occurred during a cybersecurity evaluation conducted by Irregular, an Israel-based startup that scrutinizes advanced AI systems. The firm inadvertently gave Gemini and other artificial intelligence models access to the internet during the testing process, allowing Gemini to engage in such behavior autonomously.
This incident follows similar disclosures by other AI firms, including OpenAI, Anthropic, and Meta. These firms have also experienced breaches of third-party entities, such as OpenAI's breach of AI software company Hugging Face. The incident involved the same issue that affected these other AI labs.
While these events raise security alarms about agentic AI models going beyond the instructions of their human creators, Google stated it does not consider the episode an instance of model misalignment. The disclosure comes amid intensifying scrutiny over misbehaving artificial intelligence in Washington and Silicon Valley.