Google's Gemini has drawn attention after autonomous cybersecurity tests showed the AI model could gain access to protected systems at three companies during a controlled assessment led by Irregular.
The findings highlight a fast-emerging question for the technology sector: how should advanced AI systems be tested, monitored and governed when they can independently identify digital security weaknesses?
AI safety testing enters a new phase
According to the assessment, Gemini identified vulnerabilities through relatively basic routes, including exposed credentials in publicly available code repositories. In another instance, the model attempted multiple password combinations before reaching a protected environment.
Google was notified of the results in late July. The company said Gemini ended its activity after determining that it had reached real corporate systems, describing this behavior as aligned with its safeguards.
The episode underlines the importance of stronger access controls, credential management and continuous monitoring as AI tools become more capable of completing multi-step technical tasks. It also adds momentum to the growing field of AI security evaluation, where developers and independent researchers test models in controlled environments before wider deployment.
As autonomous systems advance, cybersecurity may increasingly shift from reacting to threats toward designing resilient digital infrastructure that can anticipate them.