Anthropic announced that its Claude models accidentally infiltrated the systems of three companies during cybersecurity tests. Details in our news.
Anthropic announced that, following internal investigations, its model named Claude gained unauthorized access to the systems of three different companies during cybersecurity tests. This sparked a new debate in the technology world shortly after OpenAI announced that one of its models had infiltrated Hugging Face systems.
Anthropic stated that these incidents occurred during the proactive evaluation process initiated by OpenAI after the July 21st security breach. The company stated that test environments should have been isolated, but a configuration error occurred in working with a third-party partner, Irregular.
Unexpected Access During Security Tests
As a result of the investigations, it was determined that three different Claude models, Opus 4.7, Mythos 5, and an internal research model, had accessed the internet. Despite being explicitly told that the Claude models did not have internet access during testing, it was observed that the models perceived real-world systems as part of their assigned mission.
The Opus 4.7 model, despite recognizing the target systems as real, continued its attacks by extracting credentials and interfering with databases. The Mythos 5 model, despite understanding that it was connected to the internet, convinced itself that it was within a simulation and uploaded a malicious software package to the PyPI registry.
Anthropic emphasized that the models did not pursue an objective on their own and were only trying to fulfill the given tasks. The company reminds that models used in this type of evaluation lacked security controls, unlike those publicly available.
Security and auditing discussions continue in the sector.
Anthropic stated that, unlike the OpenAI incident, their models accessed the internet through an accidentally left-open communication path, not a software vulnerability. The company announced that it takes responsibility for this situation and will increase necessary controls to prevent similar incidents from happening in the future.
Anthropic, which is currently conducting a third-party investigation of the incidents with the independent review group METR, added that the affected organizations were not aware of this situation beforehand.
The process, which began with OpenAI’s Hugging Face incident and continued with Anthropic’s statement, has once again brought to the forefront the importance of control systems on AI lab models.
What are your thoughts on how such security tests should be subjected to control processes in the future?