Key takeaways
- After reviewing 1.41 lakh cybersecurity evaluation runs, Anthropic disclosed that three Claude models accessed real-world production…
- 41 lakh cybersecurity evaluation runs, Anthropic disclosed that three Claude models accessed real-world production systems because of a…
- 41 Lakh cybersecurity evaluation runs following the OpenAI disclosure and identified three incidents where Claude accessed the public…
What happened
41 lakh cybersecurity evaluation runs, Anthropic disclosed that three Claude models accessed real-world production systems because of a misconfigured third-party testing environment, unlike OpenAI's sandbox escape caused by a zero-day exploit The incidents saw Claude breach a production database, upload a malicious package to PyPI that was downloaded by 15 systems, and an internal model scan 9,000 internet-connected systems before halting after recognising a real production environment Anthropic suspended the evaluations, tightened safeguards and warned other AI labs to review testing environments, with the incidents adding to growing concerns among enterprises and governments over autonomous AI agents' cybersecurity risks Days after OpenAI disclosed that one of its AI models escaped a sandboxed testing environment and attempted to breach Hugging Face’s production infrastructure, Anthropic has revealed that its Claude models also gained unauthorised access to the production systems of three organisations during cybersecurity evaluations.
The disclosure follows OpenAI’s confirmation earlier this month that one of its AI models exploited a previously unknown software flaw to break out of a testing environment, gained internet access and attempted to breach Hugging Face’s production infrastructure during an internal cybersecurity evaluation.
According to OpenAI, the activity was detected and contained before causing wider impact, prompting both companies to launch a joint investigation and tighten safeguards around future AI security testing. Anthropic said it is strengthening security around its evaluation infrastructure, improving monitoring of testing environments and working with independent AI evaluation organisation METR to review the incidents.
It also urged other AI labs to review their cybersecurity evaluation processes as increasingly autonomous AI systems become more capable. These incidents further underscore the cybersecurity risk posed by AI, which has led to governments across the world reconsider their safety protocols. Earlier in the month of April, Anthropic’s AI model Mythos also led to several cybersecurity flags.
Why it matters
41 Lakh cybersecurity evaluation runs following the OpenAI disclosure and identified three incidents where Claude accessed the public internet while running in a testing environment operated by evaluation partner Irregular. Unlike the OpenAI incident, where the AI exploited a previously unknown zero-day vulnerability to escape its isolated environment before attempting to access Hugging Face, Anthropic said its model reached the internet because of a misconfiguration in a third-party evaluation environment.
The AI model had been assigned capture-the-flag cybersecurity challenges and was explicitly told it was operating in a simulated environment without internet access. However, because the evaluation systems had live internet connectivity, it treated real-world infrastructure as part of the exercise. 7, Mythos 5 and an unreleased internal research model. 7 mistook a real company for the intended test target.
It used weak passwords and exposed online services to gain access to this unanimous company’s systems and entered a production database containing a few hundred records. In the second incident, Claude Mythos 5 uploaded a malicious Python package to the public PyPI repository, believing it was part of the test.
The package remained online for about an hour and was downloaded by 15 real systems, including a cybersecurity firm’s automated malware scanner. The model then used credentials from that system to access more infrastructure. In the third incident, the internal research model scanned around 9,000 internet-connected systems and breached one organisation using exposed credentials and an SQL injection vulnerability.
It halted the attack after recognising it had reached a real production environment. Anthropic said none of the models attempted to escape the evaluation environment or pursued goals beyond completing their assigned tasks. It attributed the incidents to an operational failure rather than a model alignment issue, saying the models acted on the mistaken belief that every accessible system formed part of the simulation.
The company said it suspended all cybersecurity evaluations on July 23 after detecting signs of internet access, identified the three incidents the following day and notified its evaluation partner and the affected organisations on July 27.
What to watch
During internal red teaming — the process of evaluating AI models against specific tasks and security threats — researchers found Mythos to be highly proficient in cybersecurity. The model detected dormant bugs in decades-old code, identified critical vulnerabilities across major operating systems and web browsers, and even suggested potential methods to exploit those flaws.
Unlike traditional software, AI agents can independently plan tasks, execute a series of actions and adapt their behaviour when they encounter obstacles. As a result, enterprises and governments are rethinking how they approach cybersecurity. Contrails AI’s cofounder Ami Kumar told Inc42 that the biggest mistake organisations can make is treating AI agents like chatbots.




