Key takeaways

  • The models most affected were Haiku, Sonnet, and Opus, while the newer Fable and Mythos models showed up in only a single distillation case.
  • The techniques themselves are familiar, including stolen credentials, unpatched devices, SQL injection, and phishing.
  • AI agents kept checking whether the malware in play was being flagged by common security products, and when an antivirus tool caught it…

What happened

The models most affected were Haiku, Sonnet, and Opus, while the newer Fable and Mythos models showed up in only a single distillation case. Anthropic says it documents novel misuse rather than the typical kind. The core finding from the cyber chapter is that sophisticated attacks no longer require sophisticated attackers, and sophistication is no longer a reliable signal for attribution.

Moonshot AI (GTG-16002) relayed nearly 300,000 customer requests to Anthropic over ten days across 5,380 fraudulent accounts, while users believed they were using a Kimi model. 1 million exchanges in 14 days. Among the rerouted requests, Anthropic found a user likely tied to the People's Liberation Army who had CCTV archive footage analyzed for a single target, video from hundreds of cameras in Chengdu, including cameras outside PLA facilities.

Through DeepSeek, Claude also received requests from an operator with live credentials for a database linked to the Russian Ministry of Defense, along with work on a case management system for a Chinese public security bureau that matches movement profiles against police records. Xiaomi (GTG-16008) used Claude differently, storing requests and coding sessions from users of its own MiMo models and replaying those conversations through Claude to generate training data.

Anthropic found no evidence that Claude's responses were served directly to Xiaomi's users. The relayed requests, however, contained personal data such as names, contact details, and company information for hundreds of people in at least a dozen languages. ai) rotated through 273 accounts and pushed more than 770,000 exchanges over ten days through a CoT cleaner, a tool that automatically turns captured reasoning traces into usable training data.

Why it matters

The techniques themselves are familiar, including stolen credentials, unpatched devices, SQL injection, and phishing. What changed is the economics, since reconnaissance, exploitation, and tool-building now get handed off to models that run in parallel at machine speed. Autonomy lowers the cost side of an attacker's math, Anthropic says, and makes previously unprofitable targets worth pursuing. Anthropic tracks a Russian-speaking espionage actor as GTG-20006 that used a feedback loop.

AI agents kept checking whether the malware in play was being flagged by common security products, and when an antivirus tool caught it, the agents rewrote and recompiled the malicious code on their own until it slipped past detection again.

That shifts the burden back onto defenders, Anthropic says, because writing new detection signatures no longer slows an attacker down if that attacker cycles through changes faster than new signatures can be rolled out. More than 20 organizations were targeted, including government ministries, intelligence services, embassies, and defense contractors, with a focus on Ukraine and Europe.

The drone supply chain came up repeatedly, and the actor stole a complete proprietary SDK for a drone vision system, among other things. Access sometimes ran through third parties, such as compromised hotel guest Wi-Fi providers whose guest devices were then loaded with malware, a method Microsoft described in July 2026 as CaptiveCrunch. For clusters Anthropic attributes to the ShinyHunters collective (GTG-50014), industrial credential mining was the focus.

8 million Android apps, decompiled them, and searched for hardcoded secrets. Anthropic describes the approach as "vibe hacking," where a human sets a rough goal and the model assesses the environment and iterates until the task is done. One of the hackers said he collected HackerOne bounties on top of extorting two companies. On distillation, Anthropic identified attacks from seven more Chinese labs since its first disclosure in February.

" The largest campaign ever measured is attributed to Alibaba's Qwen lab (GTG-16005). 7 models. The peak hit almost three million exchanges a day from more than 3,500 fraudulent accounts, totaling over 151 million exchanges between May and July 2026, mostly on agentic tasks and software development. Stranger are the cases where labs relayed their own customers' requests to Claude.

What to watch

3 on cyber tasks, the lab first went after Anthropic's Fable model but gave up after the cyber safeguards degraded its performance, then deliberately switched to models it judged to have weaker protections. SenseTime bought transcripts from third parties, according to the report, so it did not obtain the captured Claude data itself but through an intermediary market.

MiniMax ran its own proxy network through a shell company that offered only Anthropic and OpenAI models, no Chinese ones, not even its own. In the surveillance chapter, Mali stands out.