Key takeaways
- The British AI Security Institute (UK AISI) and the U.S.
- It uses 41 vulnerabilities found in Chrome's V8 engine after 2023 to track how far a model advances through the software exploitation…
- Only a small group of models can solve TLO at all.
What happened
S. Center for AI Standards and Innovation (CAISI) jointly evaluated Moonshot AI's latest model, Kimi K3. S. 2, setting a new benchmark among open-weight models. Its safeguards didn't block exploit development or offensive cyber operations, and the model assisted with both without pushback. The institutes used ExploitBench, a benchmark developed by Carnegie Mellon University, to test exploit development skills.
Hugging Face fended off the attack, though it took real effort and the use of open-weight models. S. and Chinese models since early 2025 on an Elo-based scale. S. counterparts. In a previous analysis, the British institute pegged the performance gap for open models at four to seven months, compared with six to ten months at the start of 2025. The new results fit this pattern. S. systems.
AISI warns that this gap shouldn't breed complacency. " The Kimi findings also lend support to distillation allegations against Chinese model developers. S. science advisor Michael Kratsios recently accused Moonshot AI of "distilling" Anthropic's Fable by using Fable's best outputs as training data to boost Kimi K3's performance. S. export controls.
Why it matters
It uses 41 vulnerabilities found in Chrome's V8 engine after 2023 to track how far a model advances through the software exploitation process. S. 2. Kimi K3 didn't reach the highest level, known as Arbitrary Code Execution (ACE), on any of the 41 tasks. ACE is the most severe exploit level because it gives attackers full control over a target system. S.
models achieved ACE in 20 of the 41 tasks. S. closed-weight models with their system-level safeguards disabled to measure their maximum capabilities. Those safeguards are enabled in the publicly available versions. The second test, "The Last Ones" (TLO), simulates a corporate network attack with a 32-step attack path across four subnets and about 20 hosts. A human expert would need roughly 20 hours to complete it, according to the institutes.
Only a small group of models can solve TLO at all. Four publicly available closed-weight models have passed the test so far, with the strongest succeeding six or seven times out of ten. S. 2. It completed the entire attack path in one of ten attempts while staying within the 100 million token limit, showing that it has the capability but can't call on it reliably.
"Kimi K3 is capable of autonomously attacking small, weakly defended and vulnerable enterprise systems, when directed to do so and given initial network access", the institute writes. TLO doesn't account for active defense, so it isn't fully realistic. But the results would raise red flags in real-world scenarios. A fresh example showed up this week when OpenAI models tried to autonomously hack into Hugging Face.
What to watch
One explanation for the gap between strong general benchmarks and weak cyber scores is that Kimi K3 may have been trained mostly on Claude outputs covering general knowledge, programming, and agent tasks. Anthropic's safety classifiers specifically block advanced offensive cyber queries, so those outputs would be underrepresented in a distillation dataset built from Claude responses.
Kimi K3 could therefore match leading Western models on standard benchmarks without picking up their deeper exploit capabilities. The AISI results support this reading. S. models, revealing cyber capabilities that are nearly impossible to access through public interfaces and therefore largely unavailable for distillation.

