Key takeaways

  • A year and a half ago, DeepSeek R1 caused a shock.
  • In DeepSeek's own report, R1 beat o1 on individual tests like AIME 2024 but trailed clearly on others, such as factual knowledge (SimpleQA).
  • They handle long knowledge tasks, code across many steps, and coordinate tools far more reliably than their predecessors.

What happened

A year and a half ago, DeepSeek R1 caused a shock. A Chinese lab was suddenly competing with OpenAI's o1, the first commercial reasoning model, and had reportedly done it for far less money. Markets got nervous. Billions in market value evaporated within days. Was the planned infrastructure buildout overblown? The picture back then was murkier than the headlines suggested.

" Individual capabilities advance at different rates, which is why the much-quoted months-long gap always depended on who measured what. What's new is where a Western lead is still measurable at all. Three areas remain. The first is abstract specialty tests. Shortly after K3's launch, Opus 5 retook the top of the index with 61 points, a small gap over K3's 57. 5 percent. 2 percent.

But ARC-AGI deliberately measures abstract pattern recognition far removed from everyday tasks. There's no guarantee this gap predicts practical differences. It could simply stay economically irrelevant in most cases. The second is reliability. The AA-AnalystAgent benchmark, launched August 12, tests agentic data analysis on real tables and documents.

It only counts a task as solved if a model gets it right in five out of five independent runs, a measure the operators call pass^5. 5 at 50. K3 is the best open model at 39 percent. The gap comes almost entirely from poor repeatability. K3 solves 73 percent of tasks at least once in five attempts, practically even with Opus 5 at 74.

Why it matters

In DeepSeek's own report, R1 beat o1 on individual tests like AIME 2024 but trailed clearly on others, such as factual knowledge (SimpleQA). Later benchmarks exposed more gaps. Chinese models only reached the top in individual disciplines, not across the board. 2 still showed the same pattern. 3. Chinese models now sit near the top of almost every broad, demanding evaluation.

They handle long knowledge tasks, code across many steps, and coordinate tools far more reliably than their predecessors. Measured by common benchmarks, the often-cited gap of a few months has shrunk enough to become an investor problem. According to the Wall Street Journal, Anthropic is fielding uncomfortable questions ahead of its upcoming IPO and points to its remaining lead at the top in its defense.

Below that tier, the field belongs largely to open, far cheaper models from China. Investors worry that raw model performance can barely carry a business anymore. Whatever a model can do exclusively today, a freely downloadable one can do a few months later. Two accusations are in play: Chinese labs allegedly tapped Western models as teachers, a practice known as distillation.

And they allegedly tune their models for strong benchmark scores without the broad capabilities to match - so-called benchmaxxing. The American lead hasn't disappeared. But it has retreated to a few, ever-narrower areas of the so-called frontier, the leading edge of what's technically possible. Where that lead still sits, and what it's worth economically, is the first question of this issue. The second follows from it.

If the model alone can't make a clear difference anymore, what does the lead rest on? Our thesis: less and less on the model, and more and more on the overall system around it, meaning the system where ongoing work produces the next models. 8. K3 improved most on agentic tasks, so exactly the hands-on work that matters in enterprise use.

On AutomationBench-AA, K3 even debuted in first place until Anthropic answered with Opus 5. 15 million. 7, had regularly failed these long hauls. 8-Max reaches a similarly high overall level. One caveat remains: Newer Chinese models sometimes burn far more tokens than Western ones, which eats into part of their price advantage. The cost-per-task math still applies.

What to watch

Reliability doesn't follow the usual intelligence-index rankings either. 6 Sol falls behind its own predecessor here. Commercially, reliability weighs heavily, as the operators explain in their launch article. An analyst agent only saves work when its answers hold up without review. Multiple runs and checkers can compensate, but they drive up the cost per accepted result. The third is cybersecurity. Here the gap is best documented.

A joint assessment by the UK's AISI and the US CAISI found that K3 lags far behind leading US models in offensive cyber capabilities. On ExploitBench, a test for developing exploits, K3 scored 32 percent. The top US models averaged about 76. K3 failed all 41 tasks that required executing code on a target system; leading US models solved 20 on average.