Key takeaways
- Anthropic's new flagship model Claude Opus 5 posts top scores in coding and knowledge work while approaching Claude Fable 5's performance…
- The 1 million-token context window and token rates remain unchanged.
- In its prompting guide, Anthropic recommends making broad use of the "low" and "medium" settings.
What happened
Anthropic's new flagship model Claude Opus 5 posts top scores in coding and knowledge work while approaching Claude Fable 5's performance at half its token rates. 6 Sol and Chinese competitors. Its new Opus 5 model is designed to close the price-performance gap with the much pricier Fable 5. Opus 5 becomes the default model on Claude Max and the most capable model available on Claude Pro.
In a Frontier-Bench task, Opus 5 received a drawing of a machine part and had to create a 3D model in FreeCAD. The catch was that the model intentionally had no way to view the drawing directly. Opus 5 responded by writing its own computer vision pipeline to extract the geometry from raw pixels, then reconstructed the complete machine part.
No other model solved this task after five attempts, Anthropic says. Opus 5 also worked on a real bug in a popular open-source package manager. According to Anthropic, it found the root cause and fixed an edge case that the community patch had missed. A competing model fixed only the surface symptom before reporting the bug as resolved.
Why it matters
The 1 million-token context window and token rates remain unchanged. 8, Opus 5 costs $5 per million input tokens and $25 per million output tokens. 5x but doubles the price. Users can trade off performance against token use through five effort settings called low, medium, high, xhigh, and max. Anthropic says Opus 5 offers better value than its predecessor at every effort level.
In its prompting guide, Anthropic recommends making broad use of the "low" and "medium" settings. The company says they deliver good results with a fraction of the token use and latency while beating the same settings on earlier Opus models. " Opus 5 scores slightly worse at the max effort setting than at the second-highest setting on two benchmarks, despite costing more. 1 and the Artificial Analysis Coding Agent Index.
According to Anthropic's own benchmarks, Opus 5 sets records across several evaluations. 1 percent) by wide margins. 6 Sol (1,736). Opus 5 doesn't win everywhere. 8 percent). On health tasks and legal benchmarks, Fable 5 and Mythos 5 outperform Opus 5, respectively. The ARC-AGI-3 result is likely the biggest surprise and outlier in Anthropic's benchmarks. 2 percent. 8 percent. That's nearly a 4x gap over the next-best model.
There's no Fable 5 result for this benchmark, and it's unclear whether such a large lead on a test will show up in actual use. Opus 5 falls behind Mythos 5 on cybersecurity tasks. Anthropic says it deliberately didn't train Opus 5 on cyber tasks, as was also the case with its predecessor.
The model comes close to Mythos 5 at finding vulnerabilities but performs much worse when asked to exploit them. Anthropic also says Opus 5 has improved at generating visual outputs and analyzing visual content such as charts and diagrams. Anthropic describes Opus 5 as much better at checking its own work and improving it through iteration.
What to watch
An engineer at a trading firm reportedly used Opus 5 to build a market data feed for a new exchange in one session. Previous models couldn't complete the task, even with detailed plans. Opus 5's safety setup allows source code vulnerability research but blocks binary-based vulnerability scanning, penetration testing, and exploit generation, according to Anthropic. The cyber classifiers trigger about 85 percent less often than on Fable 5.
Fable 5's frequent interventions drew heavy criticism. 8 as a fallback, the same approach used with Fable 5. Anthropic calls Opus 5 the most capable generally available model for scientific research. 7 percentage points). Alongside Opus 5, Anthropic is releasing two beta features. Mid-Conversation Tool Changes on the Claude Platform let developers swap available tools during a conversation without invalidating the prompt cache. Automatic Fallbacks on the API route blocked requests to a different model automatically.

