Key takeaways

  • Anthropic's Claude Opus 5 is the most capable AI model available today, according to several benchmarks, outperforming Fable 5 while…
  • In coding, Claude Opus 5 at "xhigh" paired with Claude Code shares first place on the Artificial Analysis Coding Index, which measures how…
  • No single model can pull away or claim a clear advantage.

What happened

Anthropic's Claude Opus 5 is the most capable AI model available today, according to several benchmarks, outperforming Fable 5 while costing less. Opus 5 scored 61 on the Artificial Analysis Intelligence Index, which combines nine tests covering knowledge work, coding, scientific reasoning, and factual accuracy. 8 (56). Artificial Analysis worked with Anthropic to test the model before its public release.

That matches Anthropic's guidance, with "high" set as the default tier in both the API and Claude Code. Token pricing stays at $5 per million input tokens and $25 per million output tokens. 50 per million tokens.

Opus 5 performs especially well on the AA-Briefcase benchmark, which measures how well AI models handle typical office tasks like writing research reports, building presentations, and analyzing spreadsheets based on thousands of input files. Performance is scored across correctness, analytical quality, and presentation quality, then rolled into an Elo rating similar to chess rankings.

Opus 5's biggest gains show up in analytical quality. At "max," it reaches an Analytical Quality Elo of 2016, almost 300 points ahead of Fable 5. 2 percent ("xhigh"), and 56 percent ("high"). Presentation quality is a different story. 6 Sol at "max" (1666). 8 at 24 minutes and 55 passes.

Why it matters

In coding, Claude Opus 5 at "xhigh" paired with Claude Code shares first place on the Artificial Analysis Coding Index, which measures how well AI models handle programming tasks on their own, including finding and fixing bugs. 6 Sol. For scientific reasoning, Opus 5 scored 53 percent on Humanity's Last Exam, a very difficult knowledge test covering many academic fields. That ties it with Fable 5. 6 Terra.

8 but still trails Fable 5. Opus 5 also answers more often when it's uncertain, pushing its hallucination rate up 14 points to 50 percent. Epoch AI has also tested Claude Opus 5. The research institute gave it an overall Epoch Capability Index score of 159, just below Fable 5 at 161. 8. 6 Sol leads in both categories. Overall, this confirms that the race among frontier models is tight.

No single model can pull away or claim a clear advantage. That lends weight to the argument that AI models will eventually become commoditized. 75. 53. 8 and Sonnet 5 while costing less. ai tested Claude Opus 5 across all five reasoning tiers using Vibe Code Bench, a benchmark for programming tasks. 4 percent despite much higher costs.

ai found that the highest tiers tend to produce more complex solutions that contain errors more often. The "high" tier produces simpler solutions that meet the requirements more reliably. 1 shows a similar pattern. The "high" tier beats "max" because the model spends more time on each attempt at the top tier, leaving fewer attempts within the time limit.

What to watch

At max reasoning, Opus 5 reaches an Elo of 1720, a full 146 points ahead of Claude Fable 5 (1574). Its three highest tiers (max, xhigh, high) sweep the top three spots. 8, Anthropic now holds the vast majority of top-10 positions. 30 for Fable 5. 41, less than half of Fable 5. Both still beat Fable 5 in the Elo ranking. 6 Sol (max, 1505). 2 (max, 1254).