Key takeaways

  • Meta has released Muse Spark 1.3, its fourth model in five months.
  • 3, its fourth model in five months.
  • 1 for 16 percent, and τ³-Bench Banking for 14 percent, and Meta's biggest gains land in exactly those three tests.

What happened

3, its fourth model in five months. Independent testing shows solid gains on agentic tasks, but a gap to the top stays. What sells the model is the price. 55. 23. 40. On the Intelligence Index, max scores 62 points and xhigh 61, up from 57 in August and 53 in July. The jump comes down to how the index weights its tests.

Why it matters

1 for 16 percent, and τ³-Bench Banking for 14 percent, and Meta's biggest gains land in exactly those three tests. On τ³-Banking, where agents operate tools in a simulated banking scenario, max hits 52 percent. That's number one right now, according to Artificial Analysis, and it's the only outright lead the model holds. 3-Flash rather than leading. 2 sat at 35 percent. 9 at high.

On the index's highest-weighted test, GDPval-AA v2, Meta improves from 1,615 to 1,709 and 1,754 on a scale calibrated to human expert performance at 1,000 across 220 real-world professional tasks. 1 (max) sits at 1,853. Meta buys the max variant's edge with compute, burning 62 percent more reasoning tokens than xhigh. On GPQA Diamond, which poses expert-level science questions, Muse Spark rises from 90 to 94 percent. 9. 1 percent.

What to watch

2. AA-LCR falls from 83 to 79 percent. Factual accuracy in AA-Omniscience slips by up to three points, because the model more often declines to answer when it's unsure. Neither Meta nor Artificial Analysis has named a price for the max variant yet. Larger models and an open-weights release are on the way.