Key takeaways

  • On ulam.ai's ErdosBench, Astra took first place (update: Fable 5.1 has since retaken it with fewer solved problems but better overall…
  • 1 has since retaken it with fewer solved problems but better overall performance).
  • In some cases, it actually understated its own results.

What happened

1 has since retaken it with fewer solved problems but better overall performance). The benchmark covers 226 open math problems inspired by the famous Erdős problems. 23 and solved 106 problems, 43 of them completely. It also disproved 27 others. Compared to Sol, which solved 78 problems at maximum reasoning, Astra showed stronger scientific writing and was less prone to overblown claims.

Pachocki's point that OpenAI deliberately skipped math optimization for Astra is evidence for the "spiky" thesis. Even the leading AI lab can't push maximum progress in every direction at once. Capabilities grow where you optimize, and making math stronger means cutting back somewhere else. Behind the RSI priority, though, is the hope that the model will eventually make those optimizations itself, scaling faster across the board, including in math.

Regardless of whether Astra is 5 or 50 percent better than its predecessor, the deeper question remains for a field that's grappling with a new world of compute power that helps solve problems but doesn't necessarily help understand them.

Mathematician Terence Tao raised it at the 2026 International Congress of Mathematicians: if AI models keep producing proofs faster than humans can check them, the field risks shifting from proof scarcity to proof overload. The critical task would then no longer be solving problems but deciding which results actually matter.

Why it matters

In some cases, it actually understated its own results. Overall, benchmark developer Przemek Chojecki called it "a solid 5%-10% gain on various math-research skills tested," but the benchmark is far from saturated. OpenAI could have made Astra much stronger in math research but decided against it, even though the company had put math wins front and center in its first announcement.

" That means the most capable math model right now isn't the result of targeted optimization. It's a byproduct of other priorities. OpenAI is pouring its resources into recursive self-improvement and securing future AI systems, since "we believe it is the only way to remain at the frontier of AI research moving forward," Pachocki writes.

The alternative scenario describes an increasingly "spiky" trajectory, with extreme strength in a few domains like coding and math but stagnation or even regression in areas like language quality, common sense, or social reasoning. That would give us a highly specialized model, not something most people would call Artificial General Intelligence. Of course, you can call anything AGI if you feel like it.

What to watch

Tao says math faces a crisis of its values and practices, one he compares to the foundational upheaval of the early 20th century. The hardest problems in math remain unsolved for now, which should buy the discipline some time. The direction still seems set, though, even if some mathematicians don't think language models can deliver real breakthroughs because they lack human-like creativity.

For similar reasons, there are also doubts about AI's potential for genuine self-improvement.