Key takeaways
- Astra is clearly more expensive than its own predecessor.
- The reason is how sparing the model is.
- At the same time, the model loses about 80 Elo points on GDPval-AA v2 and slips on banking support, SciCode, and long-context reasoning…
What happened
Astra is clearly more expensive than its own predecessor. OpenAI charges two and a half times as much per unit of processed text, which makes a task cost roughly 75 percent more than it did with Sol. Compared with Anthropic, the picture flips. On coding tasks, Astra hits the same score as Claude Fable 5 according to Artificial Analysis, but costs less than half as much per task.
According to ARC Prize, the reason is that Astra solves the games in fewer moves, which means fewer model calls and fewer tokens. 5 percent, worse than running with no reasoning at all. ARC Prize doesn't comment on the outlier, but GPT-6 Astra in other benchmarks showed that it can solve longer-horizon tasks without reasoning due to its new architecture which presumably loops processing internally before generating the first token.
78 per game. But that mostly pays for time and willingness to take part. 067 cents per game. More interesting than the raw score is the efficiency. Before the launch, ARC Prize had about 500 testers play with no pre-screening and recorded, for each level, the median number of moves among those who solved it.
On the run with the OpenAI scaffold, Astra cleared 96 percent of levels in fewer moves than that median, on average with a little over half. Unlike the usual cost measures, this figure doesn't track compute consumed. It tracks how much experience with an environment the model needed before it mastered it. This is exactly where the organizers had expected humans to hold a lasting edge.
To get there, Astra keeps its own notes and works out a self-invented, algebra-like shorthand in which it records objects, coordinates, rules, and open plans, for example extend8 to3; retract10 to2 as an ordered sequence of moves or Turn 5: P=(24,20), empty, facing west as a state note. ARC Prize saw similar behavior from other models, but singles out Astra for its precision and information density.
On the standard harness, that's an important skill, because everything the model doesn't save into its own visible notes is lost.
Why it matters
The reason is how sparing the model is. It needs only a third of the compute steps Sol uses and a fifth of what Opus 5 uses. 5 percent. ARC-AGI-1 is now considered largely saturated. 1 leads with 70. The hallucination rate on AA-Omniscience drops from 92 to 51 percent.
At the same time, the model loses about 80 Elo points on GDPval-AA v2 and slips on banking support, SciCode, and long-context reasoning tasks. Epoch AI reports that on the new FrontierMath Erdős, GPT-6 Astra was the only model to solve two of 68 open Erdős problems with Lean-verified proofs, on a budget of $300 per attempt.
Three more solutions came out of non-standardized extra runs that burned through more than $220,000 in compute, but Epoch says those don't count toward the score. 1 have so far been measured on different ground. Astra leads on math, knowledge, and puzzles. 1 holds the top marks on nearly every coding test.
But Epoch has recorded only a single coding score for Astra so far, and that one comes from a run at a medium reasoning level. The clearest jump comes on ARC-AGI-3. The test drops an AI into unfamiliar game worlds whose rules and goals nobody explains to it. The model has to figure out what to do by trial and error. 7 percent at a test cost of roughly $26,000.
16 percent. 1 aren't on the benchmark yet. 9 percent OpenAI reported came under different conditions. In that setup Astra got to use the harness OpenAI built, which keeps reasoning chains between individual requests and automatically summarizes long runs. 66 times faster and used 49 percent fewer tokens than runs on the in-house harness, compared across 167 game-reasoning pairs that both setups solved. 6 Sol.
7 percent, run on the internal ARC harness, allowed a fair comparison between vendors, though it plans to publish the numbers from vendor harnesses in the future as well. The relationship between thinking effort and cost is unusual. Normally a higher reasoning level makes a test run more expensive. With Astra it's the opposite. 7 percent.
What to watch
That still holds for brute-force approaches, but with top models ARC Prize sees an almost binary pattern. Once the model has figured out the mechanics, its execution lands in the human efficiency range.

.gif)