Key takeaways
- Progress in AI compounds fastest when the entire system improves together.
- More capable models unlock better products, which generate more demand, usage, and learning.
- It also performed strongly on DeepSeek R1 and Kimi K2, showing that its gains extend across model families.
What happened
Progress in AI compounds fastest when the entire system improves together. That is how I think about OpenAI’s compute strategy: one integrated system spanning data centers and chips, frontier models, our developer platform, consumer and enterprise products, and AI-native devices, with each layer strengthening the next. Better software makes hardware more productive. Hardware designed for our workloads improves speed and efficiency.
Preserving credible choice across providers, hardware, and deployment models lets us direct demand toward the strongest performance per dollar, maintain pricing discipline as market conditions change, and move with the frontier as stronger technology emerges. Direct control adds leverage where tighter integration can improve the entire system. We partner where the ecosystem helps us move faster and build where co-design creates a meaningful advantage.
Data centers create another point of leverage. Project Camellia in Georgia shows how we can design facilities around customer workloads while creating jobs, supporting local businesses, covering project infrastructure and energy costs, conserving water through a closed-loop system, and subjecting its commitments to an annual independent public audit. The value of this system is measured by what it produces: more useful intelligence from every unit of compute.
Better models reach the right answer with fewer attempts. Smarter routing and context management reduce wasted work. Optimized software and purpose-built hardware improve speed and energy efficiency. 6 Sol with max reasoning reached a new high while using 54% fewer output tokens than another leading model.
For customers, improvements like these mean faster results, more dependable products, fewer retries, agents that complete longer workflows, and a lower total cost for successful work. The best economics come from useful intelligence per dollar. As useful intelligence becomes more capable and affordable, more work becomes economically practical. A company can provide tailored analysis to every customer, review every contract, run live financial scenarios, and help engineers test more ideas.
This is Jevons paradox: greater efficiency makes more uses worthwhile, expanding consumption and creating new economic activity through more work completed, better decisions, more products launched, and more revenue generated. More productive compute and a more competitive supply base help us serve more customers at lower cost and carry those efficiency gains through to users. Growth funds continued investment in research, infrastructure, and safety.
Why it matters
More capable models unlock better products, which generate more demand, usage, and learning. Those signals flow back through the system and help us improve it again. Today, we shared the first measured performance results from Jalapeño, OpenAI’s first custom inference chip. On InferenceX, a public benchmark using GPT‑OSS 120B, Jalapeño delivered more peak throughput per kilowatt and lower token latency than the commercial systems in the comparison.
It also performed strongly on DeepSeek R1 and Kimi K2, showing that its gains extend across model families. Jalapeño gives us greater control over how our models run and over the economics of serving them. By developing the model, serving software, chip, memory, and network together, we can improve throughput, latency, energy efficiency, and cost as one system.
It creates a credible first-party path alongside the accelerators we use from other partners, expanding our ability to match each workload to the strongest system at the right economics. We now have working first-party silicon with measured results, and future generations are already underway. Different workloads place different demands on the system. Frontier training, high-volume inference, and always-on agents have different requirements across chips, software, networks, power, and latency.
Our goal is to stay on the Pareto frontier: continually seeking the strongest mix of capability, speed, reliability, efficiency, and cost for each workload. Different chips and providers lead on different dimensions, and the frontier keeps moving. Our portfolio gives us the range to meet those needs. Microsoft’s compute and NVIDIA’s chips have been foundational to OpenAI’s growth.
Today, our portfolio also includes AWS, AMD, Broadcom, Cerebras, CoreWeave, Oracle, SB Energy and SoftBank. Each brings different strengths across cloud infrastructure, accelerated computing, low-latency inference, data-center development, and energy delivery. We actively manage this portfolio for both capability and economics. We use premium systems where capability matters most and optimize for efficiency where scale and cost matter more.
What to watch
That is OpenAI’s compounding advantage: better technology creates better economics, better economics fund the next wave of progress, and every gain makes the whole system stronger.




