Key takeaways
- As we enter the next era, what will be the defining measure of our progress?
- For more than 60 years, the semiconductor industry has asked that question, relentlessly maximizing the number of usable chips produced…
- Today, we need to apply the same principle to the unprecedented resources that the world is pouring into AI: capital on a scale once…
What happened
As we enter the next era, what will be the defining measure of our progress? Every industry has a word that shapes how it thinks. For pilots, it’s safety. For insurers, it’s risk. For the semiconductor industry, it’s yield. Yield does not ask how elegant the solution is, how many years it took or what the roadmap promised. Rather, it asks one simple question: What useful output did we produce?
For more than 60 years, the semiconductor industry has asked that question, relentlessly maximizing the number of usable chips produced from every wafer. Generation after generation, wafer after wafer, it is precisely that discipline that turned the transistor from a laboratory curiosity into the foundation of modern life.
Today, we need to apply the same principle to the unprecedented resources that the world is pouring into AI: capital on a scale once reserved for nations, gigawatts of power and record-breaking fabs and datacenters. The question that will define this decade is the same one this industry has always asked: What actually comes out?
AI has proliferated with remarkable speed, at a rate of adoption faster than the internet, the PC or even the smartphone. Yet, global penetration still stands at just 18% of the working population, and the vast majority of that usage is chat-based. As systems move from answering individual prompts to reasoning, planning, using tools and executing longer agentic workflows, the infrastructure equation changes dramatically.
A single agentic task can use more than 3,400 times as many tokens as a typical chat interaction. We are only in the early innings of agentic adoption, and the infrastructure is already strained. Power is setting the limits on what we can build and when. Packages and racks are growing larger and denser. Memory is becoming an even tighter constraint.
For years, the industry’s rational answer to each new requirement resulted in more: more silicon in the package, more memory beside it, more power to feed it and more fiber to connect it. Each generation delivered meaningful progress. But when each new gain requires more input than the one before it, we are on a treadmill. It moves only as long as we keep adding to it.
I believe we need to pursue two paths forward. The first is evolutionary: we continue improving the architectures we have today, driving incremental efficiency, utilization and economics within each generation. The second is transformational: changing the curve itself with innovation in new architectures, new materials and new approaches to system and model design. The history of our industry is defined by transformations like these.
When increasing CPU clock speeds ran into the power wall, we moved to multicore processors. When planar NAND reached its limits, memory went vertical. And now, once again, we have an opportunity to challenge our assumptions and rethink the fundamentals. Because the next chapter of AI won’t be defined simply by how much infrastructure we build, it will be defined by how much intelligence we can create from it.
For decades, the computing industry has optimized yield in the context of manufacturing. Today, that discipline has to extend across layers, from datacenters and silicon through models and the agentic harnesses that orchestrate them. And the work does not stop once the technology is built. We must then deploy and scale it faster, while developing new tools and systems to maximize utilization across our fleet.
From development through execution, each layer has a yield of its own, and losses and gains compound across them. Capacity at any one layer is only a starting point. The real measure is how effectively those layers work together to produce useful output from the system as a whole. Our experience at Microsoft building and operating AI infrastructure at scale has reinforced two key lessons.
First, the biggest constraints are rarely solved in the layer where they appear. Second, when we attack a constraint across the whole stack, tradeoffs that seemed inherent to the problem often turn out to be artifacts of the architecture. The greatest advances often come when we apply these learnings through co-design, working across layers to turn apparent limits into solvable system constraints.
Our experience building the Azure Maia platform demonstrates that memory bottlenecks are not resolved by a single layer. Model architecture, data science and compression can reduce the amount of KV cache, which stores the model’s working context during generation.
Why it matters
Innovations in memory, networking and power show what this approach looks like in practice. Today, memory is viewed as a supply problem or a component problem. In reality, it is a system problem. In AI inference, memory is now setting the limits on system performance. It must hold larger models, preserve longer contexts and deliver data fast enough to keep the compute fed. And agents raise the bar even further.
What to watch
Generation, retrieval, tool use and persistent memory run together in loops that can last minutes or hours. The result is a much longer memory horizon, with far more information kept close to the compute and available across an expanding sequence of turns. Doing that efficiently at scale will define the next generation of AI infrastructure.




