Inference clusters are scaling to tens of thousands of chips faster than the infrastructure to support them can follow
Gavin Uberti sees inference moving from eight-chip configurations to thousands of chips, and fast. The demand figures, capital commitments, and supply constraints already in evidence suggest the gap between what is being built and what is needed will widen before it narrows.
Gavin Uberti, who works on AI hardware, puts the infrastructure shift in plain terms. The inference side, he says, is moving from eight-chip clusters or NVL72 scaleup domains to thousands and tens of thousands of chips, and it will happen quickly. That framing, which might have sounded speculative two years ago, now has corroboration across both the demand and supply sides of the industry.
Nathan Labenz describes what that trajectory looks like at the frontier today. “Some of my friends at labs right now are training models across data centers. It’s not even across nodes anymore. It’s not even racks.” The unit of compute has outgrown the rack, then the building, and is now measured in campuses. Harry Stebbings adds a specific data point from the model side: Jensen Huang credited OpenAI’s latest model as having been trained on more than 100,000 Nvidia chips, with 400,000 more coming.
The demand numbers driving this expansion are steep. Ben Horowitz, a co-founder of Andreessen Horowitz, estimates token demand is growing close to 1,000 percent per year. Lin Qiao, who leads inference infrastructure work, puts the range for daily token volume growth at 20 to 100 times by the end of next year. Simon Mo reports that vLLM, the open-source inference engine, is already running on half a million GPUs at any given moment. These are not projections drawn from surveys. They are operational figures from systems already in production.
Some of my friends at labs right now are training models across data centers. It's not even across nodes anymore. It's not even racks. Nathan Labenz
The capital commitment trying to meet that demand is itself substantial. Stebbings estimates a trillion dollars of capital expenditure has already been committed across major players for the coming year alone. Horowitz notes that every GPU being manufactured is already pre-sold. Nvidia, for its part, guided to 70 percent revenue growth for its fiscal year ending January 2028, well above the 44 percent the market had expected, according to Stebbings.
Supply, however, faces constraints that capital alone cannot resolve quickly. Rene Haas, who leads Arm Holdings, is direct: “I think we’re going to be in this constrained environment for 3 to 5 years at least.” Data center construction itself, he adds, will become the next major bottleneck. On power, Horowitz cites a stark imbalance: by 2028, new data centers will need something like 44 gigawatts of additional power against perhaps 25 gigawatts of expected grid additions. Demand, he says, is growing 10 times per year, a pace that grid infrastructure cannot match on any near-term timeline.
Hardware obsolescence compounds the pressure on operators trying to plan multi-year buildouts. Qiao frames the lifecycle problem this way: if three new hardware generations arrive every year, hardware that is three years old sits nine generations behind current capability. Running a three-year-old model on that hardware, she says, is questionable. That pace of generational turnover makes long-horizon capital planning genuinely difficult, because the asset being purchased today may be economically obsolete before it is fully depreciated.
What emerges from these data points is a supply-demand gap that money has been committed to close but that physical and logistical constraints will prevent from closing on the timeline demand requires. Wafer capacity, advanced packaging, memory production, power infrastructure, and physical construction are all constrained simultaneously, and none of them respond to capital injection within a single budget cycle. Rory O’Driscoll, a venture investor, draws the structural conclusion: inference providers will become far more capital-intensive businesses as the scale of the underlying infrastructure grows. The companies best positioned in that environment are those that secured compute commitments early. Patrick O’Shaughnessy notes, for instance, that Google has already made a deal to sell roughly 20 percent of its TPU capacity to Anthropic. Reservation, not spot purchasing, is the game being played at the frontier. The cluster sizes Uberti describes are not a distant endpoint. They are the near-term destination, arriving faster than the infrastructure needed to support them.