Citation Bureau
XV SEPTEMBER MMXXVI
· 3 min read · Vol. I · No. 373

Hyperscalers are spending half their capital budgets on memory, and the shortage will not ease before 2028

Hyperscalers are allocating 50 percent of capital expenditure to memory while supply runs 30 percent short of demand. Prices are rising, consumer markets are being crowded out, and the structural constraint on AI inference is bandwidth, not compute.

Hyperscalers are spending half their capital budgets on memory. That figure, reported by Dwarkesh Patel, is not a forecast: it describes this year’s allocation. It sits alongside a separate observation from Patrick O’Shaughnessy that supply across memory markets is already running about 30 percent short of demand. Together, those two data points frame a shortage that is acute now and, by most accounts, will not ease for years.

The capital commitment is extraordinary on its own terms. Martin Casado, a general partner at Andreessen Horowitz, notes that Microsoft, Meta, and Google are each on track to spend over 50 percent of their revenue on capital expenditure this year. Memory has gone from a line item to the dominant cost in building AI infrastructure, and the transition happened within a single budget cycle.

Prices are moving accordingly. Andrew Feldman points to Micron producing numbers with 80 to 85 percent gross margins. Nathan Labenz adds that memory chip makers are operating at 80 to 90 percent gross margins, above even Nvidia’s roughly 70 percent. Caitlin Kalinowski, who has been advising startups to pre-buy memory and hold enough stock to ride out price spikes, puts it plainly: prices are probably going to double. Jake Cooper has already seen it on the balance sheet: the servers his company runs have appreciated in value as RAM prices rise.

Dylan Patel supplies the timeline that explains why none of this resolves quickly. Despite capacity growing 20 to 30 percent annually, true incremental supply from new demand signals will not arrive until 2028. Rene Haas frames the same horizon from the supply chain side, expecting the constrained environment to persist for at least three to five years. The shortage is not a momentary demand spike waiting for fabricators to catch up. It is a structural condition with a multi-year runway.

Going massively beyond that would be cost-prohibitive. Not because of the compute cost, because of the memory bandwidth. Reiner Pope

The scarcity is already crowding out consumer hardware. Patel expects smartphone volumes to fall 30 percent because there is not enough memory to go around. When the same physical substrate serves both a frontier AI cluster and a mid-range handset, and the cluster operator has the capital to outbid every consumer electronics buyer in the market, the consumer market loses.

What drives the demand is the nature of AI inference itself. Kunle Olukotun, a computer architecture researcher, describes the inference problem as fundamentally one of data movement: as models grow larger, the weights and what is called the KV cache must be moved into compute units continuously, and that movement is the bottleneck. Labenz states the conclusion directly: the AI hardware constraint is memory and bandwidth, not compute. Reiner Pope, a researcher who works on large-scale model deployment, points to a specific empirical signal. Context lengths in production models shot up from around 8,000 tokens to between 100,000 and 200,000 tokens in an earlier transition, and then stalled. They have hovered in that range for the past year or two. Pope’s explanation: “going massively beyond that would be cost-prohibitive. Not because of the compute cost, because of the memory bandwidth.” The plateau in context length is a direct artifact of the memory ceiling.

Pope also notes that Blackwell, Nvidia’s most recent GPU generation, finally delivers scale-up memory on the order of 10 to 20 terabytes, enough to hold a five-trillion-parameter model plus KV cache. That matters, but it does not dissolve the bandwidth constraint. Holding the weights is a necessary condition, not a sufficient one. Sean Lie observes that running a frontier-level model of a few trillion parameters would require thousands of specialized chips just to hold the weights. The compute exists. The memory infrastructure to run models efficiently, at scale, at bandwidth, does not yet exist in sufficient quantity.

The shortage has also stretched beyond silicon. Yaroslav Azhnyuk reports that optic fiber prices rose from roughly $4 per kilometer to around $32 per kilometer in a matter of months earlier this year, a sign that scarcity is propagating through the full infrastructure stack. The direction of events confirms what the supply-side data already shows: demand is consistently outrunning what buyers anticipated, and the gap between what AI systems require and what the market can deliver is not narrowing on any near-term schedule.

The Editor, for the readers of Citation Bureau

AI Hardware DemandData Center CapexMemory ChipsSemiconductorsSupply Chains



From the Archive