Source: raw/newsletter-epoch-ai-add82097a6.md (“Is a compute crunch coming?”, Luke Emberson & Jaime Sevilla) · raw/newsletter-epoch-ai-e05a84c055.md (“Frontier labs don’t use most AI compute (yet)”, Josh You) · raw/newsletter-epoch-ai-a6ad3f7054.md (“Diversion and resale: estimating compute smuggling to China”, Isabel Juniewicz) · raw/newsletter-epoch-ai-6fb94702d7.md (“What we learned from 1,604 Chinese AI job postings”, Cheryl Wu, JS Denain, Anson Ho) · Epoch Brief data-insights: raw/newsletter-epoch-ai-61525dda28.md (June 1), raw/newsletter-epoch-ai-66b49f0f8d.md (May 22), raw/newsletter-epoch-ai-45df197d73.md (May 15), raw/newsletter-epoch-ai-ac7e42b77d.md (June 12), raw/newsletter-epoch-ai-b3167bb772.md (June 26), raw/newsletter-epoch-ai-b9855d3cc7.md (May 8) Publisher: Epoch AI. The three lead pieces are Gradient Updates — Epoch’s explicitly opinionated/informal series representing the named authors’ views, not an Epoch institutional position. The smuggling piece excerpts a fuller Epoch report; the bracketed Data Insights are Epoch’s standard (non-opinion) data work.

A consolidation of Epoch AI’s mid-2026 quantitative work on the cost and physical limits of the AI buildout, distilled for the one question a practitioner actually cares about: what does this imply for the price and availability of frontier intelligence? The through-line across all of it is a squeeze — token demand appears to be outrunning inference supply, the two labs driving demand (Anthropic and OpenAI) are on track to absorb the world’s compute headroom within a few years, and the physical inputs (memory, data centers, chips) are the binding constraints. This is capability-and-cost forecasting, not marketing advice — but it grounds decisions about model choice and when to expect frontier access to get more expensive.

Key Takeaways

  • A compute crunch is plausibly near — and it hits long-context agentic workloads first. Emberson & Sevilla model that today’s Blackwell GPUs could serve 500 million–20 billion output tokens/second, with global inference capacity more than tripling each year (~3–4×). But token demand — estimated at 200 million–4 billion tokens/second at current prices — appears to be growing ~10× per year, plausibly outpacing supply already. ^[The supply and demand ranges are wide and the authors flag them as highly uncertain.]
  • The practical consequence: frontier long-context access gets pricier; everyday users get pushed to smaller/cheaper models. Epoch’s own read — “the price of access to frontier capabilities [rises] for those willing to pay, while everyday users shift to cheaper, smaller models.” Because inference/training efficiency keeps improving fast, this need not mean worse quality: tomorrow’s smaller, cheaper models quickly match today’s frontier. This is the cost lever behind the wiki’s cheap-workhorse-vs-daily-driver split.
  • Frontier labs don’t yet use most AI compute — but Anthropic and OpenAI are on track to. Josh You estimates the five most compute-rich developers (OpenAI, Anthropic, xAI, Google DeepMind, Meta) together held under half of world compute at end-2025; OpenAI alone was ~10–15%. Yet OpenAI and Anthropic are growing compute ~4×/year vs the industry’s ~3× — enough to consume the headroom and reach ~80% of world compute within ~5 years, at which point scaling stalls unless total capex accelerates dramatically.
  • The buildout is already ~1T annualized in 2026 (~1% of gross world product, ~3% of US GDP); hyperscaler capex has quadrupled since GPT-4 and is on track to overtake hyperscalers’ operating cash flows by end of 2026, forcing external financing. GPU-hour rental prices rose ~30% this year — a supply-tightness signal you can already feel in API pricing.
  • Memory, not logic, is the bottleneck. High-bandwidth memory (HBM) grew from 52% to 63% of AI-chip component cost (Q1 2024→Q4 2025), spend rising ~32B — faster than any other component. Servers account for 60% of the total cost of owning a 1 GW data center (8.5B/year), dwarfing energy.
  • China’s compute is heavily smuggled, and Chinese labs lean on domestic chips only for the easy jobs. Epoch estimates 660,000 H100-equivalents smuggled into China through 2025 (90% CI 290k–1.6M) — roughly a quarter to a third of China’s total AI compute. Separately, Chinese job postings suggest domestic chips are used frequently for inference, rarely for pretraining large models (GLM-Image, at 16B params, was trained entirely on domestic chips — but that is 10–100× smaller than their frontier models).
  • Open-weight models trail the closed frontier by ~4 months in Epoch’s capability index (an 8-point gap, ≈ the GPT-5→GPT-5.5 difference) — the durable lag to price into any open-weight routing decision.

1. Is a compute crunch coming? (the demand/supply squeeze)

Gradient Update by Luke Emberson & Jaime Sevilla, 2026-05-26. The authors model the supply side precisely (how many tokens the world’s chips could serve) and compare it to imperfect demand proxies.

  • Supply. Serving a Kimi K2.6-class open model on all the world’s Nvidia GB200 + GB300 chips (1.9M and 1.5M GPUs at Q4 2025, ~40% of aggregate FLOP/s), and accounting for real-world inefficiencies (calibrated against SemiAnalysis’s InferenceX data → ~65% compute / 30% bandwidth efficiency), yields ~400,000 tok/s per GB200 NVL72 rack and a global total of 500 million–20 billion output tokens/second — depending heavily on request context length. That’s 150,000–7 million tokens per month for every person on earth. Capacity is more than tripling each year.
  • Demand. Proxies (Google’s ~1.2B tok/s across its platforms ≈ 130M output tok/s; Exponential View’s ~40 quadrillion tokens/quarter ≈ 5B tok/s industry-wide) suggest current demand of 200M–4B tokens/second, growing ~10×/year — a faster slope than supply’s ~3–4×.
  • The crunch and who feels it. If these slopes hold, demand overtakes supply soon (if not already), “particularly for the long-context workloads that drive agentic AI.” Long context is the culprit because attention compute grows quadratically — the batch size at which inference becomes compute-bound shrinks as context grows (e.g. from ~870 concurrent users per rack at 8K context down to ~130 at 25K). ^[inferred: the “agentic = long context = hits the compute wall first” causal chain is the article’s argument, restated here as the practical takeaway.]
  • Why quality need not regress. Inference and training efficiency are improving fast enough (Epoch’s own price-trend data) that “the smaller, cheaper models of tomorrow will quickly match today’s frontier.” The crunch reprices frontier access without necessarily degrading what everyday users get.

Practical read: budget for frontier long-context/agentic work to get relatively more expensive over time, and lean on the “this year’s mid-tier ≈ last year’s flagship” migration for routine work. This is the supply-side mechanism underneath Cost & Intelligence Levers.


2. Frontier labs don’t use most AI compute — yet (the consolidation clock)

Gradient Update by Josh You, 2026-05-21. The author flags these lab-compute estimates as more tentative than Epoch’s standard data work.

  • Today’s split. Cumulative AI compute sold ≈ 20M H100-equivalents (H100e) by end-2025; ~16M operational after a one-quarter install lag. OpenAI ≈ 1.7M H100e / 1.9 GW (up from 0.6 GW in 2024, 0.2 GW in 2023). Anthropic ≈ 1.4 GW (~70% of OpenAI), much of it ~500k H100e of Trainium2 at Amazon’s Project Rainier. OpenAI + Anthropic + xAI together < 4M H100e ≈ 20–30% of world compute; Google + Meta own ~1/3 of world compute but allocate perhaps half of that to their frontier labs (~15% of world total).
  • The consolidation clock. OpenAI and Anthropic grew compute ~4×/year in 2025 vs the industry’s ~3×. A naive extrapolation (top-two at 20% share today, growing 33% faster than the world) has them doubling their share in 2.5 years and reaching ~80% within five. Anthropic’s revenue run-rate went 30B annualized in Q1 2026 (an acceleration on last year’s ~10× growth), funding aggressive compute deals — including renting xAI’s entire ~300k-H100e Colossus 1 cluster (up to $15B/year).
  • The wall isn’t capability — it’s capex. Sustaining 4×/year once the headroom is gone would require more than doubling capex annually from a ~$1T-in-2027 base — feasible “only if AI starts to dramatically accelerate economic growth.” Otherwise frontier compute growth converges down to overall compute growth. Epoch is explicit this is not a hard 2029 “wall”: flat capex still grows the stock, chips keep improving. But the compute-scaling driver of progress is “not sustainable unless the world fundamentally changes soon.”

Practical read: the two vendors most likely to tighten supply and raise prices are also the two you most likely depend on. This is the supply-consolidation case for keeping a portable harness and an open-weight fallback (Lever 4).


3. The physical inputs: capex, memory, data centers

Epoch’s (non-opinion) Data Insights quantify the binding constraints:

ConstraintFindingSource
Hyperscaler capexQuadrupled since GPT-4; ~1T in 2027; Q1 2026 came in at 155.1B projection)June 1 brief
Capex vs cash flowOn track to overtake hyperscalers’ operating cash flows by end-2026 → external financing (Alphabet’s $80B raise, Amazon bond offerings)June 26 brief (also MirrorCode’s source)
Memory (HBM)Rose from 52% → 63% of AI-chip component cost (Q1’24→Q4’25); ~32B spend; the dominant cost and primary supply bottleneckMay 22 brief
Data-center servers60% of total cost of a 1 GW data center — 8.5B/year, dwarfing energyMay 15 brief (via May 22 ICYMI)
Single-site scaleRecord for compute at one data center has doubled every 7 months since Colossus 1 (Aug 2024)June 12 brief
Macro footprintAI infra ≈ 1.5% of US GDP (data-center construction alone ~0.8% in Q1 2026), up from a ~0.7% 2015–22 average; now the leading driver of US private-investment growthJune 12 brief
Compute ownershipFive hyperscalers own >2/3 of global AI computeMay 8 brief
Unit economicsAnthropic and OpenAI earn ~5.5M revenue per employee — higher than any public tech company on the Forbes Global 2000May 8 brief

Practical read: HBM is the choke point, so supply relief depends on memory fabs, not just logic nodes — a slower lever. The revenue-per-employee figures are a useful sanity check when sizing what these labs can afford to spend on compute.


4. China: smuggled chips + a domestic-silicon workaround

Two Epoch pieces bear on Chinese compute — relevant for anyone forecasting the open-weight frontier (GLM-5.2, Kimi K3).

  • Smuggling (Isabel Juniewicz, full report). Cumulative allegations of diverted/missing chips total ~300,000 H100e by end-2025 — about a quarter of what China acquired legally. A Monte-Carlo estimate puts true smuggling at 660,000 H100e (90% CI 290k–1.6M)3% of the global compute stockpile (comparable to what xAI held), with the upper bound implying most of China’s compute was smuggled. Context: China imported ~2.5B Supermicro diversion scheme. Cloud compute outside China used by Chinese customers is a separate, generally-permitted channel not counted here.
  • Domestic chips, from 1,604 job postings (Cheryl Wu, JS Denain, Anson Ho). Scraping DeepSeek, MiniMax, Moonshot, Z.ai, ByteDance, and Alibaba postings: Chinese labs still depend on Nvidia/CUDA/TensorRT-LLM for inference, but are actively hiring for Ascend/Cambricon “heterogeneous computing.” Best guess: domestic chips are used frequently for inference, rarely for pretraining large models — the exception being post-training or small models (Z.ai’s GLM-Image, 16B params, was trained entirely on domestic chips, but that’s 10–100× smaller than their frontier models). Other findings worth keeping: Chinese labs have distinct “personalities” (Z.ai is B2B-heavy like Anthropic — 73.7% of 2025 revenue from running models on customers’ infra; MiniMax/Moonshot are consumer/international — 73% of MiniMax revenue is international vs Z.ai’s 9.8%), require far less experience (US labs avg 5.5 years vs 1.6 in China), and cluster in Beijing/Hangzhou/Shanghai (93% of located postings) rather than one hub the way US labs cluster in SF (85%).
  • The open-closed gap. Epoch’s index puts the best open-weight models ~4 months behind the closed frontier since Jan 2026 — an 8-point capability gap ≈ GPT-5→GPT-5.5.

Practical read: export controls are leakier than headline policy implies, but China’s frontier training is still Nvidia-bound; the domestic-chip story is real for inference and small models, not yet for frontier pretraining. Price the ~4-month open-weight lag into any cost-driven routing to GLM/Kimi-class models.


Try It

  • Treat “frontier gets pricier, mid-tier catches up” as a planning assumption, not a hope. The compute crunch predicts exactly the pattern the wiki’s cost work already recommends: route routine, center-of-distribution work to cheaper/smaller models and reserve frontier spend for genuinely novel-capability tasks. See Cost & Intelligence Levers.
  • Watch long-context pricing specifically. Agentic and long-context workloads are the first to hit the compute wall. If your workflow depends on 100K+ token contexts, expect that to be where price pressure and rate limits land first — design for context economy.
  • Keep an open-weight fallback, and price the lag. The supply-consolidation thesis (Anthropic + OpenAI absorbing headroom) is the structural case for a portable harness. Budget open-weight routing at a ~4-month capability lag, not parity.
  • Use HBM as your supply-relief indicator. Because memory is the binding constraint, watch HBM capacity announcements (not logic-node news) as the leading signal for whether the crunch eases.
  • Sanity-check “AI bubble” claims against the capex-vs-cash-flow line. Epoch’s data (capex overtaking operating cash flow, forcing external financing in 2026) is the concrete number to cite when the buildout’s sustainability comes up.

Open Questions

  • How wide are the true supply/demand slopes? Epoch’s own ranges (supply 500M–20B tok/s; demand 200M–4B tok/s) are order-of-magnitude wide and depend on unknown average model size behind current demand. The crunch timing is directional, not precise.
  • Will Anthropic/OpenAI actually sustain ~4×/year? The consolidation clock assumes continued 4× compute growth; a capex slowdown after 2026 (not guaranteed) would push the “headroom exhausted” date out.
  • Does the ~$1T-capex-can’t-double-again ceiling bind, or does AI-driven growth relieve it? The entire “scaling stalls” conclusion hinges on whether AI accelerates economic growth enough to fund the next doubling — Epoch leaves this open.
  • How much undetected smuggling? The 660k median is bracketed by heavy uncertainty on detection rate (median 24.5%) and whether allegedly-diverted chips actually reached China.
  • Not folded here — Epoch’s AGI-transition economics tail. Four adjacent Epoch Gradient Updates (superstar-researcher pay, capital control after AGI, the “missing half” of futurism, an O*NET for AI R&D) were read during this ingest but judged too abstract for this wiki’s applied bar. Flagged for the main session in case a dedicated “AGI-transition economics” synthesis is ever wanted.