Source: raw/newsletter-epoch-ai-add82097a6.md (“Is a compute crunch coming?”, Luke Emberson & Jaime Sevilla) · raw/newsletter-epoch-ai-e05a84c055.md (“Frontier labs don’t use most AI compute (yet)”, Josh You) · raw/newsletter-epoch-ai-a6ad3f7054.md (“Diversion and resale: estimating compute smuggling to China”, Isabel Juniewicz) · raw/newsletter-epoch-ai-6fb94702d7.md (“What we learned from 1,604 Chinese AI job postings”, Cheryl Wu, JS Denain, Anson Ho) · Epoch Brief data-insights: raw/newsletter-epoch-ai-61525dda28.md (June 1), raw/newsletter-epoch-ai-66b49f0f8d.md (May 22), raw/newsletter-epoch-ai-45df197d73.md (May 15), raw/newsletter-epoch-ai-ac7e42b77d.md (June 12), raw/newsletter-epoch-ai-b3167bb772.md (June 26), raw/newsletter-epoch-ai-b9855d3cc7.md (May 8)
Publisher: Epoch AI. The three lead pieces are Gradient Updates — Epoch’s explicitly opinionated/informal series representing the named authors’ views, not an Epoch institutional position. The smuggling piece excerpts a fuller Epoch report; the bracketed Data Insights are Epoch’s standard (non-opinion) data work.
A consolidation of Epoch AI’s mid-2026 quantitative work on the cost and physical limits of the AI buildout, distilled for the one question a practitioner actually cares about: what does this imply for the price and availability of frontier intelligence? The through-line across all of it is a squeeze — token demand appears to be outrunning inference supply, the two labs driving demand (Anthropic and OpenAI) are on track to absorb the world’s compute headroom within a few years, and the physical inputs (memory, data centers, chips) are the binding constraints. This is capability-and-cost forecasting, not marketing advice — but it grounds decisions about model choice and when to expect frontier access to get more expensive.
Key Takeaways
- A compute crunch is plausibly near — and it hits long-context agentic workloads first. Emberson & Sevilla model that today’s Blackwell GPUs could serve 500 million–20 billion output tokens/second, with global inference capacity more than tripling each year (~3–4×). But token demand — estimated at 200 million–4 billion tokens/second at current prices — appears to be growing ~10× per year, plausibly outpacing supply already. ^[The supply and demand ranges are wide and the authors flag them as highly uncertain.]
- The practical consequence: frontier long-context access gets pricier; everyday users get pushed to smaller/cheaper models. Epoch’s own read — “the price of access to frontier capabilities [rises] for those willing to pay, while everyday users shift to cheaper, smaller models.” Because inference/training efficiency keeps improving fast, this need not mean worse quality: tomorrow’s smaller, cheaper models quickly match today’s frontier. This is the cost lever behind the wiki’s cheap-workhorse-vs-daily-driver split.
- Frontier labs don’t yet use most AI compute — but Anthropic and OpenAI are on track to. Josh You estimates the five most compute-rich developers (OpenAI, Anthropic, xAI, Google DeepMind, Meta) together held under half of world compute at end-2025; OpenAI alone was ~10–15%. Yet OpenAI and Anthropic are growing compute ~4×/year vs the industry’s ~3× — enough to consume the headroom and reach ~80% of world compute within ~5 years, at which point scaling stalls unless total capex accelerates dramatically.
- The buildout is already ~1T annualized in 2026 (~1% of gross world product, ~3% of US GDP); hyperscaler capex has quadrupled since GPT-4 and is on track to overtake hyperscalers’ operating cash flows by end of 2026, forcing external financing. GPU-hour rental prices rose ~30% this year — a supply-tightness signal you can already feel in API pricing.
- Memory, not logic, is the bottleneck. High-bandwidth memory (HBM) grew from 52% to 63% of AI-chip component cost (Q1 2024→Q4 2025), spend rising ~32B — faster than any other component. Servers account for 60% of the total cost of owning a 1 GW data center (8.5B/year), dwarfing energy.
- China’s compute is heavily smuggled, and Chinese labs lean on domestic chips only for the easy jobs. Epoch estimates 660,000 H100-equivalents smuggled into China through 2025 (90% CI 290k–1.6M) — roughly a quarter to a third of China’s total AI compute. Separately, Chinese job postings suggest domestic chips are used frequently for inference, rarely for pretraining large models (GLM-Image, at 16B params, was trained entirely on domestic chips — but that is 10–100× smaller than their frontier models).
- Open-weight models trail the closed frontier by ~4 months in Epoch’s capability index (an 8-point gap, ≈ the GPT-5→GPT-5.5 difference) — the durable lag to price into any open-weight routing decision.
1. Is a compute crunch coming? (the demand/supply squeeze)
Gradient Update by Luke Emberson & Jaime Sevilla, 2026-05-26. The authors model the supply side precisely (how many tokens the world’s chips could serve) and compare it to imperfect demand proxies.
- Supply. Serving a Kimi K2.6-class open model on all the world’s Nvidia GB200 + GB300 chips (1.9M and 1.5M GPUs at Q4 2025, ~40% of aggregate FLOP/s), and accounting for real-world inefficiencies (calibrated against SemiAnalysis’s InferenceX data → ~65% compute / 30% bandwidth efficiency), yields ~400,000 tok/s per GB200 NVL72 rack and a global total of 500 million–20 billion output tokens/second — depending heavily on request context length. That’s 150,000–7 million tokens per month for every person on earth. Capacity is more than tripling each year.
- Demand. Proxies (Google’s ~1.2B tok/s across its platforms ≈ 130M output tok/s; Exponential View’s ~40 quadrillion tokens/quarter ≈ 5B tok/s industry-wide) suggest current demand of 200M–4B tokens/second, growing ~10×/year — a faster slope than supply’s ~3–4×.
- The crunch and who feels it. If these slopes hold, demand overtakes supply soon (if not already), “particularly for the long-context workloads that drive agentic AI.” Long context is the culprit because attention compute grows quadratically — the batch size at which inference becomes compute-bound shrinks as context grows (e.g. from ~870 concurrent users per rack at 8K context down to ~130 at 25K). ^[inferred: the “agentic = long context = hits the compute wall first” causal chain is the article’s argument, restated here as the practical takeaway.]
- Why quality need not regress. Inference and training efficiency are improving fast enough (Epoch’s own price-trend data) that “the smaller, cheaper models of tomorrow will quickly match today’s frontier.” The crunch reprices frontier access without necessarily degrading what everyday users get.
Practical read: budget for frontier long-context/agentic work to get relatively more expensive over time, and lean on the “this year’s mid-tier ≈ last year’s flagship” migration for routine work. This is the supply-side mechanism underneath Cost & Intelligence Levers.
2. Frontier labs don’t use most AI compute — yet (the consolidation clock)
Gradient Update by Josh You, 2026-05-21. The author flags these lab-compute estimates as more tentative than Epoch’s standard data work.
- Today’s split. Cumulative AI compute sold ≈ 20M H100-equivalents (H100e) by end-2025; ~16M operational after a one-quarter install lag. OpenAI ≈ 1.7M H100e / 1.9 GW (up from 0.6 GW in 2024, 0.2 GW in 2023). Anthropic ≈ 1.4 GW (~70% of OpenAI), much of it ~500k H100e of Trainium2 at Amazon’s Project Rainier. OpenAI + Anthropic + xAI together < 4M H100e ≈ 20–30% of world compute; Google + Meta own ~1/3 of world compute but allocate perhaps half of that to their frontier labs (~15% of world total).
- The consolidation clock. OpenAI and Anthropic grew compute ~4×/year in 2025 vs the industry’s ~3×. A naive extrapolation (top-two at 20% share today, growing 33% faster than the world) has them doubling their share in 2.5 years and reaching ~80% within five. Anthropic’s revenue run-rate went 30B annualized in Q1 2026 (an acceleration on last year’s ~10× growth), funding aggressive compute deals — including renting xAI’s entire ~300k-H100e Colossus 1 cluster (up to $15B/year).
- The wall isn’t capability — it’s capex. Sustaining 4×/year once the headroom is gone would require more than doubling capex annually from a ~$1T-in-2027 base — feasible “only if AI starts to dramatically accelerate economic growth.” Otherwise frontier compute growth converges down to overall compute growth. Epoch is explicit this is not a hard 2029 “wall”: flat capex still grows the stock, chips keep improving. But the compute-scaling driver of progress is “not sustainable unless the world fundamentally changes soon.”
Practical read: the two vendors most likely to tighten supply and raise prices are also the two you most likely depend on. This is the supply-consolidation case for keeping a portable harness and an open-weight fallback (Lever 4).
3. The physical inputs: capex, memory, data centers
Epoch’s (non-opinion) Data Insights quantify the binding constraints:
| Constraint | Finding | Source |
|---|---|---|
| Hyperscaler capex | Quadrupled since GPT-4; ~1T in 2027; Q1 2026 came in at 155.1B projection) | June 1 brief |
| Capex vs cash flow | On track to overtake hyperscalers’ operating cash flows by end-2026 → external financing (Alphabet’s $80B raise, Amazon bond offerings) | June 26 brief (also MirrorCode’s source) |
| Memory (HBM) | Rose from 52% → 63% of AI-chip component cost (Q1’24→Q4’25); ~32B spend; the dominant cost and primary supply bottleneck | May 22 brief |
| Data-center servers | 60% of total cost of a 1 GW data center — 8.5B/year, dwarfing energy | May 15 brief (via May 22 ICYMI) |
| Single-site scale | Record for compute at one data center has doubled every 7 months since Colossus 1 (Aug 2024) | June 12 brief |
| Macro footprint | AI infra ≈ 1.5% of US GDP (data-center construction alone ~0.8% in Q1 2026), up from a ~0.7% 2015–22 average; now the leading driver of US private-investment growth | June 12 brief |
| Compute ownership | Five hyperscalers own >2/3 of global AI compute | May 8 brief |
| Unit economics | Anthropic and OpenAI earn ~5.5M revenue per employee — higher than any public tech company on the Forbes Global 2000 | May 8 brief |
Practical read: HBM is the choke point, so supply relief depends on memory fabs, not just logic nodes — a slower lever. The revenue-per-employee figures are a useful sanity check when sizing what these labs can afford to spend on compute.
4. China: smuggled chips + a domestic-silicon workaround
Two Epoch pieces bear on Chinese compute — relevant for anyone forecasting the open-weight frontier (GLM-5.2, Kimi K3).
- Smuggling (Isabel Juniewicz, full report). Cumulative allegations of diverted/missing chips total ~300,000 H100e by end-2025 — about a quarter of what China acquired legally. A Monte-Carlo estimate puts true smuggling at 660,000 H100e (90% CI 290k–1.6M) ≈ 3% of the global compute stockpile (comparable to what xAI held), with the upper bound implying most of China’s compute was smuggled. Context: China imported ~2.5B Supermicro diversion scheme. Cloud compute outside China used by Chinese customers is a separate, generally-permitted channel not counted here.
- Domestic chips, from 1,604 job postings (Cheryl Wu, JS Denain, Anson Ho). Scraping DeepSeek, MiniMax, Moonshot, Z.ai, ByteDance, and Alibaba postings: Chinese labs still depend on Nvidia/CUDA/TensorRT-LLM for inference, but are actively hiring for Ascend/Cambricon “heterogeneous computing.” Best guess: domestic chips are used frequently for inference, rarely for pretraining large models — the exception being post-training or small models (Z.ai’s GLM-Image, 16B params, was trained entirely on domestic chips, but that’s 10–100× smaller than their frontier models). Other findings worth keeping: Chinese labs have distinct “personalities” (Z.ai is B2B-heavy like Anthropic — 73.7% of 2025 revenue from running models on customers’ infra; MiniMax/Moonshot are consumer/international — 73% of MiniMax revenue is international vs Z.ai’s 9.8%), require far less experience (US labs avg 5.5 years vs 1.6 in China), and cluster in Beijing/Hangzhou/Shanghai (93% of located postings) rather than one hub the way US labs cluster in SF (85%).
- The open-closed gap. Epoch’s index puts the best open-weight models ~4 months behind the closed frontier since Jan 2026 — an 8-point capability gap ≈ GPT-5→GPT-5.5.
Practical read: export controls are leakier than headline policy implies, but China’s frontier training is still Nvidia-bound; the domestic-chip story is real for inference and small models, not yet for frontier pretraining. Price the ~4-month open-weight lag into any cost-driven routing to GLM/Kimi-class models.
Try It
- Treat “frontier gets pricier, mid-tier catches up” as a planning assumption, not a hope. The compute crunch predicts exactly the pattern the wiki’s cost work already recommends: route routine, center-of-distribution work to cheaper/smaller models and reserve frontier spend for genuinely novel-capability tasks. See Cost & Intelligence Levers.
- Watch long-context pricing specifically. Agentic and long-context workloads are the first to hit the compute wall. If your workflow depends on 100K+ token contexts, expect that to be where price pressure and rate limits land first — design for context economy.
- Keep an open-weight fallback, and price the lag. The supply-consolidation thesis (Anthropic + OpenAI absorbing headroom) is the structural case for a portable harness. Budget open-weight routing at a ~4-month capability lag, not parity.
- Use HBM as your supply-relief indicator. Because memory is the binding constraint, watch HBM capacity announcements (not logic-node news) as the leading signal for whether the crunch eases.
- Sanity-check “AI bubble” claims against the capex-vs-cash-flow line. Epoch’s data (capex overtaking operating cash flow, forcing external financing in 2026) is the concrete number to cite when the buildout’s sustainability comes up.
Related
- Cost & Intelligence Levers for Agent Workflows — the demand-side operator playbook; this article is its supply-side macro backdrop.
- MirrorCode (Epoch AI + METR) — sibling Epoch benchmark; its June-26-brief source also carries the capex-vs-cash-flow insight used here.
- AI Competition Shifts Beyond Model Quality — the strategy-side read on the same hyperscaler capex buildout.
- xAI Colossus Compute Deal — the concrete Anthropic compute-securing move behind the consolidation thesis.
- Stanford HAI 2026 — Chapter 4 (Economy) — independent macro backdrop on AI investment and diffusion.
- GLM-5.2 and Kimi K3 — the open-weight Chinese models whose compute story this article contextualizes.
- AI Industry Research — topic hub; Epoch AI is a named source-quality anchor.
Open Questions
- How wide are the true supply/demand slopes? Epoch’s own ranges (supply 500M–20B tok/s; demand 200M–4B tok/s) are order-of-magnitude wide and depend on unknown average model size behind current demand. The crunch timing is directional, not precise.
- Will Anthropic/OpenAI actually sustain ~4×/year? The consolidation clock assumes continued 4× compute growth; a capex slowdown after 2026 (not guaranteed) would push the “headroom exhausted” date out.
- Does the ~$1T-capex-can’t-double-again ceiling bind, or does AI-driven growth relieve it? The entire “scaling stalls” conclusion hinges on whether AI accelerates economic growth enough to fund the next doubling — Epoch leaves this open.
- How much undetected smuggling? The 660k median is bracketed by heavy uncertainty on detection rate (median 24.5%) and whether allegedly-diverted chips actually reached China.
- Not folded here — Epoch’s AGI-transition economics tail. Four adjacent Epoch Gradient Updates (superstar-researcher pay, capital control after AGI, the “missing half” of futurism, an O*NET for AI R&D) were read during this ingest but judged too abstract for this wiki’s applied bar. Flagged for the main session in case a dedicated “AGI-transition economics” synthesis is ever wanted.