Source: raw/ai_index_report_2026.pdf
Publisher: Stanford Institute for Human-Centered Artificial Intelligence (HAI)
Citation: Sha Sajadieh, Loredana Fattorini, Raymond Perrault, Yolanda Gil et al. “The AI Index 2026 Annual Report,” AI Index Steering Committee, Institute for Human-Centered AI, Stanford University, Stanford, CA, April 2026.
License: CC BY-ND 4.0 International
Length: 423 pages, 9 chapters
Co-chairs: Yolanda Gil (USC/ISI) · Raymond Perrault (SRI)
Lead editor: Sha Sajadieh (Stanford)
Supporting partners: Google, NSF, OpenAI, Open Philanthropy, Quid, Infosys
Data partners: Digital Policy Alert, Center for Research on Foundation Models, CSTA, Epoch AI, GitHub, IFR (International Federation of Robotics), Kapor Foundation, Lightcast, LinkedIn, McKinsey, NSF ECEP, Schmidt Sciences, RAISE Health, zeki
The ninth edition of the Stanford HAI AI Index. First edition to include standalone chapters on AI in Science and AI in Medicine. Adds AI sovereignty as an analytical framework. The co-chairs’ framing: “The data does not point in a single direction. It reveals a field that is scaling faster than the systems around it can adapt.” Authoritative independent measurement — the primary artifact for sanity-checking vendor claims and framing strategy conversations.
The Core Finding
Generative AI hit nearly 53% population-level adoption within three years — faster than the personal computer or the internet. Organizational adoption is 88%. 4 in 5 university students now use generative AI. The US-China model performance gap has effectively closed. Safety benchmarks are lagging capability. Talent flow into the US has dropped 89% since 2017.
The 15 Top Takeaways (Verbatim from Report)
-
AI capability is not plateauing. Industry produced over 90% of notable frontier models in 2025, and several meet or exceed PhD-level baselines in PhD-level science, multimodal reasoning, and competition mathematics. SWE-bench Verified rose from 60% to near 100% in a single year. Organizational adoption reached 88%, and 4 in 5 university students use generative AI.
-
The US-China AI model performance gap has effectively closed. US and Chinese models have traded the lead multiple times since early 2025. In February 2025, DeepSeek-R1 briefly matched the top US model; as of March 2026, Anthropic’s top model leads by just 2.7%. The US still produces more top-tier models and higher-impact patents; China leads in publication volume, citations, patent output, and industrial robot installations. South Korea leads in AI patents per capita.
-
The US hosts the most AI data centers, with chips fabricated by one Taiwanese foundry. US hosts 5,427 data centers — more than 10× any other country, consuming more energy than any other. A single company, TSMC, fabricates almost every leading AI chip, making the global hardware supply chain dependent on one foundry in Taiwan (TSMC-US expansion began operations in 2025).
-
AI models can win gold at the International Mathematical Olympiad but cannot reliably tell time. Gemini Deep Think earned a gold medal at IMO, yet the top model reads analog clocks correctly just 50.1% of the time. AI agents jumped from 12% to ~66% task success on OSWorld but still fail roughly 1 in 3 attempts on structured benchmarks. This is the “jagged frontier”.
-
Robots still fail at most household tasks. Robots succeed in only 12% of household tasks. On RLBench, simulation manipulation hits 89.4% success — the sim-to-real gap is wide.
-
Responsible AI is not keeping pace with capability; safety benchmarks are lagging; AI incidents rising sharply. Documented AI incidents rose to 362 in 2025, up from 233 in 2024. Frontier labs report capability benchmarks consistently; responsible-AI reporting is spotty. Improving one dimension (e.g. safety) can degrade another (e.g. accuracy).
-
The US leads AI investment but is losing ability to attract global talent. US private AI investment reached **12.4B (headline figure; China’s government guidance funds likely understate total). US led in entrepreneurial activity with 1,953 newly funded AI companies (10× the next country). But AI researchers/developers moving to the US dropped 89% since 2017 — 80% drop in the last year alone.
-
AI adoption is spreading at historic speed; consumers deriving substantial value from free tools. GenAI hit 53% population adoption in three years, faster than PC or internet. Adoption varies by country: Singapore 61%, UAE 54%, US 28.3% (ranked 24th). Estimated consumer value of GenAI tools to US consumers reached $172B annually by early 2026, with median per-user value tripling between 2025 and 2026.
-
Productivity gains from AI appear in the same fields where entry-level employment is declining. Studies show 14-26% productivity gains in customer support and software development; weaker/negative effects in high-judgment tasks. AI agent deployment remains single-digit across most business functions. US software developers ages 22-25 saw employment fall nearly 20% from 2024 even as older developer headcount grew.
-
AI’s environmental footprint is expanding. Grok 4 training emissions reached 72,816 tons CO₂-equivalent. AI data center power capacity rose to 29.6 GW at peak demand — comparable to New York State. Annual GPT-4o inference water use alone may exceed drinking water needs of 12 million people.
-
AI models for science can outperform human scientists, but bigger ≠ better. Frontier models outperform chemists on average on ChemBench but score <20% on astrophysics replication and 33% on Earth observation. A 111M-parameter protein model (MSAPairformer) beat previous leading methods on ProteinGym; a 200M genomics model (GPN-Star) outperformed a model nearly 200× larger. Most AI-for-science foundation models come from cross-sector collaborations, not industry alone.
-
AI is transforming clinical care, but rigorous evidence remains limited. AI clinical-note-generation tools saw substantial 2025 adoption. Physicians reported up to 83% less time writing notes and significant burnout reductions. But a review of 500+ clinical AI studies found nearly half relied on exam-style questions rather than real patient data; only 5% used real clinical data.
-
Formal education is lagging but people are learning AI skills at every stage. Over 80% of US high school and college students use AI for school. Only half of middle/high schools have AI policies; just 6% of teachers say those policies are clear. AI engineering skills are accelerating fastest in UAE, Chile, South Africa. New AI PhDs in US+Canada rose 22% (2022→2024), but those PhDs are taking academia jobs, not industry.
-
AI sovereignty is becoming a defining feature of national policy. National AI strategies are expanding, particularly among developing economies; state-backed supercomputing investments rising in parallel. Model production remains concentrated in US+China, but open-source development is redistributing participation — the rest of the world now outpaces Europe and approaches the US on GitHub.
-
AI experts and the public have very different perspectives; global trust is fragmented. 73% of experts expect positive impact on how people do their jobs vs. only 23% of the public — a 50-point gap. US has the lowest trust in its own government to regulate AI, at 31%. Globally, the EU is trusted more than US or China to regulate AI effectively.
Chapter Map (423 pages total)
| Chapter | Page | Focus |
|---|---|---|
| 1. Research and Development | 12 | Notable models, compute, data centers, energy, open source, publications, patents, talent |
| 2. Technical Performance | 68 | Benchmark progress, agents, jagged-frontier examples |
| 3. Responsible AI | 126 | Safety benchmarks, incidents, disclosure gaps, tradeoffs |
| 4. Economy | 171 | Investment, adoption, productivity, labor markets, consumer value |
| 5. Science | 231 | AI for science, foundation models, cross-sector collab |
| 6. Medicine | 255 | Clinical AI adoption, evidence gaps, ambient scribes |
| 7. Education | 288 | Student/teacher use, AI policies, skill diffusion |
| 8. Policy and Governance | 323 | EU AI Act, US deregulation, national strategies, AI sovereignty |
| 9. Public Opinion | 360 | Trust gaps, expert/public divergence, global sentiment |
Chapter 1 Highlights (R&D)
- Industry produced >90% of notable models in 2025; frontier models are less transparent — training code, dataset sizes, parameter counts withheld by OpenAI, Anthropic, Google.
- China: 50 notable models in 2025 (vs US 30 in the organization view — but US retains 50 notable overall, China 30). US retains higher-impact patents; China has grown top-100 cited papers from 33 (2021) to 41 (2024).
- Parameter counts stayed near 1 trillion for three years; frontier labs stopped reporting. Training compute (independently estimable) continued rising.
- Synthetic data not replacing real pre-training data yet, but post-training + data-quality techniques showing promise (OLMo 3.1 Think 32B with 90× fewer params than Grok 4 achieves comparable results on some benchmarks via pruning/dedup/curation).
- Global AI compute: 3.3× per year since 2022, reaching 17.1M H100-equivalents. Nvidia ≥60%; Google+Amazon supply most of the rest; Huawei small but growing.
- Energy: Grok 4 training ≈72,816 tons CO₂-eq. AI data center power = 29.6 GW (NY State at peak).
- Open source: 5.6M projects on GitHub; Hugging Face uploads tripled since 2023. US projects still dominate engagement (30M cumulative stars across 10+ star projects).
- Talent migration to the US dropped 89% since 2017, 80% in last year alone. Switzerland + Singapore lead per-capita researcher density.
- Gender gap: deeply entrenched; no country approaches parity. Saudi Arabia (32.3%), Canada (29.6%), Australia (30.1%) lead.
Key Takeaways (Distilled for Practical Use)
- Cite-worthy adoption stats: 53% pop / 88% org / 80% student — use these when framing “AI is mainstream” slides.
- Jagged-frontier examples are the best way to communicate “AI is powerful but unreliable” to non-technical stakeholders. IMO gold medal + 50.1% analog clock reading is the canonical pairing.
- US-China gap closed is a strategic signal — if you’re sourcing models, open-weight Chinese models (DeepSeek, Qwen, Z.ai) are competitive with frontier.
- Consumer value of free GenAI: $172B/yr US, tripling year-over-year — frames willingness-to-pay analysis for any SaaS wedge.
- Productivity gains 14-26% in support + dev — this is the defensible stat for marketing-automation ROI decks.
- 22-25 year-old dev employment -20% — the “entry-level is collapsing first” data point.
- AI incidents +55% YoY (233→362) — the specific number to cite when making the case for governance/safety investment.
- Safety benchmarks are saturating or missing — frontier labs disclose capability, not responsible-AI, consistently.
Try It
- For strategy decks: pull the chart images from the HAI public data ([Google Drive link in report]) and cite Figure numbers directly. License is CC BY-ND — attribution required, no derivatives of the charts themselves.
- For policy work: the HAI report adds national/regulatory context (EU AI Act, US deregulation, sovereignty framing) to any private-sector governance frame you already use.
- For sales conversations: the “expert 73% vs public 23%” gap is the single most useful data point for framing objections about AI-assisted work.
- For the Karpathy wiki itself: this is the baseline for annual “state of the field” file-back articles — when HAI 2027 lands, we compare.
Implementation
Tool/Service: AI Index 2026 public data (Google Drive) + interactive Global AI Vibrancy tool (36 countries, updating end of 2026).
Setup: Report is freely downloadable from aiindex.stanford.edu.
Cost: Free; CC BY-ND 4.0.
Integration notes:
- Epoch AI is the data partner for compute + notable models. Query it directly rather than waiting for the annual report:
epoch.ai/data/ai-models?subset=notable&view=table(browsable), or pullepoch.ai/data/notable_ai_models.csvdirectly (CC-BY, synced daily). No REST API; a Python client (pip install epochai) reads their underlying Airtable base. Full detail in Chapter 8 deep-dive. AI Index is a February snapshot (data cutoff 2026-02-12). - LinkedIn (via Lightcast) is the labor-market data source — the 22-25-year-old developer employment decline actually traces to Brynjolfsson et al. (2025) ADP payroll data, not LinkedIn directly; see Chapter 4 deep-dive for the full study bibliography.
- McKinsey & Company (“The State of AI in 2025: Agents, Innovation, and Transformation,” fielded Jun-Jul 2025, n=1,993 across 105 nations) is the org-adoption source; full methodology in the Chapter 4 deep-dive.
- OECD.AI Index (oecd.ai, DOI 10.1787/32c01014-en, Feb 2026) is the international policy-implementation-scoring complement — 38 OECD countries, 28 indicators; Stanford HAI sits on its Expert Group. Full detail in the Chapter 8 deep-dive.
External Report Comparisons
McKinsey State of AI, OECD.AI Index, Epoch AI’s own trackers — each now has a full treatment in the relevant deep-dive: McKinsey’s exact methodology and its own agent-adoption quote live in the Chapter 4 deep-dive; the OECD.AI Index (a real, distinct, DOI-registered Feb 2026 publication — not to be confused with the OECD.AI Policy Observatory or the separate OECD AI Capability Indicators) and Epoch AI’s query surfaces + self-acknowledged non-English coverage gap live in the Chapter 8 deep-dive.
Gartner Hype Cycle for AI, 2025/2026 — researched 2026-07-02. Gartner’s most recent general “Hype Cycle for Artificial Intelligence” is still the 2025 edition (press release dated August 5, 2025); no 2026 successor to the general AI Hype Cycle was found despite a direct search of Gartner’s newsroom through late June 2026. Instead, Gartner split 2026 coverage into narrower named cycles — a first-ever “Hype Cycle for Agentic AI, 2026” (April 2, 2026) and a “Hype Cycle for Generative AI, 2026” (May 20, 2026, largely paywalled). In the 2025 general cycle, generative AI itself sat in the Trough of Disillusionment (average $1.9M spend per initiative in 2024, <30% CEO satisfaction with ROI), while AI agents and AI-ready data were the fastest-advancing items, both at the Peak of Inflated Expectations. The 2026 Agentic AI cycle keeps agentic AI overall at the Peak (17% of organizations had deployed agents, 60%+ expected to within two years) and separately projects that more than 40% of agentic AI projects will be canceled by the end of 2027.
Note: the phrase “accessibility inflection point” used to frame this comparison in the original research-agenda entry does not appear verbatim anywhere in HAI’s own materials — it is a paraphrase, not a direct quote, and should not be attributed to HAI as a named term. The two frameworks measure genuinely different axes and are not in direct contradiction: Gartner’s Trough placement is an enterprise/IT-leader lens on hype-versus-delivered-ROI (spend, CEO satisfaction, governance friction), while HAI’s 53%-in-three-years figure is population-level “has anyone used it” adoption speed. Gartner has not shifted toward a “mainstream/accessible” framing for GenAI or agents — if anything, the emergence of a brand-new, separately-hyped “Agentic AI” cycle in 2026 shows the classic Gartner pattern of hype migrating to the next-named subcategory rather than the underlying capability graduating to the Plateau of Productivity.
“73% expert / 23% public” trust gap — who counts as “expert”? Researched 2026-07-02, resolved directly from Chapter 9 (p. 372). Source: McClain et al. (2025), How the U.S. Public and AI Experts View Artificial Intelligence, Pew Research Center. For this survey, “AI experts” were defined as U.S.-based authors or presenters at AI-related conferences in 2023 or 2024 who reported that their work or research relates to AI — a professional-activity-based definition, not self-identification. Sample sizes: 1,013 AI experts (survey) plus 30 individual AI experts for in-depth interviews (Oct-Nov 2024), compared against 5,410 U.S. adults in the general-public survey (Aug 2024). The gap is genuinely large and consistent across topics, not just jobs: economy 69% (experts) vs. 21% (public), K-12 education 61% vs. 24%, medical care 84% vs. 44% — all Pew Research 2025, Figure 9.2.1.
Open Questions
This article is a top-takeaways + chapter-map summary, not a full ingestion of the 423-page report. Status as of the 2026-07-02 drain pass:
Chapter 2 (Technical Performance) — agent benchmarks, OSWorld, jagged-frontier dossier.Resolved 2026-07-03 — see Chapter 2 Technical Performance Deep-Dive. Short answer: SWE-bench Verified rose ~60%→~100% (2024→2025); OSWorld agents rose 12%→66.3%; robots still fail ~88% of real household tasks (12.4% full-task success) vs. 89.4% in controlled simulation; the report’s own jagged-frontier catalog spans at least 8 documented pairings beyond the canonical IMO-gold-vs-clock-reading example.Chapter 3 (Responsible AI) — documented incidents catalog, safety-vs-accuracy tradeoffs.Resolved 2026-07-03 — see Chapter 3 Responsible AI Deep-Dive. Short answer: AI Incident Database recorded 362 incidents in 2025 (+55% YoY from 233); Foundation Model Transparency Index average score fell from 58 to 40 (upstream training-data disclosure averages just 15% vs. 69-75% for downstream policy categories); safety benchmarks show “Good”/“Very Good” ratings under normal conditions but collapse under the Jailbreak T2T v0.5 adversarial benchmark; § 3.10 documents 3 empirical studies quantifying RAI-dimension tradeoffs. Pairs directly with WEO AI Governance.Chapter 4 (Economy) — detailed productivity study list and methodology.Resolved 2026-07-02 — see Chapter 4 Economy Deep-Dive.Chapter 8 (Policy and Governance) — EU AI Act first prohibitions, US deregulation shift, national strategy diff.Resolved 2026-07-02 — see Chapter 8 Policy and Governance Deep-Dive.HAI doesn’t publish a methodology for “notable AI models” — relies on Epoch AI curation. Bias risk: what does Epoch miss from non-English sources?Resolved 2026-07-02 — see the Epoch AI Methodology section of the Chapter 8 deep-dive. Short answer: Epoch AI’s formal criteria are silent on language/geography, but Epoch’s own blog repeatedly acknowledges non-English (especially Chinese-language) models are harder for its team to discover; HAI’s own Chapter 8 separately names specific sub-Saharan African language models excluded entirely from its regional model counts.Consumer-value figure ($172B/yr) methodology not visible in top-takeaways section; requires reading Chapter 4 to assess.Resolved 2026-07-02 — see Chapter 4 Economy Deep-Dive. Short answer: it’s a stated-preference willingness-to-accept choice experiment (Brynjolfsson et al. 2026), not revenue and not revealed-preference.“AI experts 73% vs public 23%” gap — who exactly are “experts” in the survey sample?Resolved 2026-07-02 — see the External Report Comparisons section above.The “88% org adoption” figure is McKinsey-sourced and includes very light use.Resolved 2026-07-02 — see Chapter 4 Economy Deep-Dive. Short answer: 88% = any AI in any function (bar has loosened since 2020); the comparable agent figure, in McKinsey’s own words, is “no more than 10 percent… scaling AI agents” in any given business function.
Related
- Chapter 4 — Economy Deep-Dive — investment, McKinsey adoption methodology, consumer-value study, full productivity-research bibliography
- Chapter 8 — Policy and Governance Deep-Dive — EU AI Act, US deregulation, AI sovereignty, Epoch AI + OECD.AI Index methodology
- AI Agents Unleashed — 2026 Playbook (Mindstream × Futurepedia) — complementary: HAI gives the macro data, Mindstream gives the implementation playbook
- Claude Agent Hierarchy
- Claude AI — vendor-specific landing
- AI Marketing Automation Use Cases
- WalkMe State of Digital Adoption 2026