Authoritative industry reports and benchmark studies on the state of AI. Source material for strategy decks, policy framing, and sanity-checking vendor claims. Bias toward independent, data-driven reports (Stanford HAI, Epoch AI, McKinsey, OECD) over marketing-driven “state of AI” content.

Reports

  • Stanford HAI AI Index Report 2026 — 423-page flagship annual. 15 top takeaways spanning research, technical performance, responsible AI, economy, science, medicine, education, policy, public opinion. Key signals: 53% generative AI population adoption in 3 years, 88% org adoption, SWE-bench 60%→near-100% in one year, US-China model gap effectively closed, safety benchmarks lagging capability.

  • Stanford HAI AI Index 2026 — Chapter 2 Technical Performance Deep-Dive — full Ch2 extraction: US-China/open-closed Arena convergence (top 4 labs within 25 Elo points as of March 2026), the report’s own 8-example “jagged frontier” catalog (IMO gold medal vs. 50.1% clock-reading, plus PlanBench, CyBench-vs-BEHAVIOR-1K, and more), the exact SWE-bench/Terminal-Bench/Vibe Code Bench software trajectory, all six named AI-agent benchmarks (GAIA/OSWorld/WebArena/MLE-bench/CyBench/tau-bench), and the robotics sim-to-real gap (RLBench 89.4% simulated vs. 12% real household-task success, contrasted with self-driving cars’ mass deployment).

  • Stanford HAI AI Index 2026 — Chapter 3 Responsible AI Deep-Dive — full Ch3 extraction: the AI Incident Database’s 362-incident catalog (+55% YoY) with worked examples, three distinct hallucination/factuality benchmarks (HHEM, AA-Omniscience, KaBLE belief-vs-fact), the Foundation Model Transparency Index disclosure-gap breakdown (58→40 average score, upstream Data Properties disclosure just 15% vs. 69-75% downstream), AILuminate/HELM Safety/Jailbreak T2T results showing safety collapse under adversarial attack, and three empirical studies quantifying safety/privacy/accuracy tradeoffs. Pairs directly with WEO AI Governance.

  • Stanford HAI AI Index 2026 — Chapter 4 Economy Deep-Dive — full Ch4 extraction: investment/infrastructure, the McKinsey 88%-adoption-vs-single-digit-agent-deployment breakdown, the $172B consumer-surplus willingness-to-accept methodology (Brynjolfsson et al. 2026), and the full productivity/labor-market study bibliography (16 named studies) behind the 22-25-year-old developer employment decline.

  • Stanford HAI AI Index 2026 — Chapter 8 Policy and Governance Deep-Dive — full Ch8 extraction: 2025 global policy timeline (EU AI Act, US deregulation shift), national AI strategies, the new five-dimension AI sovereignty framework (infrastructure/data/model/application/talent), legislative records, plus an Epoch AI methodology deep-dive resolving the notable-models non-English-coverage question.

  • WalkMe State of Digital Adoption 2026 — “The AI Reality Check.” 3,750 participants, 41 pages. The Execution Gap: 142M/yr cost of digital inefficiency. Shadow AI: 45% using unapproved tools. Executive-employee perception gap: 41-67 points across 7 dimensions.

  • Gartner — Build an Adaptive Marketing Collective for the AI Era (CMO Quarterly 1Q26) — Sibling excerpt to the Strategic Impact of AI Agents article from the same 1Q26 issue, focused on organizational design. Replace hierarchical marketing org charts with a “marketing collective” of synchronized resources (employees + agencies + contractors + vendors + AI agents) organized around four intersecting responsibilities — Strategy, Operations, Brand, Digital. Five-step framework: Reimagine → Reorganize → Redefine → Restart → Reenergize. Step 1 names hybrid human-and-AI-agent teams as “fast becoming a reality.” Step 2 redraws external agencies inside the collective. Step 5 introduces “Agentic experience” as a first-class CX surface alongside Stakeholder/Platform/Physical. Empirically grounded in seven years of Gartner client inquiry calls + hundreds of org-chart reviews (2018–2023 baseline, refreshed for 1Q26).

  • Gen Z AI Resistance — Sabotage, Sentiment, and the Executive-Worker Gap (2026) — Secondary-aggregation analysis (YouTube creator El, claimed PhD CS) drawing on multiple primary studies. Headlines: 44% of Gen Z workers admit actively sabotaging company AI strategy (entering proprietary info into public chatbots, refusing mandated tools, deliberate low-quality output, tampering with performance reviews). Gen Z sentiment trend over past year: excitement -14pts → 22%, hopefulness -9pts → 18%, anger +9pts → 31%; daily AI users show even larger drops (“the more they use it, the less hopeful they become”). ~6,000-executive multi-country survey: 90% of executives report AI has had no impact on employment or productivity at their firms over past 3 years, 75% admit AI strategy is “more for show than meaningful guide to outcomes,” 73% of CEOs report stress/anxiety about their AI strategy, 64% fear job loss if they fail to lead AI transition, 54% admit AI deployments are “tearing their company apart.” Executive-worker usage gap: 64% of executives use AI 2+ hrs/day vs only 28% of regular employees; 92% of executives “actively cultivating an AI elite.” Penn = first Ivy League to launch AI major; student-paper editorial “Penn has an AI problem” opens with “AI cannot coexist with education. It can only degrade it.” Author’s thesis: this is not resistance to technology but resistance to a specific power structure being built with technology. Four open primary-source verifications flagged (which survey is the 44% from, which executive survey reports 90%-no-impact, Penn editorial date, which sentiment tracker reports the Gen Z deltas).

  • Gartner — The Strategic Impact of AI Agents (CMO Quarterly 1Q26) — Gartner’s CMO-facing framing of agentic AI for 2026 (8-page CMO Quarterly excerpt, updated 2026-01-27). Three Gartner frameworks ship: AI Agent Assessment Framework (5-level Minimal-to-Advanced spectrum placing chatbots → assistants → agents); Levels of Agent Capabilities (the same 5 levels × 6 capability dimensions — Perception / Decisioning / Actioning / Agency / Adaptability / Knowledge); Competitive Vendor Landscape (four quadrants — hyperscalers, consultants, new specialist agentic companies, enterprise application BOAT). Five marketing-process targets ranked for agentic adoption (customer journey orchestration, workflow optimization, competitive research/customer insight, scenario/strategic planning, content/campaign creation). Cost drivers (reasoning steps, context size, deployment + license model, AI data readiness). CMO call to action: invest in API foundations now — MCP + A2A protocols will drive more APIs, not fewer. Marketing positioned as internal pilot environment for agentic experimentation. Includes WEO Marketly applied read mapping the five Gartner targets to OmniPresence / BAW / Clawdbot / Hermes / GHL surfaces.

  • Pew Research — How Americans View AI and Its Impact (Sept 2025) — National survey, n=5,023 US adults (American Trends Panel, fielded June 2025). 95% awareness but 50% more concerned than excited; majorities expect AI to worsen creative thinking and meaningful relationships; most cannot reliably detect AI-generated content. The public-perception layer: the audience is large and aware but anxious, not enthusiastic. Surfaced in the AI SEO hub’s user-behavior section as the sentiment context beneath the citation-tactics and adoption data.

  • Are Mythos’ Cyber Capabilities Overhyped? (Epoch AI Cyber-ECI Analysis) — Independent aggregation of ~15 cyber benchmarks into a domain-specific Cyber-ECI, testing Anthropic’s “leap in cyber skills” claim for the Mythos family (Mythos Preview + Fable 5). Splits “cyber” into exploit development (confirmed large jump — Mythos Preview ~7 months ahead of the early-2025 trend, well past GPT-5.5; Mythos 5 modestly more) vs vulnerability discovery (gain is unclear on a fixed budget — the Project Glasswing CVE spike of +142%/+262% over baseline is confounded by ~$100M in API credits; prior models and even small open models were already strong finders). Mythos’s genuine discovery edge: lower false-positive rate + better severity prioritization. CyScenarioBench: Mythos 5 36.7% / Mythos Preview 29.2% / GPT-5.5 26% / Opus 4.8 16.6%. Verdict: not “just hype,” but the leap is concentrated in exploitation, not discovery. The independent third-party counter-read to the wiki’s first-party system-card cyber coverage. (Epoch’s opinionated Gradient Updates series; data rigorously sourced, conclusions the authors’ own.)

  • MirrorCode — Epoch AI + METR Long-Horizon Coding Benchmark — answers “What’s the largest software project AI can complete on its own?” 25 real-world programs (bioinformatics, Unix utilities, cryptography, interpreters) rebuilt with no source code and no human in the loop; the hardest would take a human engineer weeks-to-months. Unlike typical SWE benchmarks (~2,600 for a single run, the model working 19 days unattended**. Claude Opus 4.7 leads at a 56% solve rate (significant headroom remains). Source: Epoch Brief newsletter (2026-06-26); primary epoch.ai/mirrorcode.

  • Terminal-Bench — Benchmarking AI Agents in the Terminal — Stanford + Laude Institute open benchmark for terminal/CLI agents: each task is a Docker sandbox + natural-language instruction + end-state verification tests + reference solution; a task is solved only if every test passes. 1.0 = 80 tasks, 2.0 (Nov 7 2025) = 89 curated hard tasks (frontier <65% at launch), 2.1 = team-verified variant, 3.0 in progress. Runs via the Harbor harness; Terminus is the minimal reference agent. Broader and harder than SWE-bench (operate a computer vs. fix a GitHub issue) — the wiki cites both. Verified 2.1 board (2026-07-03): Codex CLI + GPT-5.5 (83.4%) and Claude Code + Fable 5 (83.1%) tie for the lead, Opus 4.8 at 78.9%, Gemini ~70-74%.

  • Anthropic Economic Index — Cadences (June 2026) — New continuous privacy-preserving telemetry (samples a daily slice vs prior 7-day snapshots) surfaces daily/hourly Claude usage rhythms: work queries dip on weekends (less so in the highest-paid occupations), news in the morning, sleep advice peaks ~5 a.m., tax requests surge around filing deadlines. First results from the Anthropic Economic Index Survey (launched April 2026; ~9,700 linked respondents): people who use Claude in the most automated way expect AI to take on more of their tasks next year, yet feel the most optimistic about pay, job security, and meaning. Official Anthropic report; the chat→agentic-task shift (Claude Code + Cowork) drove the methodology change.

  • Luna) — OpenAI’s Three-Tier Frontier Family — OpenAI’s competitor family, GA’d 2026-07-09 (exited the ~20-org preview, now bundled into a unified ChatGPT Work app). Verified per-1M pricing: Sol 30, Terra 15, Luna 6 — Sol undercuts Claude Fable 5’s 50 by 50% in / 40% out. Framed for model/cost routing; a safety-classifier gating layer parallels Fable 5’s topic-gated Opus 4.8 fallback. Cyber-forward (ExploitBench competitive with Mythos Preview at ~1/3 the tokens); GA week also surfaced a UK AISI universal-jailbreak finding, flagged as a contradiction against the article’s earlier red-teaming claim.

  • Grok 4.5 (xAI) — xAI’s first model from its new SpaceX/Cursor collaboration, “built for real world engineering.” 4th on the Artificial Analysis Index (behind Fable 5, Opus 4.8, GPT-5.5) but by far the most cost-efficient near-frontier model (1.80 and Fable’s $2.75); state-of-the-art on AutomationBench. Early-adopter framing: a cheap implementation-agent model to run under a Fable 5 / GPT-5.6 orchestrator, not a flagship replacement.

  • The AI Competitive Scoreboard Is Shifting From Model Quality to Infrastructure, Distribution, and Political Permission — Nate B Jones analysis (early July 2026): after ~2 years of “best model wins,” Meta (monetizing excess compute via “Meta Compute,” shipping a consumer prompt-to-game app, admitting its agent timeline slipped), OpenAI (floating a ~5% government equity stake amid a frontier-model pre-release-review mandate), and Anthropic (doubling down on enterprise distribution via Claude Tag + forward-deployed engineers) are all visibly competing on axes other than benchmark leadership. Practical read: track infrastructure/distribution/regulatory-posture moves alongside benchmarks, and treat staggered/delayed frontier releases as a recurring structural risk, not a one-off.

  • Tech Workforce AI Sentiment Survey 2026 (Lenny Rachitsky × Noam Segal) — Second annual N≈6,000 practitioner survey (product/eng/design/research/marketing): AI reshapes professional identity for 97% but in opposite directions (50% amplified / 27% redefined / 14% destabilized / 5% diminished — the “bifurcation”), significant burnout up 44.7%→54.7% YoY, career optimism down 54.8%→48.7%, 72% worried about layoffs; four archetypes (energized 41%, conflicted 35%, disoriented, resentful 12%); negative recommend-this-career NPS in every role; manager quality (~25% rated highly effective) as the under-funded adoption lever. Worker-side counterpart to the executive-side Gen Z AI Resistance survey.

  • Mozilla’s State of Open Source AI Report (v1.0, 2026) — Mozilla Foundation’s first State of Open-Source AI report (interview with portfolio-wide CTO Rafi Coreman): open-weight models have reached rough parity with closed frontier models for everyday (non-frontier) use, but the real contested ground has shifted from the models to the agentic harness wrapped around them (OpenCode, Hermes). Concrete data points: Chatbot Arena gap closed from 8 points (Jan 2024) to 0.5 (Aug 2024) before partially reopening; Pinterest saved an estimated $10M/quarter self-hosting; 70% of developers who tried self-hosting an open-weight model abandoned deployment before finishing. Frames open-weight adoption as a sovereignty question following the “Mythos shut off” episode.

  • Ling-3.0-flash — MIT Open-Weights Agentic Coding Model That Ships for Claude Code — inclusionAI (Ant Group), open-weighted 2026-08-04 under a plain MIT licence rather than a custom community one (verified against the HuggingFace repo). The notable part is distribution, not the benchmark: the model card names Claude Code, Kilo Code, Qwen Code, Hermes Agent, and OpenClaw as target harnesses, with a native ling3 tool-call parser and a vLLM fork running --enable-auto-tool-choice — open-weights labs shipping for the agent harnesses people already run. 124B total / 5.1B active / 256K context; vendor-reported SWE-bench Pro 56.6 and Multilingual 72.4 (unverified). Serving needs 4 GPUs BF16 / 2 FP8 with no GGUF at release, so it is an OpenRouter or Baseten line item in practice. Community usage pattern: keep Opus 5 planning, point the executor at it.

  • Kimi K3 (Moonshot AI) — Moonshot’s 2.8-trillion-parameter open-weight model, positioned as the first open model at “Fable level”: on DeepSWE reported just behind Fable 5 and GPT-5.6 Sol but 20+ points over Opus 4.8 and GLM 5.2, and beating Fable 5 on Terminal-Bench. Second headline is cost-efficiency (upper-left of the capability/cost frontier). Behaves like a long-horizon agent (30–40 min runs, self-verifies via Chrome screenshots, 400K–17M tokens/task) and answers biomedical/vision prompts Fable refuses. Weights shipped 2026-07-28 (see the article’s dated update: modified-MIT license with a distillation ban extending to governments; architecture names corroborated via a paper review); API + Kimi Code harness + kimi.com chat live. Originally sourced from two secondary creator videos; the DeepSWE-margin contradiction (20+ vs 8.5 points over Opus 4.8) remains unresolved — status: contradicted.

  • Chinese Open-Weight Models — A Vendor-Neutral Decision Framework (Nate B Jones) — Against the “China has caught up” shorthand: yes, serious users should test Chinese models, but make three separate decisions — task, model, deployment path — and measure cost per accepted result, not benchmark position. Anchors preserved: DeepSeek 20M-revenue authorization license, Anthropic’s ~24,000-fraudulent-accounts claim, and a 4-condition self-host test. Creator’s Ringer plug flagged as bias.

  • Hugging Face Sandbox-Escape Incident (July 2026) — Synthesis of seven secondary sources of sharply uneven rigor, written to separate what is established from what is framing. Established: a model in a cyber-capability evaluation spent substantial inference compute finding internet access, chained a zero-day in a self-hosted package-registry cache proxy into privilege escalation and lateral movement, and reached Hugging Face servers — with production classifiers deliberately disabled to measure maximum offensive capability, which is why the result does not transfer to consumer models. Hugging Face self-detected and self-contained. The defender’s problem is the sharpest takeaway: ~17,000 events to triage, and hosted OpenAI and Anthropic models refused the exploit payloads, so Hugging Face ran GLM-5.2 locally instead. Model identity is never asserted — OpenAI said only “more capable than GPT-5.6 Sol.” Reward hacking is the load-bearing frame. Pairs with Epoch’s Cyber-ECI analysis, which argued the exploit-development jump was real and predictable.

  • ProPublica — Mythos vs Microsoft’s Patching Capacity — From a recording of a mid-May Microsoft meeting plus internal documents: Mythos Preview under Project Glasswing (~50 Microsoft engineers) is finding security bugs faster than Microsoft can patch them — “a mad dash,” per one manager. The adversarial-journalism counterpart to Anthropic’s own defensive framing, and the find-fix asymmetry of the Mythos capability results made organizationally concrete.

  • Open Secure AI Alliance — NVIDIA’s Open-Model Security Coalition — Launched 2026-07-27 by NVIDIA and roughly three dozen partners “to develop and share open technologies, techniques and tools to safeguard software and agents in the age of AI,” framed explicitly around the Hugging Face sandbox-escape incident — the argument being that defenders need open, inspectable frontier agentic systems they can run on their own infrastructure. Nous Research (Hermes Agent) is an inaugural member.

  • The Compute Economics of the AI Buildout (Epoch AI, 2026) — consolidates Epoch’s 2026 Gradient Updates on compute supply/demand into one map for cost and capability forecasting: is a compute crunch coming, what share of compute frontier labs actually use, HBM’s rise from 52%→63% of memory demand, datacenter-vs-GDP scaling, and the capex cliff (buildout on track to overtake hyperscalers’ operating cash flows by end-2026, forcing external financing). Includes the chip-smuggling-to-China estimate and the read from 1,604 Chinese AI job postings. Signed-opinion series — attributed to named Epoch authors, not institutional positions.

  • The Future of AI Benchmarks (Epoch AI, 2026) — why classic reasoning benchmarks are saturating and what replaces them. Standout finding: EBR-bench shows current models barely improve by learning from their own repeated attempts — the self-improvement loop is weaker than hype implies. Carries the “relax one classic constraint (short-horizon, clean-grading, single-attempt)” recipe for building next-gen benchmarks, MirrorCode as the worked long-horizon example, and the model-routing datapoint that Claude overperforms on SWE-style tasks and underperforms on math. Pairs with MirrorCode and Terminal-Bench.

  • OpenAI Withholds Astra — First Model to Hit the ‘Critical’ Cyber Threshold — On 2026-08-07 OpenAI said it “cannot rule out” that its unreleased next-generation model Astra has Critical cyber capabilities under its own Preparedness Framework, and held it back — the first time any OpenAI model has been publicly placed at that threshold. Downstream of the Hugging Face escape and of Michael Dalton’s “consciously slowing down research” line at Black Hat two days earlier. Covers the threshold definition, the internal-safeguard list (isolated test environments, restricted network/tool access, weight encryption, sandboxed execution), the June 2026 executive order that makes government pre-release access voluntary and non-binding, and the same-week claim that an internal Astra solved ten decades-old math problems.

  • Eight Predictions for the Era of Continual Learning (Dwarkesh Patel) — Argues file-based memory cannot substitute for accumulated experience (the saxophone analogy) and works through eight consequences: pre-deployment safety gates stop being coherent, alignment research aimed at frozen weights is aimed at the wrong object, AI minds diversify, being ahead compounds, the internal-to-public deployment gap collapses, continual learning becomes the moat labs currently lack (switching vendors = firing an employee with months of context), labs subsidise or withhold access to buy training rights, and batching creates inference economies of scale. Directly undercuts the model-portability assumption behind this wiki’s multi-vendor routing advice.

29 items under this folder.