Source: ahrefs-schema-ai-citations-study-2026-05-11.md — Louise Linehan & Xibeijia Guan, reviewed by Ryan Law. Published 2026-05-11 on the Ahrefs Blog.
Ahrefs ran a difference-in-differences causal study tracking 1,885 pages that added JSON-LD schema between August 2025 and March 2026 against a 4,000-page control set. They measured whether the schema additions moved AI citations across Google AI Overviews, Google AI Mode, and ChatGPT in the 30 days following the change. Headline: schema did not produce a statistically meaningful lift on any of the three surfaces. AI Overviews showed a 4.6 percent decline that was statistically significant but inside a larger declining trend, so the authors stop short of calling it a schema-caused drop. It is Ahrefs’s follow-up to its own correlational study of 6 million URLs, which found AI-cited pages almost three times more likely to have JSON-LD than non-cited pages. Scope matters: every studied page already had 100+ AI Overview citations before treatment, so the findings apply only to pages “already inside the consideration set.”
Per-surface results from the matched difference-in-differences test, which Ahrefs calls “the test we trust most”:
| AI surface | Effect on citations vs. matched controls | Verdict |
|---|---|---|
| Google AI Overviews | −4.6% | Small but statistically significant decline. Both groups were already declining; treated pages fell slightly faster. |
| Google AI Mode | +2.4% | Statistically indistinguishable from zero. |
| ChatGPT | +2.2% | Statistically indistinguishable from zero. |
Key Takeaways
- Matched difference-in-differences design. 1,885 pages that added JSON-LD (Aug 2025-Mar 2026), matched against 4,000 control pages. For each treated URL, Ahrefs picked 3 control URLs from different domains, with similar pre-period citation levels, that had never added JSON-LD. Citations were counted 30 days before and 30 days after the treatment date.
- Applies only to pages already receiving AI citations. Studied pages had 100+ AI Overview citations pre-treatment (the source dates this to February 2025). The study does not test whether schema helps a page that isn’t cited yet get into the consideration set.
- Three surfaces, no significant positive effect. None showed a statistically significant positive effect (AI Mode +2.4%, ChatGPT +2.2%, both indistinguishable from zero). In Ahrefs’s words, “they can’t tell whether the schema did a tiny bit of good or nothing at all.”
- AIO −4.6% is significant but small, and may be coincidence. It works out to around 12 daily citations lost per page, in a sample where most pages were getting hundreds; the odds of a gap that large by chance are about 1 in 2,500. Treated and control pages were both “already on a steep downward trajectory” before schema was added. Ahrefs says the gap suggests a small negative effect “but it could also just be coincidence” and it “cannot tell which from this data alone.”
- All schema types pooled. No per-type analysis was run. Ahrefs lists “Article, FAQ, Product, HowTo, Organization …” and flags the gap itself: “It’s possible some types help more than others. This may be worth digging into.” Only JSON-LD was tested, not Microdata or RDFa.
- Consistent with FLUQs + Google’s official position, for already-cited pages. Both FLUQs and Google’s Generative AI Search Optimization guide argue that AI search ranks on the same core signals as classical search. Ahrefs’s null result on already-cited pages fits that view.
- Practical implication (pages already receiving AI citations). Ahrefs: “If you’re already doing the rest of the SEO work well, JSON-LD isn’t going to be the unlock.” Schema still has value for “rich results, voice assistants, knowledge graphs, downstream entity recognition.” Don’t add schema to an already-cited page expecting an AI-citation lift.
- Why Ahrefs ran it: to test its own correlation. The earlier 6M-URL finding (cited pages ~3x likelier to have JSON-LD) “gets shared in LinkedIn carousels and conference slides as proof that schema is an AI visibility lever.” Ahrefs’s caution: schema “tends to live on better-maintained, more technically sophisticated sites”, so the 3x could be “schema riding the wave of every other signal.” This study was designed to isolate the effect of adding schema.
Ahrefs causal null vs. four correlational studies showing schema lift
Ahrefs says (this article, matched DiD, 1,885 pages, three surfaces) — adding schema produces no statistically meaningful AI citation lift on pages already receiving AI citations (100+ AIO citations pre-treatment). Four correlational studies say schema-using pages ARE cited more:
- AirOps (16,851-query ChatGPT study, stratified): +6.5pp citation advantage for JSON-LD pages (38.5% vs 32.0%).
- Digital Applied (1,000-AIO study, regression-style DA control): 2.3× lift for Article + BreadcrumbList; 2.8× for HowTo.
- GEO-16 arXiv (1,702 citations across Brave/AIO/Perplexity, cross-sectional): Structured Data r=0.63, +39% citation impact, 95% CI [0.59, 0.67].
- Zyppy meta-analysis (54 studies aggregated): Structured Data 5.6 / 10 (#20 of 23) — mid-tier but present.
Reconciliation: All four correlational studies cannot match on unobserved publisher characteristics (editorial maturity, technical SEO depth, content team composition) the way Ahrefs’s matched DiD does. The most parsimonious interpretation^[inferred]: schema is a marker of editorial / technical / publication-infrastructure maturity that correlates strongly with citation, not the lever itself. Adding schema to an existing page (Ahrefs’s intervention) doesn’t cause the lift because the page’s underlying characteristics determine citation candidacy. Practitioner: keep your schema (still helps Google’s classical surfaces — rich snippets, knowledge panel eligibility — and is the right table-stakes signal per AirOps); don’t expect adding it to an already-cited page to be the lever that moves AI citations. Whether adding schema helps a page that isn’t cited yet has no causal test here: Ahrefs’s design excludes such pages, and the four studies above are correlational. ^[inferred] Status: resolved (2026-05-19) — methodological-difference, not factual.
Study Design Details
- Pages tracked: 1,885 treated pages; 4,000 control pages.
- Treatment definition: Pages that introduced JSON-LD between Aug 2025 and Mar 2026, found by examining HTML history in Ahrefs’s crawler database. The treatment date is the first day the crawler detected JSON-LD, after the last check that found none.
- Control matching: For each treated URL, 3 control URLs from different domains, with similar pre-period citation levels, that had never added JSON-LD.
- Sample scope: Pages with 100+ AI Overview citations pre-treatment (the source dates this to February 2025).
- Outcome: Citation counts in the 30 days before and 30 days after the treatment date on Google AI Overviews, Google AI Mode, and ChatGPT, tracked with Brand Radar (Ahrefs’s AI citation tracker).
- Four statistical tests: a two-sample t-test; difference-in-differences, which strips out platform-wide trends (“the test we trust most”); an event study plotting week-by-week citations to check the two groups tracked together before treatment; and a symmetrical-window DiD that excludes recrawling periods. The per-surface numbers above come from the matched DiD.
- Schema types: All types pooled; no per-type analysis. JSON-LD only (no separate Microdata or RDFa analysis).
- Caveats Ahrefs flags:
- Pages had 100+ AI Overview citations before treatment, so findings apply only to pages “already inside the consideration set.”
- An external searchVIU experiment found “every system extracted only visible HTML content. JSON-LD, hidden Microdata, and hidden RDFa were all ignored” during direct retrieval.
- Schema often coincides with other simultaneous changes (links, content updates, technical fixes).
- The 30-day measurement window may miss slow-burn effects.
- Only HTML-embedded schema was tested; JavaScript-injected markup is untested.
Open Questions
- Long-window effect. Ahrefs flags that the 30-day window may miss slow-burn effects. The source reports no longer-window result.
- Per-type effects. Ahrefs pooled all schema types and says some types may “help more than others. This may be worth digging into.”
- AEO-specific schema (Speakable, ClaimReview, QAPage). Ahrefs’s list of pooled types doesn’t name the AEO-specific subtypes that some practitioners argue are the actual lever.
- Pages not yet cited. The sample was limited to pages with 100+ AI Overview citations. Whether schema helps a page that isn’t cited yet is untested.
- JavaScript-injected schema. Untested; only HTML-embedded JSON-LD was studied.
- Control-set size. The source gives both “4,000 control pages” and 3 control URLs per treated URL (1,885 × 3 = 5,655), which suggests some controls were matched to more than one treated page. The source doesn’t say.
Related
- AirOps + Kevin Indig Fan-Out Effect ChatGPT Study — Largest single-engine correlational study. +6.5pp schema lift in stratified analysis. Same direction-of-effect, methodologically weaker than Ahrefs’s matched DiD.
- Digital Applied 1,000 AIO Citation Pattern Study — AIO-only correlational study with regression-style DA control. 2.3× lift claim. See contradiction callout for reconciliation.
- GEO-16 Framework (arXiv 2509.10762v1) — Academic cross-sectional study across Brave/AIO/Perplexity. Structured Data r=0.63, +39%, 95% CI [0.59, 0.67]. Authors explicitly self-flag observational design.
- Zyppy AI Citation Ranking Factors — Cyrus Shepard’s parallel meta-analysis of 54 studies. Scores Structured Data 5.6 / 10 (#20 of 23 factors). The most-aggregated view across the literature.
- FLUQs Framework — Core SEO first, AEO-specific tactics second. Ahrefs’s null result on already-cited pages is consistent with it.
- Google’s Generative AI Search Optimization Guide — Google’s official position: AIO/AIM rely on the same Search index, so AI search optimization is core SEO. Ahrefs’s study is consistent with it for pages already receiving AI citations.
- Similarweb Most-Cited Domains in LLMs — What domains LLMs actually cite, useful for benchmarking what “winning AI citations” looks like.
Try It
- Don’t pull schema if you have it. Ahrefs says it retains value for rich results, voice assistants, knowledge graphs and downstream entity recognition, and the one negative result (AIO −4.6%) “could also just be coincidence.”
- On pages already receiving AI citations, don’t add schema primarily to chase more. The study found no lift there. Spend the hours on higher-ranked factors from Zyppy’s analysis instead.
- Run Ahrefs’s own small test before committing budget: pick 5-10 test pages with baseline AI citations, match them to 5-10 control pages with similar metrics, add schema to the test pages only, record baseline citations across platforms, and compare movement after 30+ days (Ahrefs runs this in Brand Radar).
- Scope what you tell a client or stakeholder. For pages already cited in AI answers, a matched 2026 test found adding JSON-LD didn’t boost citations. It did not test pages outside the consideration set. Ask schema-focused vendors whether their case studies used matched controls.