Source: zyppy-ai-citation-ranking-factors-2026-05-07.md + raw/newsletter-zyppy-signal-203893f8bc.md — Cyrus Shepard, Signal by Zyppy. Published 2026-05-07.
Corrected 2026-07-24
The 23-factor table and several takeaways were rebuilt from the recovered primary source. A prior version of this article carried fabricated middle-tier factors (e.g. “Direct Quote Density,” “Page Speed,” “Mobile UX,” “Anchor Text,” “Image Alt Text,” “Backlinks”) and invented per-engine ChatGPT/Gemini/Perplexity score columns that do not appear in Shepard’s post. Shepard publishes a single score per factor (9.5 down to 2.0), not per-engine scores. The table below now matches the source exactly.
Cyrus Shepard (Zyppy SEO) downloaded nearly every published AI-citation experiment, study, explainer, and patent from the prior ~2 years (across ChatGPT, Gemini, and Perplexity), narrowed to the 54 most salient sources, cross-referenced their findings, and hand-scored a 23-factor ranking on a single 0-10 scale. The strongest signal: AI citation engines re-rank on top of classical search relevance, so winning organic SEO is the precondition. The most-hyped 2025 tactics — schema and LLMs.txt — score near the bottom (5.6 and 2.0). Shepard’s thesis: “win SEO, win AI citations (most of the time, with extra steps).”
This is a meta-analysis / secondary synthesis, not new primary data. Scores are Shepard’s manual judgment, weighted by his three criteria (below); he states plainly “these aren’t ‘Ranking Factors’ in the traditional sense… Correlation is not causation.” Treat it as an evidence map over the cluster, not as an independent study.
Key Takeaways
- Top tier (9.0+): URL Accessibility (9.5 — can the bot reach the page?), Search Rank (9.4 — does it already rank on Google?), Fan-out Rank (9.3 — does it rank for the query’s fan-out expansions?), Preview Control (9.2 —
nosnippet/data-nosnippetcan suppress your own visibility), Query-Answer Match (9.2 — page content is semantically close to the query and the answer), Intent-Format Match (9.0 — listicle for “best,” step-by-step for “how-to”). - Upper-mid tier (8.0-8.9): Topic Cluster Ranking (8.9), Answer Near the Top (8.8 — Gemini applies a strict per-URL retrieval cap, so top-of-page content is likelier to be cited), AI-ready Structure (8.6 — headings/sections/tables), Factually Specific (8.3), Explicit Phrasing (8.1 — definitive claims beat hedged ones), Cites Sources (8.0), Self-Contained Passages (8.0).
- Lower-mid tier (6.3-7.6): Content Visibility (7.6 — visible HTML text, not JS-hidden), Freshness (7.0), Brand / Entity Trust (6.8), Length (6.7 — longer tended to help but evidence was inconsistent, and length reduces the share of a page that gets retrieved), Language (6.3 — engines bias toward the query’s language/locale).
- Bottom tier (under 6): Entity Consistency (5.8), Structured Data (5.6), Known Source (5.4 — already in the model’s training data), Domain Authority (5.0 — relationship found but “often weak”), LLMs.txt (2.0 — no credible evidence it influences citations at all; the most overhyped tactic of 2025-26 by Shepard’s read).
- Structured Data scores 5.6 (#20 of 23). Shepard’s read: “practically every study that looks at schema and AI citations finds a positive relationship. The effect is typically small, but it’s amazingly consistent.” Empirically corroborated by the Ahrefs schema causal study (2026-05-11), which found no statistically meaningful AI-citation lift from adding JSON-LD — consistent with schema being a small, correlational marker rather than a lever.
- The through-line is classical SEO. Shepard’s own summary: the factors reduce to Relevance, Trust, Topical Authority, and Extractability — signals that “should align with current SEO thinking.” “Win SEO, win AI citations (most of the time, with extra steps).”
- Why a meta-analysis and not a single study. Each of the 54 underlying studies has small-sample or selection issues; cross-referencing which findings recur across studies, surfaces, and methodologies is more robust than trusting any one case study — at the cost of the scores being subjective aggregate judgments rather than measured effect sizes.
The 23 Ranking Factors
Scored on three criteria: Repeatability (how often a similar finding recurs across studies + consistency of direction), Strength of Evidence (a 50-million-query study outweighs a 10-query case study), and Official Support (docs, technical specs, patents). Shepard assigned each score by hand, using AI to fine-tune. A single score per factor; no per-engine breakdown exists in the source.
| # | Factor | Score | Definition |
|---|---|---|---|
| 1 | URL Accessibility | 9.5 | Page is available and crawlable during training/grounding |
| 2 | Search Rank | 9.4 | How the URL ranks for the exact query |
| 3 | Fan-out Rank | 9.3 | How the URL ranks for related fan-out queries |
| 4 | Preview Control | 9.2 | Preview directives (nosnippet, data-nosnippet) can suppress visibility |
| 5 | Query-Answer Match | 9.2 | Page content closely matches the query (primary/fan-out) |
| 6 | Intent-Format Match | 9.0 | Page type matches query intent (listicle for “best,” etc.) |
| 7 | Topic Cluster Ranking | 8.9 | Site ranks for multiple related queries (primary + fan-out) |
| 8 | Answer Near the Top | 8.8 | Content near the top of the page is likelier to be cited |
| 9 | AI-ready Structure | 8.6 | Formatted so AI can extract sections (headings, tables) |
| 10 | Factually Specific | 8.3 | Shows specific, verifiable facts |
| 11 | Explicit Phrasing | 8.1 | Definitive claims over vague/hedged statements |
| 12 | Cites Sources | 8.0 | Facts backed with referenced sources |
| 13 | Self-Contained Passages | 8.0 | Key statements stand alone without extra context |
| 14 | Content Visibility | 7.6 | Important text in visible HTML, not hidden/JS-gated |
| 15 | Freshness | 7.0 | How current the information is (varies by query) |
| 16 | Brand / Entity Trust | 6.8 | How much the engine knows about and trusts the brand |
| 17 | Length | 6.7 | Word count (longer tended to help, but inconsistent) |
| 18 | Language | 6.3 | Language/locale of the content vs the query |
| 19 | Entity Consistency | 5.8 | Consistent naming for brands, people, products |
| 20 | Structured Data | 5.6 | Schema to identify entities and support content |
| 21 | Known Source | 5.4 | URL already known to the engine via training data |
| 22 | Domain Authority | 5.0 | Link-based popularity measure (relationship “often weak”) |
| 23 | LLMs.txt | 2.0 | Hosting an LLMs.txt file (no credible supporting evidence) |
Tactical Implications
- Win SEO first. Search Rank is #2 (9.4). If you aren’t ranking organically, AI citations are near-impossible regardless of everything else. This collapses the “AEO is separate from SEO” narrative. Supporting data Shepard cites: Ahrefs found 38% of AI Overview citations come from Google’s top 10; AirOps found a strong retrieval-rank→ChatGPT-citation relationship; Semrush found Perplexity answers had ~82% overlap with Google’s top 10.
- Optimize for query fan-out. Fan-out Rank is #3 (9.3). Engines expand the query into sub-queries; pages ranking for the head term and its expansions get cited more. Cluster content around topical concepts, not just exact-match keywords. (See the companion Zyppy AI Citation Playbook for Shepard’s step-by-step fan-out workflow.)
- Control your preview. Preview Control (#4, 9.2) — a
nosnippet/data-nosnippeton important text can suppress your own AI visibility. Audit for accidental suppression. - Put the answer near the top and make passages self-contained. Answer Near the Top (8.8) + Self-Contained Passages (8.0): engines don’t retrieve the whole page (Dan Petrovic’s Gemini retrieval-cap research), so front-load the citable claim and make it stand alone.
- Format-match the intent. Intent-Format Match (#6, 9.0). “How to” → numbered steps; “what is” → definition + example; “compare” → table. Wrong format = no citation even when the content is correct.
- Skip LLMs.txt. Score 2.0 (#23). No credible evidence. Reallocate the hours to factors 1-9.
- Schema for the right reasons. Score 5.6 (#20). Add schema because Google still rewards it on classical surfaces — not for AI-citation lift. The Ahrefs causal study is the independent confirmation.
Open Questions
- The exact 54 underlying studies. Shepard links a public spreadsheet of all 54; worth pulling for citation-chain verification.
- How scores decay over time. Engine retrieval changes monthly. The 2.0 for LLMs.txt assumes engines aren’t using it as of May 2026 — if any engine starts honoring it, this jumps.
- Weighting by vertical. All 23 factors are averaged across studies. Per-vertical weights (medical, legal, e-commerce) would likely shift the ranking — medical content is cited more conservatively, so Brand/Entity Trust and Topical Authority probably weight higher there.
Related
- Zyppy AI Citation Playbook (Fan-out Framework + 7-Step Audit) — Shepard’s actionable companion: how to operationalize factors 2-6 (Search Rank, Fan-out Rank, Preview Control, Query-Answer Match, Intent-Format Match). This meta-analysis is the “what matters”; the playbook is the “how.”
- Ahrefs Schema → AI Citations Causal Study — Matched-DiD study that empirically confirms Shepard’s 5.6 for Structured Data: no statistically meaningful lift from adding schema.
- AirOps + Kevin Indig Fan-Out Effect ChatGPT Study — Largest single-engine dataset. AirOps’s retrieval-rank finding (rank-1 cited 58.4% vs rank-10 14.2%) is the deepest empirical backing for Shepard’s 5 weights.
- Digital Applied 1,000 AIO Citation Pattern Study — Concrete AIO-only correlational data on schema lift (2.3×-2.8× after DA control) — a much larger magnitude than Shepard’s meta-analytic 5.6, illustrating correlational-vs-meta-analytic divergence.
- GEO-16 Framework (arXiv 2509.10762v1) — Academic cross-sectional study; likely one of Shepard’s 54 underlying sources.
- FLUQs Framework — The “core SEO first” thesis. Shepard’s 23 factors are the empirical backing.
- Google’s Generative AI Search Optimization Guide — Google’s official position aligns: AI Overviews + AI Mode use the same Search index, so AI search is still SEO.
Try It
- Pull your own top-10 ranking pages from GSC. Cross-reference against the 23 factors. The 1-2 lowest-scoring factors on your pages are your highest-leverage fixes.
- Audit title tags and meta descriptions for Preview Control (#4). Rewrite any preview that doesn’t literally answer the page’s primary query, and check for accidental
nosnippet/data-nosnippeton important text. - Check fan-out coverage. Run your head term through Google AI Mode / a fan-out tool and capture the expanded sub-queries. Does your page address them? If not, expand sections or add an FAQ block. (Full workflow in the playbook.)
- Stop new LLMs.txt projects (score 2.0). Reallocate to factors 1-9.
- Re-frame schema work. Keep existing schema (still helps classical Google). Don’t expand it chasing AI citations.