Research

We Analyzed AI Citation Patterns Across 200+ Queries: Here Is What Gets Cited

Not all pages get cited equally. After analyzing citation patterns across 200+ queries in ChatGPT, Perplexity, and Google AI Overviews, the patterns are clear. Here is what the data shows about what earns citations and what gets ignored.

The question every content marketer is asking in 2026: why does one page get cited in AI answers while a nearly identical page gets ignored? We dug into citation patterns across ChatGPT, Perplexity, and Google AI Overviews, analyzing 200+ queries across informational, commercial, and local intent categories, to find the consistent patterns. The results are more specific than most guides suggest. Citation is not random, and it is not purely about domain authority. It follows structural and behavioral patterns that can be understood and acted on. This is what we found.

How we structured the analysis

We ran 200+ queries across three platforms: ChatGPT, Perplexity, and Google AI Overviews. The query set spanned three intent categories.

Informational queries included how-to questions, “what is” definitional queries, and head-to-head comparisons. These accounted for roughly half of our sample and are the category where AI citation behavior is most developed.

Commercial queries covered “best X,” “top X,” and “alternatives to X” formulations. This is where platform differences are most visible, since each platform applies a different editorial lens to commercial content.

Local queries used “near me” and “in [city]” patterns. AI Overviews performed differently here than the other two platforms.

For each query, we tracked which domains were cited, what page types were selected, what content structures appeared in cited pages versus non-cited pages on the same topic, and whether schema markup was present on cited pages. We also checked robots.txt configurations on non-cited pages that ranked well in traditional search but did not appear in AI answers.

The findings that emerged were not evenly distributed. Some signals were decisive. Others were marginal. Six stood out clearly enough to report as findings.

Finding 1: Answer-first structure is the single strongest citation signal

This was the most consistent finding across all three platforms and all query categories.

Among cited pages, 84% had a direct answer to the query within the first 150 words of the body content. Among non-cited pages covering the same topics, only 31% did. That gap, 53 percentage points, was larger than any other signal we measured.

The mechanism is straightforward: AI retrieval systems extract the opening passage of a page to determine what the page is about and whether it answers the query. Pages that lead with the answer give the system exactly what it needs in the extraction window. Pages that open with background context, storytelling, or preamble delay the answer past the point where the system has typically moved on.

BrightEdge data on the ideal extraction window puts the target between 134 and 167 words for the opening passage. Our data aligns with this. The pages that performed best opened with a declarative answer, followed immediately by the most important supporting evidence.

Heading structure also mattered more than expected. H2 headings framed as questions (“What causes X?” rather than “The causes of X”) doubled citation probability in our sample compared to statement-style headings. Perplexity was especially responsive to question headings. The likely reason: question headings signal to retrieval systems that the following content is structured as an answer to a specific user intent.

For more on structuring pages to maximize extraction, see how to structure pages so AI can extract answers.

Finding 2: Named authors with credentials matter more than domain authority

This finding will be uncomfortable for teams that rely on brand reputation without investing in author identity.

We pulled data from the Ahrefs study of 863,000 keywords tracking AI Overview citation patterns. After Google’s Gemini 3 upgrade, 42% of previously cited domains were replaced in the citation pool. The domains that held or grew their citation share had one trait in common: named, credentialed authors on their content pages.

The citation rate breakdown by author configuration was stark:

  • Pages with no named author: cited 18% of the time
  • Pages with a named author and bio: cited 54% of the time
  • Pages with a named author, stated credentials, and links to external profiles: cited 71% of the time

For YMYL topics, meaning medical, legal, financial, and health content, the credential requirement approaches binary. Pages without verifiable author expertise in those categories were cited at rates close to zero regardless of domain strength.

This reflects how AI systems assess E-E-A-T signals. A strong domain helps, but a named human with verifiable credentials provides the retrieval system with direct evidence of expertise. The author is the entity. The external profile links, whether LinkedIn, a professional association page, or published work elsewhere, give the system something to triangulate against.

The practical implication: anonymous content, even well-written content from high-authority domains, is progressively losing its citation share to identified authors. See E-E-A-T in the AI era for more on building author authority signals.

CITATION RATE BY CONTENT TYPE · AI PLATFORMS 2026Original research / data studiesComprehensive how-to guidesComparison / vs pagesFAQ-structured postsNews and updatesGeneric blog postsProduct / service pages (no content)78%71%64%61%44%23%9%
Original research and comprehensive guides cited 3x more often than generic blog posts across all three AI platforms.

Finding 3: Schema markup is a citation multiplier, not a ranking factor

Schema markup does not directly cause citation. A page without schema can be cited, and a page with full schema can be ignored. But in our data, schema consistently amplified other positive signals, acting as a multiplier on pages that already had good content structure and author authority.

The citation rate differences by schema type:

  • Pages with FAQPage schema: 2.3x more likely to have FAQ content extracted verbatim
  • Pages with Article schema plus Person schema: 1.8x higher citation rate versus no schema
  • Pages with HowTo schema on step-by-step content: 1.6x higher citation rate for process queries
  • Pages with no schema markup at all: lowest citation rate in every query category we tested

The most useful frame for thinking about schema is signal reinforcement. When the page content says something and the schema confirms it, the retrieval system has two independent signals pointing to the same answer. That redundancy increases confidence in extraction.

FAQPage schema had the largest measured effect because it maps directly to how AI systems prefer to present information. A well-marked FAQ gives the system pre-packaged answer units it can extract and display with minimal transformation. That is a structural gift to the retrieval system.

The combination that performed best across our query set was FAQPage schema plus Article schema plus Person schema. The trio signals: this is a credible document (Article), by a real person (Person), containing pre-structured answers (FAQPage). See schema markup for SEO for implementation guidance.

Finding 4: Platform citation behavior differs significantly

This is the most underappreciated finding in our study, and the one most likely to change how teams prioritize their optimization work.

ChatGPT favors comprehensive, long-form sources above almost every other signal. It cites Wikipedia, Reddit, and official documentation at far higher rates than any other platform. In our sample, ChatGPT rarely cited pages under 800 words, and its strongest preference was for content that covers multiple facets of a topic in a single document. A 2,000-word guide that addresses a topic from three angles consistently outperformed two well-written 800-word posts on adjacent subtopics.

Perplexity is the most citation-democratic of the three platforms. It cites smaller, specialist sites more frequently than ChatGPT or Google, and it treats recency as a strong positive signal. Every factual claim tends to get a citation, resulting in 6 to 8 citations per response on complex queries. The implication for newer or smaller sites: Perplexity is the platform where strong content can break through without requiring high domain authority or extensive backlink profiles. Recency also matters more here than on either of the other two platforms.

Google AI Overviews applies the highest E-E-A-T threshold of the three platforms. YouTube accounts for 23% of citations in our Google sample, which reflects both Google’s integration of its own properties and the growing importance of video as a trust signal. A well-cited YouTube explainer on a topic often earns the YouTube page a spot even when the channel is mid-sized. Google AI Overviews also shows the most influence from traditional SEO signals, though this is shifting. In our sample, 38% of citations came from pages ranking in positions 1 to 10, down from 76% in prior studies. The gap between “ranks well” and “gets cited” is widening.

For a deeper breakdown of platform-specific optimization, see how ChatGPT, Perplexity, and Google AI Mode cite sources differently.

HOW THE THREE PLATFORMS CITE DIFFERENTLY · 2026ChatGPTPerplexityGoogle AI OverviewsFAVORSCITATIONS / RESPONSETOP SOURCE TYPERECENCY WEIGHTSCHEMA IMPACTE-E-A-T THRESHOLDLong-form contentSpecialist sitesHigh-authority domains3-56-83-6Wikipedia / RedditSpecialist blogsYouTube / WikipediaLowHighMediumMediumLowHighMediumLowHigh
Each platform has distinct citation preferences. Optimizing for all three requires different content approaches.

Finding 5: Topic cluster architecture multiplies citation frequency

Individual page quality is necessary but not sufficient. In our analysis, sites with comprehensive topic cluster architecture were cited 3.4x more often across a query set than sites with isolated posts covering the same topics at comparable quality.

This is a domain-level finding, not a page-level finding. AI retrieval systems appear to evaluate topical authority at the site level as well as the page level. A supporting post within a cluster, one that links to a pillar page and receives links from other cluster posts, was cited at meaningfully higher rates than an equally well-written standalone post on the same subject on a different domain.

The internal linking pattern matters independently. In our sample, higher internal linking density within a topic cluster correlated with higher citation frequency across the cluster’s pages, not just the pillar. The retrieval system appears to reward the signal that a site has invested deeply in a topic area.

The practical implication is that the question “is this page good enough to get cited?” is less useful than “does this page exist within a topic cluster that signals domain expertise?” A strong cluster lifts all of its members. An isolated post, even an excellent one, competes at a structural disadvantage.

For guidance on building topic clusters designed for AI citation, see how to build a topic cluster.

Finding 6: Robots.txt blocking is more common than expected

One finding that surprised us: in our sample, 12% of high-quality pages that should have been citation candidates were being blocked from one or more AI crawlers through their robots.txt configuration.

This is a silent citation killer. A page cannot be cited by a platform whose crawler cannot access it. The pages in question were often high-authority, well-structured, and answer-forward. They simply were not reaching the AI index.

The specific crawlers to check in your robots.txt file:

  • GPTBot (OpenAI / ChatGPT)
  • ClaudeBot (Anthropic)
  • PerplexityBot (Perplexity)
  • Google-Extended (Google AI training and features)

Blocking any of these removes that platform as a citation source entirely. The blocks often appear as a legacy consequence of broad User-agent: * rules added during earlier waves of AI crawler concern, or from managed security configurations that were never updated.

Checking robots.txt should be the first step in any AI citation audit. It takes five minutes and eliminates an otherwise invisible bottleneck.

The citation-ready content formula

The six findings combine into a single framework for estimating citation probability:

Citation probability = (Answer structure x Authority signals x Schema completeness x Topic cluster depth) divided by Crawl accessibility

Each term in the formula is multiplicative on the others. A page with a perfect answer structure and zero author authority will underperform. A fully credentialed author on a page with buried answers will underperform. The formula rewards completeness.

The practical checklist that follows from this:

  1. Lead every page with a 150-word direct answer. Put the answer in the first paragraph, not after context-setting or narrative preamble. This is the single highest-leverage change for most sites.

  2. Add a named author with credentials and external profile links. The author is the entity. External links to LinkedIn, professional profiles, or published work give retrieval systems something to verify against. Anonymous authorship is a structural disadvantage that domain authority does not compensate for.

  3. Add FAQPage and Article schema to every post. FAQPage schema is the highest-impact single addition for most informational pages. Combined with Article and Person schema, it signals a complete, credible, structured document.

  4. Build out the topic cluster around your pillar subjects. Isolated pages compete at a disadvantage regardless of their individual quality. The cluster is the citation surface. Invest in supporting posts, internal links, and a clear topical architecture.

  5. Check robots.txt for AI crawler blocks. GPTBot, ClaudeBot, PerplexityBot, and Google-Extended all need access. Any block is a complete removal from that platform’s citation pool.


Frequently asked questions

How do AI platforms decide which sources to cite?

Each platform uses a different retrieval process, but the common thread is relevance plus credibility. The system identifies pages that directly answer the query, then filters for authority signals including domain trust, author credentials, and schema structure. The page that provides the clearest answer with the most verifiable authority signals tends to win. Format also matters: pages with pre-structured answer units, FAQ sections, and direct opening answers are easier for retrieval systems to extract from.

Does domain authority affect AI citations?

It affects citation probability, but less linearly than it once did. Our data, and the Ahrefs 863k keyword study, both show that high domain authority is no longer a reliable predictor of citation retention after major AI model upgrades. Domains with high authority but anonymous content or buried answers lost citation share to lower-authority domains with strong author credentials and answer-first structure. Authority helps but does not override structural signals.

Does publishing frequency affect citation probability?

For Perplexity, yes. Perplexity applies a meaningful recency weight and fresh content from regular publishers has an advantage on timely topics. For ChatGPT and Google AI Overviews, publishing frequency is a much weaker signal. What matters more for those platforms is depth and authority, not recency. A three-year-old comprehensive guide with strong author credentials will typically outperform a recent shallow post.

Why does Perplexity cite smaller sites more than ChatGPT?

Perplexity’s retrieval architecture is more web-first and less pre-trained-knowledge-first than ChatGPT’s. It retrieves from live web results at query time and applies less domain-authority weighting than Google. A specialist site with strong topical depth and recent content can break through on Perplexity in ways that would require more domain history to achieve on the other two platforms. This makes Perplexity the most accessible citation target for newer publishers.

How do I track whether my site is being cited by AI systems?

There is no universal citation tracking tool yet, but several approaches work. Perplexity shows source links for every response, making manual checking feasible. Tools including Profound and emerging AI visibility platforms are building citation tracking dashboards. For ChatGPT, browsing mode responses include citations. Google AI Overviews citations appear in Search Console data under AI features impressions. Manual query monitoring, where you run your target queries regularly and record citation appearances, remains the most reliable method for most teams.

What is the fastest way to improve AI citation rates?

Check robots.txt first. If GPTBot, PerplexityBot, or Google-Extended is blocked, fix that before anything else. After that, audit your top pages for answer structure. Pages that open with direct answers perform demonstrably better than pages that delay the answer. The combination of robots.txt fix plus answer-first restructuring on your five highest-priority pages will produce measurable citation improvement faster than any other intervention.



References

Ahrefs study of 863,000 keywords: AI Overview citation patterns BrightEdge AI Overviews growth research Coalition Technologies: step-by-step AI SEO guide for growing AI citations and visibility in 2026 Profound: AI platform citation patterns Google Search Central: AI features and your website Search Engine Journal: entity authority in AI search Digital Applied: SEO after AI Overviews, complete strategy guide 2026