Playbook

How to write content that AI engines cite

AI citations concentrate on a few dozen domains. Zumi data shows how few, and what that means for writing content AI engines cite. Read the guide.

Karthick Sreedaran7 min read

Reading brief

Decision
Whether to improve an existing page or create a new authoritative resource.
Evidence
Check extractability, source support, originality, entity consistency, update cadence, and the structural patterns associated with underperforming content.
Action
Rewrite one high-value uncited page with a direct answer, clear takeaways, and named sources, then retest it.

Getting cited by AI engines takes two jobs, and most advice covers only one. The first is writing a page an engine can quote. The second is getting that page, or a mention of the brand, onto the sites engines actually cite.

The second job is the larger one. In Zumi's measurement of 8,740 AI answers, 45 domains took half of all citations. The other 2,398 domains shared the rest.

This guide covers both jobs, in the order a team would tackle them.

Key takeaways

  • AI citations concentrate. In a Zumi run of 304 beauty-category questions over 11 days, 45 of 2,443 cited domains took half of all 40,304 citations.
  • A well-written page outside those 45 domains competes for the other half, spread thin across 2,398 domains. Coverage on the busiest domains is often worth more than another page on the brand's own site.
  • On the page itself, statistics, direct quotes, and cited sources produced the largest measured lift in the GEO study by researchers at Princeton and IIT Delhi. Keyword stuffing produced little to none.
  • Pages that rank in search but are not cited can usually be fixed in place: a takeaways block, headings that state a claim, the strongest claim moved up, and a dated source on every claim.
  • Two conditions sit outside the article: the brand described the same way on Wikipedia, Crunchbase, LinkedIn, and its own site, and a regular review of the pages most likely to be cited.

Where do AI engines' citations actually go?

Zumi asked 304 questions about beauty and personal care in India every day from 23 July to 2 August 2026. Nearly all of the 8,740 answers came from ChatGPT, Perplexity, and Google AI Overviews. Together they carried 40,304 citations spread over 2,443 domains.

Share of all citationsDomains it took
The top 10 domains28.9% of citations
Half of all citations45 domains
The other half2,398 domains

The ten most-cited domains were Instagram, YouTube, Nykaa, Reddit, NBC News, the American Academy of Dermatology, Amazon India, Allure, Smytten, and Healthline. None was a beauty brand's own website.

That matches Pew Research Center's finding for Google's AI summaries: "the most frequently cited sources in both Google AI summaries and standard search results are Wikipedia, YouTube and Reddit" (Pew Research Center, 2025).

The measurement covers one category in one country over 11 days. Other categories will differ in which domains lead. A steep concentration like this one is still the pattern to plan for.

Which content changes lift citation?

The "GEO: Generative Engine Optimization" study from researchers at Princeton and IIT Delhi (Aggarwal et al., KDD 2024) tested nine content changes across 10,000 queries.

GEO, short for generative engine optimization, is the work of getting content into AI-written answers. Three changes produced a consistent lift in how visible a page was in those answers:

  • Statistics and hard numbers: a 30 to 40% relative improvement.
  • Cited external sources: the largest single effect, up to 115.1% for pages ranked fifth in traditional results.
  • Direct quotations from a named expert or institution: the same 30 to 40% range.

Keyword stuffing offered little to no improvement. The study measured these lifts. It did not test why they happen.

What makes a passage easy for an engine to quote?

An AI engine writes its answer from passages, claims, and figures it pulls from the pages it reads. A good page that buries its point is harder to quote than a plain page that states it.

The answer comes first. The key claim belongs in the first paragraph, not after a build-up. A news story puts the key fact in its first line, which is the structure to copy; an academic paper saves the conclusion for the end.

Claims are single sentences. "The evidence suggests that..." followed by three paragraphs of context is hard to quote. One sentence naming the subject, the action, and the measured result is the form an engine can lift and credit.

Headings state what the section answers. "Statistics lift visibility by 30 to 40%" tells the reader and the engine the point. "Statistics and citations" is only a label.

A takeaways block sits near the top. A short, labeled list of the page's main claims gives an engine complete sentences it can quote without reading the rest.

Can existing content be fixed instead of replaced?

Not every citation gap needs a new page. A page on the right topic, written for reading straight through rather than for quoting, can usually be fixed in place, faster than a rewrite.

The edits, in order of impact:

  1. A key takeaways block near the top, if one is missing.
  2. Headings rewritten from topics into claims or questions.
  3. The first specific claim, statistic, or cited source moved into the first paragraph.
  4. A named, dated source attached to every claim that lacks one.

A page that ranks but is not cited usually buries its best claim rather than missing it. The facts are already there, just too far down the page.

What does original mean here?

Original does not require a large survey. It means a claim that exists nowhere else in exactly that form: an observation from a stated method, or a number from a specific analysis.

"AI search is growing" is not original. The share of last quarter's sales calls in which the buyer said an AI engine was where they first heard of the brand, counted from call records, is. That kind of claim adds information the answer did not already have.

What has to be true off the page?

The concentration data sets the priority. A page on the brand's own site competes for the half of citations spread across 2,398 domains. Coverage on the 45 domains that take the other half appears in answers far more often, domain for domain.

The brand has to be described the same way everywhere. When the Wikipedia article, Crunchbase profile, LinkedIn page, and the brand's own site describe the company differently, engines have conflicting versions to reconcile. The audit and correction steps are in entity optimization for AI.

The brand has to appear where engines already look. Articles, reviews, and creator videos on the domains an engine cites for the category carry the brand into answers its own site cannot reach. Digital PR pays off when it targets those domains.

The priority pages have to stay current. A practical starting point is the 10 to 20 pages most likely to be cited, each reviewed at least every quarter. Each update should change something real, such as new data or a corrected claim; a new date on unchanged copy does not count.

What tends to hold a page back?

Long introductions push the key claim too far down the page. Hedged language ("it may be the case that...") gives an engine no firm claim to quote. Generic advice ("create high-quality content") adds nothing the engine could not find on a dozen other pages.

Specific claims are the ones that get quoted. The version of a claim that could only come from one dataset or one team's experience is the version worth writing.

When is rewriting not the answer?

A better page on the brand's own site will not change much when the answers for its category cite almost no brand sites at all. In the beauty data above, none of the top ten domains was a brand's own site.

For a category like that, the work is coverage on the domains that are cited. Checking which domains the answers cite shows which case applies.

For the research behind these signals, see how AI engines decide what brands to mention. For the SaaS version of this problem, owning a category and building clusters of related content, see GEO for SaaS.

Writing is one playbook among several:

Zumi is an AI Search Intelligence Platform that shows whether a rewrite actually earned more citations: mention rate, share of voice, average position, and citation share, tracked per engine before and after a page changes.

Book a demo to see the before-and-after for a real page.


Sources

  • Aggarwal, P. et al. "GEO: Generative Engine Optimization." KDD 2024 / arXiv:2311.09735
  • Pew Research Center. "Google users are less likely to click on links when an AI summary appears in the results" (July 22, 2025). pewresearch.org
  • Zumi, India beauty and personal care question run, 304 questions asked daily on ChatGPT, Perplexity, and Google AI Overviews, 23 July to 2 August 2026.