The largest published analysis of AI citations found that reviews and third-party social proof accounted for 57% of cited sources, while blogs and guides accounted for 5.4%. That finding is genuinely useful and it is also widely misreported, because the study measured branded prompts — questions about specific companies — and the distribution for informational queries is not the same thing.
Both halves matter. The direction of the finding is real: AI systems disproportionately cite sources you do not own. The precise numbers describe one query type, and using them as a general law will send you in the wrong direction.
This guide covers what the evidence supports, where it does not, and which formats survive a generated summary.
What the study actually found
The analysis covered 23,387 unique cited sources drawn from 240 branded prompts, across four industries and six intent types, on five platforms — ChatGPT, Perplexity, Gemini, AI Mode and AI Overviews.
| Format | Share of citations |
|---|---|
| Reviews and social proof — reviews, listicles, forums, case studies | 57% |
| Directories — Wikipedia, documentation, profiles, support pages | 17% |
| Product pages — ecommerce, pricing, landing pages | 12% |
| Thought leadership — blogs, guides, research | 5.4% |
| Brand foundation — About (1.92%), homepage (1.82%), FAQ (0.41%), legal and contact (0.31%) | ~4.5% |
| News and press | Low |
| Video | Lowest |
The headline conclusion holds up: when someone asks an assistant about a company, the assistant overwhelmingly reaches for what other people said about it rather than what the company said about itself.
Where it gets misread
The prompts were branded. They asked about specific companies. That is a query type where third-party validation is exactly what a careful answer should cite — and it tells you very little about what gets cited when someone asks "how do I set up DMARC".
So "blogs only get 5% of citations" is not a general finding. For informational queries, documentation and explanatory content carry far more of the load, for the straightforward reason that no review site answers a how-to question.
Two smaller cautions. Five platforms behave differently and the aggregate hides that variation. And a citation is not a visit — the study counted sources named in answers, not traffic.
The honest summary: for questions about you, third-party sources dominate. For questions about the thing you do, your own content is still in play. Most sites need both and are only investing in one.
What this changes about content strategy
Three shifts, in order of how uncomfortable they are.
1. Some of your best AI visibility is not on your site
If assistants answering questions about your company mostly cite reviews, directories, forums and case studies, then the work of being cited is partly off-site work — and it looks more like PR and reputation management than SEO.
That means: legitimate directory and marketplace profiles kept accurate. Genuine review presence wherever your category's buyers look. Being interviewed, quoted, and included in other people's comparisons. A Wikidata entry where you qualify.
None of it is manufacturable, and the attempt — fake reviews, paid placements presented as editorial — is precisely what these systems are being tuned to discount.
2. Definitional content is the exposed category
A page whose value is explaining what something is competes directly with a generated summary, and the summary is free, instant, and adequate.
That does not mean deleting explainers. It means not expecting them to carry traffic, and making sure each one earns its place through something the summary cannot reproduce: your settings, your numbers, your judgement about what to do next.
3. Structure decides whether you are quotable
Retrieval works on passages. A section that states its answer in the first two sentences can be lifted and cited; one that arrives at its point in paragraph five gets skipped in favour of a page that did not make the reader wait.
This is the cheapest change on the list and it improves the page for humans at the same time, which is the test of whether an SEO recommendation is real.
The formats that hold up
Six, ordered by how hard they are to summarise away.
Original data. Your own test results, your own survey, your own analysis of numbers nobody else has. A summary can compress your data; it cannot replace it, and it has to cite you to use it. This is the single strongest format in AI search and the one fewest sites produce.
Specific procedures with real settings. Exact menu paths, actual error codes, version numbers, the thing that goes wrong at step four. Generic instructions summarise perfectly. Specific ones send people to the source.
Comparisons with a stated decision framework. Not "top 10 tools" — the reasoning that lets a reader decide for their own situation. Listicles are the most summarisable content on the internet. A framework with trade-offs is not.
Genuinely answered questions. A question a real person asked, answered completely, in their words. This is also the format that survived the FAQ schema deprecation intact — the SERP feature is gone, the value to readers and to retrieval is not.
Tables and structured comparisons. Machine-readable, quotable as a unit, and hard to paraphrase without losing the point.
Documented mistakes. What you got wrong, what it cost, what you do now. Rare, cheap to produce if you are honest, and almost impossible to synthesise from other sources — nobody else has your failures.
The formats that do not
Thin definitional posts. "What is X" as a standalone 600-word page. This is what generated answers replaced.
Undifferentiated listicles. Ten tools with a paragraph each and no position. Assembled from other pages, and therefore reassemblable without you.
Content whose only virtue is length. Two thousand words of restatement is two thousand words a summary discards.
Anything unverifiable. Unattributed statistics, undated claims, figures that trace back to another blog citing another blog. These are exactly what fact-checking passes strip, and the site publishing them loses standing when the trail collapses. The discipline is in fact-checking AI content before you publish.
Pure opinion with no evidence. Positions are valuable when they come with reasoning and specifics. Without either, there is nothing to cite.
The test worth applying
Before publishing, one question: could a competent assistant produce this page's core value from other sources?
If yes, the page is a candidate for being summarised out of relevance, and it needs something added — your data, your settings, your judgement, your named mistake.
If no, it is worth publishing regardless of what happens to search, because the value does not depend on the distribution channel.
That test also happens to be the definition of content worth writing, which is the reassuring part of this entire subject: the response to AI search is not a new technique. It is the standard that always applied, now enforced.
Frequently asked questions
What content gets cited by AI search?
An analysis of 23,387 cited sources from branded prompts found reviews and social proof at 57%, directories at 17%, product pages at 12% and blogs at 5.4%. The prompts asked about specific companies, so the distribution reflects brand queries — for how-to questions, explanatory content carries far more of the load.
Do blogs still matter for AI search?
Yes, though less for questions about your company than for questions about your subject. The 5.4% figure comes from branded prompts, where reviews and directories are what a careful answer should cite. For informational queries, documentation and guides remain a primary source.
Why do AI systems cite reviews more than company websites?
Because external validation is more credible than self-description for questions about a company. When someone asks whether a product is good, what other people said is more relevant evidence than what the seller said, and the systems are built to reflect that.
What content does AI search make obsolete?
Thin definitional pages, undifferentiated listicles, and long restatements of information available everywhere. A generated answer does exactly that job, free and instantly. Content built on original data, specific procedures or genuine judgement is not replaceable the same way.
Does FAQ content still work after the schema deprecation?
The content does; the SERP feature does not. FAQ rich results stopped appearing in Google on 7 May 2026, so the markup no longer produces a visual result. Well-answered questions remain valuable to readers and are still extracted by retrieval systems whether or not markup is present.
How do I get my content cited by AI?
Answer the question in the first two sentences of each section, keep sections self-contained so a passage survives being lifted out of context, and publish specifics — data, settings, figures, dates — that cannot be reproduced from other sources. Structure decides whether you are quotable; substance decides whether you are worth quoting.
Are listicles dead for SEO?
Undifferentiated ones are exposed, because they are assembled from other pages and can be reassembled without you. A comparison built on a stated decision framework, with trade-offs and a position, is not the same format and does not summarise away.
Should I stop publishing explainer content?
No, but stop expecting it to carry traffic on its own. An explainer earns its place when it adds something the summary cannot reproduce — your configuration, your numbers, the failure mode nobody else documents — and when it links onward to the content that does.
What to do next
Take your five highest-traffic pages and apply the test: could an assistant produce their core value from other sources? The ones where the answer is yes are your exposure, and they need something of yours added rather than more words.
Then check whether each section answers its own question in the first two sentences. That single structural pass is the cheapest AI-search improvement available, and it makes the pages better to read.
Related guides
- SEO when AI answers the question first — the wider picture this sits inside
- Entity SEO: why named things beat keywords now — being identifiable enough to cite
- How to track AI search traffic in GA4 — measuring what arrives
- llms.txt and AI crawlers: block them or not? — being reachable in the first place
- Fact-checking AI content before you publish — why unverifiable claims cost standing
- AI content production: a workflow that does not produce slop — producing this at volume
Free: The 60-Minute Email Authentication Fix
A no-fluff checklist to set up SPF, DKIM & DMARC correctly and pass Gmail & Yahoo's sender requirements.

Muhammad Basim has worked in digital marketing since 2013, focused on email deliverability and AI-assisted content production. He is the author of The Email Deliverability Playbook and The Email Copywriting Playbook.
Related Articles

Email Deliverability: Why Authenticated Emails Still Land in Spam
Email deliverability is whether your message reaches the inbox rather than the spam folder, and it is decided by three factors: technical infrastructure, list quality, and sending behaviour. Authentication is one item inside the first factor. Getting it right is necessary, and it is nowhere near sufficient. That gap explains the most common complaint in […]

Linkable Assets: What Actually Earns Links
A link is a citation, and a citation requires that someone writing about your subject needed you to make their point. That is the whole mechanism, and it is the reason most "linkable content" earns nothing. The test, before you build anything: Could a writer covering this topic finish their sentence without referencing you? If […]

Link Outreach Emails That Get Replies
Outreach fails for exactly two reasons, and only one of them gets written about. The first is that the email was not worth replying to. That is the reason every guide addresses, and the advice — personalise, be brief, offer value — is correct and insufficient. The second is that the email never arrived. Nobody […]

