TL;DR: Every AI search engine picks its sources by different rules. ChatGPT cites pages ranking beyond position 20 on Google almost 90% of the time (Semrush, 2025). Google AI Overviews halved its reliance on the organic top 10 in a single year, from 76.1% to 37.9% of citations (Ahrefs, 2025–2026). Perplexity pulls hardest from Reddit and from fresh content. Understand the three filters and you know exactly where your entry into the AI answer is, especially in small-language markets, where we ran our own test.
How does an AI engine find its sources in the first place?
An AI search engine doesn't "know" the answer, it assembles one. Your question gets split into several narrower queries (query fan-out), pages are retrieved for each, and sentences are extracted from those pages to build the answer. The citation goes to the pages whose sentences are easiest to lift and attribute.
That is where AI search departs from classic SEO. Google ranks pages; an AI engine picks sentences. Ahrefs' overlap analysis found that only 12% of URLs cited by AI assistants rank in Google's top 10 for the same prompt (the average for ChatGPT, Gemini and Copilot; Perplexity is the outlier at ~33%). Position helps. It doesn't decide.
And "AI search" isn't one system. The three biggest engines select sources in three different ways, which is why GEO (Generative Engine Optimization) isn't one recipe.

How does Google AI Overviews choose citations?
AI Overviews draws on Google's index, but less and less from page one. In July 2025, Ahrefs measured that 76.1% of AI Overviews citations came from the organic top 10. In the repeat analysis in March 2026, that share had fallen to 37.9%. Halved in eight months.
In practice: ranking no longer guarantees a citation, and missing the top 10 no longer excludes you. Google builds the AI answer as a consensus of multiple sources, so pages sitting at position 11, 30 or 50 increasingly earn citations, if they carry a clearly structured, extractable answer.
The single most-cited source? Reddit, at 21% of top-source share, per Profound's analysis of 680 million citations (August 2024 – June 2025). In AI answers, Google trusts real people's experiences over corporate blogs.
How does ChatGPT choose citations?
ChatGPT search is the furthest removed from classic rankings. Semrush's study of AI search's impact on SEO traffic puts it in one sentence worth quoting verbatim:
"When ChatGPT search cites webpages, the pages it cites rank in traditional organic search positions 21+ for related queries almost 90% of the time."
, Semrush, We Studied the Impact of AI Search on SEO Traffic
In other words: ChatGPT systematically cites pages almost nobody sees on Google. It doesn't mirror the SERP, it hunts for text it can extract an answer from, wherever that text ranks.
Its most-cited source is Wikipedia, at 47.9% of top-source share (Profound, 2025). And within a page, ChatGPT reads the top. Kevin Indig (Growth Advisor) analyzed 1.2 million ChatGPT answers and 18,012 verified citations: 44.2% of citations come from the first 30% of the content, with a sharp drop toward the footer, a shape Indig calls the "ski ramp". Same study: sentences with clear definitions ("X is...") get cited nearly twice as often as vague framing.

How does Perplexity choose citations?
Perplexity is the closest of the three to an actual search engine: it runs a live search for every query, ranks sources by relevance and authority, and footnotes every claim. It rewards two things above all.
First, communities. Reddit is its top source at 46.7% of most-cited-source share (Profound, 2025), more than double its share on Google AI Overviews.
Second, freshness. Seer Interactive measured a strong recency bias across all AI systems, roughly 65% of AI crawler visits go to content published within the last year, and Perplexity's is the strongest: about half of its citations go to content from the current year. A page that never gets refreshed quietly disappears from Perplexity.
What the data shows: three engines, three filters
| Google AI Overviews | ChatGPT search | Perplexity | |
|---|---|---|---|
| Where sources come from | Google's index, multiple parallel queries | own retrieval + query fan-out | real-time search, footnote per claim |
| Link to Google's top 10 | weakening: 76.1% → 37.9% of citations in one year (Ahrefs, 2025–2026) | weakest: ~90% of citations from positions 21+ (Semrush, 2025) | moderate; freshness beats position (Seer, 2025) |
| Most-cited source | Reddit, 21% (Profound, 2025) | Wikipedia, 47.9% (Profound, 2025) | Reddit, 46.7% (Profound, 2025) |
| Rewards most | multi-source consensus, structure | definitions, entities, first 30% of content (Indig, 2026) | freshness, community experience, direct answers |
One-sentence read: three systems, three filters, and none of them maps onto Google's classic page one anymore.
What actually increases your odds of being cited?
The firmest answer comes from Princeton's "GEO: Generative Engine Optimization" study (Aggarwal et al., KDD 2024), the first controlled test of tactics on generative engines:
- Expert quotations, the strongest single tactic: +41% position-adjusted visibility.
- Statistics with sources, +31%.
- Citing sources, combined with the above, up to 40% higher visibility, and up to +115% for a page sitting at position five.
- Keyword stuffing, the only tested tactic that performed worse than doing nothing.
Layer Indig's structural finding on top: the answer must live in the first 30% of the text, in sentences that stand on their own.
Schema markup? Evidence is split. AirOps measured a +13% citation lift for rich, multi-type schema, while Ahrefs' test on 1,885 pages found no clear effect. Worth having for Google, but schema alone won't buy you the citation.
"AI doesn't cite the best page. It cites the most citable one. Those are two different disciplines, and only the second one can be measured."
, Marko Hrnjak, Head of Marketing Technology, Risely Digital
Our experiment: what does AI cite when you ask in Croatian?
Every study above tests English prompts. So we tested our own market, Risely test, July 2026: two probes through Perplexity, in Croatian.
Probe 1: "Which agencies in Croatia offer GEO optimization?" The answer cited nine sources. Eight were the agencies' own sales pages, the pages of the very agencies named in the answer, including one Slovenian domain targeting the Croatian market. The ninth: a German website with no connection to the topic, pulled in because its URL contains "geo". Not one independent source. Not one case study. Not one data point.
Probe 2: "How do ChatGPT and Perplexity pick sources for Croatian questions?" The answer leaned on lifestyle and tech portals, one IT company's blog, a telecom's promotional page and two YouTube videos. Expert sources: zero.
The conclusion we didn't read anywhere, we measured it: in a small-language market, the source pool is so thin that AI cites sales brochures, because nothing better exists. In English you compete with Wikipedia and Reddit for the citation. In Croatian, your competition is empty space. Bad news for the quality of today's Croatian-language AI answers; very good news for whoever publishes verifiable, structured content first. If your brand operates in any small-language market, the same asymmetry almost certainly applies.
First three steps into AI answers
- Measure your baseline. Before writing anything, check who AI cites in your category today, the free GEO Audit runs it for your URL, including Croatian-language prompts no global tool tests.
- Structure content for extraction. Answer in the first 30% of the text, numbers with named sources, expert quotes, question-shaped headings. The full list of criteria lives in our GEO checklist.
- Treat it as a discipline, not a hack. GEO optimization is a loop: measure who's in the answer, fix what blocks the citation, measure again. Same as SEO, with a different yardstick: share of AI answers instead of positions.
FAQ
Does a page have to rank in Google's top 10 to be cited by AI?
No. ChatGPT cites pages from positions 21+ almost 90% of the time (Semrush, 2025), and the top-10 share of AI Overviews citations has fallen to 37.9% (Ahrefs, 2026). Position helps, but what decides is whether a clear, self-contained answer can be extracted from the text.
Why do AI engines cite Reddit and Wikipedia so much?
Profound's 30-million-citation analysis shows Wikipedia is ChatGPT's top source (47.9% of top-source share), while Reddit leads for Perplexity (46.7%) and AI Overviews (21%). Both offer exactly what models look for: structured definitions and real user experience.
How much does content freshness matter for AI citations?
A lot. Around 65% of AI crawler visits go to content published within the last year (Seer Interactive, 2025), and Perplexity is the most sensitive, about half of its citations go to current-year content. Regularly refreshing key pages is a requirement, not a bonus.
Does schema markup help you get cited by AI?
The evidence is split: AirOps measured +13% with rich schema; Ahrefs' 1,885-page test found no clear effect. Schema remains good practice for Google, but it doesn't earn the citation by itself. Extractable content does.
How do I check whether AI cites my site?
Manually: ask the AI engines the questions your buyers actually ask, in your language and in English, and record who appears in the answer, you, a competitor, or nobody. Systematically: with a tool that probes those prompts automatically and repeats the measurement, so you get a trend instead of an impression.
Keep reading: What is GEO and how it differs from SEO · Schema markup for AI search
