Where Does ChatGPT Get Its Information? We Measured
ChatGPT answers from training data, a search index, or a live fetch. We checked 23 AI Overviews: only 40% of page-one domains were cited in the answer.

The short version
ChatGPT answers from three different places — a training corpus frozen months ago, its own search index, and a live fetch made because you asked. Only the last two can include a page published this year, and none of them tell you whether your page ended up in the answer. You cannot inspect ChatGPT's retrieval, but you can inspect Google's AI Overview, which shows its sources. We checked 23 of them: 147 citations, and only 40% of the top-10 organic domains on those searches were cited at all. On 9 of the 23, the #1 result was not in the answer above it. Meanwhile 22% of citations came from sites that ranked nowhere in the top 20. A 400,000-URL analysis reaches the same place from the other side: domain authority affects whether you are retrieved, not whether you are cited. Being read and being cited are separate events, and your server log only ever proves the first.
Key highlights
- 1 ChatGPT answers from three sources: training data, its search index, and a live fetch
- 2 You cannot inspect ChatGPT's retrieval; Google's AI Overview is the one that shows sources
- 3 Only 40% of top-10 organic domains were cited in the answer above them
- 4 On 9 of 23 searches, the #1 result was not cited at all
- 5 22% of citations came from sites ranked nowhere in the top 20
- 6 YouTube and Reddit were the two most-cited domains in the sample
- 7 106 of 120 cited domains appeared exactly once — there is no fixed source list
- 8 SE Ranking: 88% of AI Mode advertisers were not cited in the answer above their own ad
- 9 A 400k-URL analysis puts domain authority on retrieval, not citation
- 10 Eight of eight AI systems said a retrieved page is not a cited one
Somebody asks ChatGPT about your category and your competitor gets named. You want to know where that came from — whether the model simply knew it, looked it up, or went and read a page thirty seconds ago.
Those are three different mechanisms with three different fixes, and the honest answer is that you can only see one of them from your side. This piece covers all three, then what you can actually measure.
The Three Places an Answer Can Come From
Every answer you see is assembled from at least one of these, and they behave nothing alike.
| Source | What it is | How old | Can you influence it now |
|---|---|---|---|
| Training corpus | text the model absorbed during training | months to years | no — the next training run, at best |
| Search index | the engine's own crawled copy of the web | days to weeks | yes, by being crawlable and worth indexing |
| Live fetch | your page, retrieved because somebody asked | seconds | yes, and this is the one nearest a citation |
Three consequences follow, and each one changes what you should do:
- A page published this year cannot be in the training corpus of a model that shipped last year. If a competitor is named from training data, no amount of publishing fixes it this quarter. That is the slowest lever there is.
- The search index is where most of the winnable ground sits. It refreshes continually, it is what the engine consults when it needs something current, and it is reachable by ordinary technical means.
- A live fetch is the strongest signal you can observe — somebody asked a question and the engine went and got your page to help answer it. It is also the agent most often confused with the training crawler, because the names look alike and the jobs are opposite.
The trap is treating these as one thing called "AI". They have different latencies, different levers, and only one of them shows up in your server log as it happens.
Can You Actually See Any of It?
Partly, and the boundary is worth being precise about.
- What you can see: which agents fetched your pages, when, and which URLs. That is your access log, and it is real evidence of being read.
- What you cannot see: whether any answer named you. ChatGPT does not publish, per answer, which pages it consulted in a form you can audit from outside.
Which leaves one useful surface. Google's AI Overview lists its sources. It is a different engine from ChatGPT with its own retrieval, so it is a proxy and not a substitute — but it is the only place where the relationship between ranking and being cited is visible from the outside at all.
So we measured it.
We Checked 23 AI Overviews and Every Source They Cited
Twenty-three searches, all in the SEO, local-SEO and AI-search space. For each one we took the AI Overview's cited domains and the organic results on the same page, and asked two questions in opposite directions.
The sample and the method are the honest part, so both first:
- 23 searches, each returning an AI Overview with at least one linked source.
- 147 unique cited domains across them; a domain cited twice in one answer counts once.
- Comparison is domain-level against the same search's organic results, top 20.
- Searches are commercial and informational queries in one category, not a random sample of the web. Read these as a category reading, not a universal rate.
citations from the top 10 organic 90/147 61%
citations ranked 11-20 only 25/147 17%
citations nowhere in the top 20 32/147 22%
top-10 domains that WERE cited 90/223 40%
searches where the #1 result was cited 14/23 61% The first three lines are the direction everybody publishes. The last two are the direction that matters to you, and they say something different.
Ranking Does Not Get You Cited
Across those 23 searches there were 223 domains sitting in the top 10 organic results. 90 of them were cited in the AI Overview. The other 133 were not.
Put plainly: six out of ten pages ranking on page one were left out of the answer printed above them.
And it goes further up the page than that. On 9 of the 23 searches, the #1 organic result was not cited at all. The single best-ranked page for the query, ignored by the summary sitting directly above it.
That reframes what a good ranking buys you:
- Ranking is a qualifier, not a ticket. 61% of citations came from the top 10, so being there matters enormously for your odds.
- But it is not sufficient. Getting into the top 10 gave a domain roughly a 40% chance of being cited, in this sample.
- The gap is a selection step. Something chooses among the pages it already has, and rank order is not that something.
If you have been reporting a page-one ranking as AI-search progress, this is the number that says how much progress. It is real, and it is not the finish line.
There are Only About Six Slots
Before asking who gets picked, it is worth knowing how many can be.
citations per AI Overview min 4 · median 6 · max 10 · mean 6.4
top-10 organic results 10, every time Six sources, drawn from a field that includes ten page-one results plus everything else on the open web. The arithmetic is unforgiving on its own — even if the answer layer only ever picked from page one, four of the ten would still be left out.
It does not only pick from page one, which is why the real figure is 40% and not 60%.
One thing the sample did not show, and it matters: every one of the 23 answers cited at least one of its own top-10 results. Page one is never ignored wholesale. On 4 of the 23, though, most of the citations came from outside it — so the balance swings, query by query, in a way rank order does not predict.
Who Actually Gets Cited
Across the 23 answers there were 120 distinct domains cited. Two things about that set are worth more than the rate:
| Domain | Cited on | Share of the 23 |
|---|---|---|
youtube.com | 9 searches | 39% |
reddit.com | 7 searches | 30% |
semrush.com | 3 searches | 13% |
| everything else | 1-2 searches each | — |
The two most-cited domains in a set of business-software and SEO searches are a video platform and a forum. Neither is a blog, and neither is competing on the traditional page-one contest in the way the other 118 are.
And the tail is almost entirely flat: 106 of the 120 domains were cited exactly once. There is no standing list of trusted sources being reused across queries. The selection is made per question, mostly from sites that appear in one answer and never again.
Two things follow for anyone planning content:
- Format is a variable, not a detail. If a video and a forum thread outrank 118 publishers for citation frequency in this category, the shape of the answer is doing work that the domain's authority is not.
- A single citation is a normal outcome, not a breakthrough. Being cited once is what 88% of cited domains experienced. Treat one appearance as a data point, not a trend.
And You Can be Cited Without Ranking at All
The inverse is just as true. 32 of the 147 citations — 22% — came from domains that appeared nowhere in the top 20 organic results for that search.
Better than one citation in five went to a site that, by traditional measurement, was not competing on that query at all.
Two readings, both worth holding:
- The optimistic one: the answer layer is a second door. A page can be selected on the strength of what it says, without first winning the ranking contest.
- The cautious one: your rank-tracking dashboard cannot see that door. A competitor gaining citations from position 40 produces no movement in any report you currently run.
Read is Not Cited, and Eight of Eight Agree
All of the above is about a surface that publishes its sources. Your server log is not that surface.
We put the question to eight different AI systems: does retrieving a page mean citing it? Eight of eight said no. They are describing their own behaviour, so it is self-report rather than measurement — but eight independent systems agreeing on the distinction, against their own commercial interest in looking useful, is worth something.
So a log entry supports exactly three claims:
- You were reachable. The request succeeded, nothing blocked it.
- You were read. Bytes were served to that agent, at that time.
- Something wanted this URL. Which URL was chosen is genuinely informative.
It supports none of these:
- that an answer named you
- that a person saw your name
- that anything followed from it
This is the honest limit of the whole exercise, and it is why a bot arriving is not the same as being mentioned. A log that looks busy can sit underneath a category where you are never named, and nothing in the file will tell you.
Why Our 61% is Not the 38% You May Have Read
An Ahrefs study circulated widely in r/SEO put the share of AI Overview citations coming from the top 10 organic results at roughly 38%. We measured 61% on our 23. Both can be right, and the difference is instructive.
- Sample composition. Ours is 23 commercial and informational queries inside one category. A broad sample spanning news, shopping, health and local intent will behave differently, because the answer layer draws on different source types per intent.
- Sample size. 23 searches is small. Treat our figure as a reading of this category at this date, not a competing universal rate.
- What both agree on. Whichever number you take, a large minority of citations comes from outside page one. That is the finding that survives both samples.
Where they genuinely disagree, prefer the larger sample. Where they agree, act on it.
Somebody Else Measured the Same Gap on Ad Spend
The strongest independent check on "ranking does not get you cited" comes from a different direction entirely: paying for placement.
SE Ranking analysed 50,032 commercial keywords in Google AI Mode. 14,733 of them returned a text ad. They then checked whether the advertiser's own domain appeared among the sources the AI cited for that same query. As the practitioner who wrote it up put it:
"So in 88% of cases, a brand bought the ad slot and was not among the sources the AI used to construct the answer sitting directly above that ad."
Their obvious objection was selection bias — maybe lower-authority sites buy ads and higher-authority sites get cited. They tested it, comparing advertising domains against non-advertising ones matched on domain trust, backlinks, referring domains and organic visibility. Advertisers showed no citation advantage.
Different surface, different method, much larger sample, same shape as our 40%: presence on the page does not buy presence in the answer.
What Appears to Get a Page Picked
Our run measures the outcome, not the mechanism. For the mechanism, the most useful analysis we have seen is one circulated in r/bigseo covering 400,000 URLs across 10,000 queries, and it is framed on exactly the distinction this piece is about — what it takes "to go from a URL retrieved (ChatGPT considers you to answer that question) to cited (your URL appears on the summary)."
After clustering 70+ content and domain features, its reported weights:
| Factor | Weight | What it affects |
|---|---|---|
| Content–answer fit | 55% | citation — how closely the page matches the answer style the model wants |
| On-page structure | 14% | citation — how easy the page is to parse and quote |
| Domain authority | 12% | retrieval, not citation |
| Query relevance | 12% | retrieval |
| Content consensus | 7% | citation — agreement with other sources |
That third row is the whole argument in one line. Authority gets you considered. It is not what gets you used. Our 40% is what that looks like from the outside.
The figures are theirs, not ours, and the underlying analysis is a vendor's — read it before betting on the exact percentages. The direction, though, matches what we measured independently.
A smaller, more human check points the same way. Someone who hand-read 20 ChatGPT citations looking for patterns reported:
"A lot of the cited pages weren't necessarily the most visually impressive or the most heavily optimized from a traditional SEO perspective. What they did well was provide clear, useful, and trustworthy information."
Their list of recurring signals — answer arrives fast with no preamble, clear heading structure, original data, mentions on sites other than your own — is not a ranking checklist. It is a description of a page that is easy to quote.
What the People Running These Tests are Actually Finding
Three things worth knowing before you plan around any of this, all from practitioners publishing their own runs.
- Single-run citation checks are close to noise. A tester who ran 12 buyer-intent questions five times each against Perplexity and Claude — 120 calls in a day — cited research finding only ~32–43% agreement between identical prompts run minutes apart. If you check a query once and see yourself, you have not learned much.
- Being in the category does not mean being cited in it. In that same test, 12 of 16 AI-visibility specialist vendors were never cited once across 120 calls, while established SEO platforms were cited at roughly 2.6× their rate. The author builds one of the tools that scored zero, and led with it.
- Referral traffic does move, sometimes. One team reported ChatGPT referrers going from 20 sessions to 120 over eight weeks after adding FAQ schema and an
llms.txt. One team, no control, self-reported — but it is the shape of evidence most people actually have, and worth reading as such.
None of these are ours, all of them are linked, and every one is a person publishing a method rather than a vendor publishing a claim.
What Actually Follows From This
Four things, in the order they are worth doing.
- Rank anyway. 61% of citations came from the top 10. Ranking is upstream of citation and remains the highest-leverage input, even though it guarantees nothing. Anybody selling you AI visibility as a replacement for search visibility has the arrow backwards.
- Stop reading rank as citation. Two different events, measured on two different surfaces. If you report one as the other you will be wrong roughly 60% of the time, in this sample.
- Check whether you are reachable at all before optimising anything else. A page the crawlers cannot read is not in the running for the 40%, and that is a technical problem with a technical fix.
- Measure the answer surface separately. The AI Overview shows its sources; check the queries you care about by hand, or track them. Your log tells you who read you. Neither one substitutes for the other.
Conclusion
ChatGPT answers from three places — training data you cannot change now, a search index you can, and a live fetch that happens because somebody asked. Only the last two are worth optimising for this quarter.
You cannot audit that retrieval from outside. You can audit Google's AI Overview, and in 23 of them, 60% of the page-one domains were not cited and the #1 result was skipped nine times. Meanwhile 22% of citations went to sites nowhere near the top 20.
Being read is not being cited. Your log proves the first, cannot see the second, and the distance between them is the whole problem.
AI Citation and Retrieval: Frequently Asked Questions
Where does ChatGPT get its information from?
Does ChatGPT cite the pages it reads?
If I rank 1 on Google will I be in the AI Overview?
Can a page be cited by AI without ranking on Google?
How can I tell if AI is citing my site?
Does my server log show whether I was mentioned in an answer?
Should I focus on SEO or on AI visibility?
Find the exact growth leak in your business — in 2 minutes.
Paste your URL. Our AI agent crawls your site, diagnoses what's broken, and ships a step-by-step fix plan. Free, no signup.
Run free audit

