Where Does ChatGPT Get Its Information? We Measured

ChatGPT answers from training data, a search index, or a live fetch. We checked 23 AI Overviews: only 40% of page-one domains were cited in the answer.

Five source websites sized by how often they were cited across 23 AI Overviews — youtube.com largest, then reddit.com, semrush.com, quora.com and a node for 117 further domains — with curved arrows feeding into a single ChatGPT answer hub

The short version

ChatGPT answers from three different places — a training corpus frozen months ago, its own search index, and a live fetch made because you asked. Only the last two can include a page published this year, and none of them tell you whether your page ended up in the answer. You cannot inspect ChatGPT's retrieval, but you can inspect Google's AI Overview, which shows its sources. We checked 23 of them: 147 citations, and only 40% of the top-10 organic domains on those searches were cited at all. On 9 of the 23, the #1 result was not in the answer above it. Meanwhile 22% of citations came from sites that ranked nowhere in the top 20. A 400,000-URL analysis reaches the same place from the other side: domain authority affects whether you are retrieved, not whether you are cited. Being read and being cited are separate events, and your server log only ever proves the first.

Somebody asks ChatGPT about your category and your competitor gets named. You want to know where that came from — whether the model simply knew it, looked it up, or went and read a page thirty seconds ago.

Those are three different mechanisms with three different fixes, and the honest answer is that you can only see one of them from your side. This piece covers all three, then what you can actually measure.

The Three Places an Answer Can Come From

Every answer you see is assembled from at least one of these, and they behave nothing alike.

SourceWhat it isHow oldCan you influence it now
Training corpustext the model absorbed during trainingmonths to yearsno — the next training run, at best
Search indexthe engine's own crawled copy of the webdays to weeksyes, by being crawlable and worth indexing
Live fetchyour page, retrieved because somebody askedsecondsyes, and this is the one nearest a citation

Three consequences follow, and each one changes what you should do:

  1. A page published this year cannot be in the training corpus of a model that shipped last year. If a competitor is named from training data, no amount of publishing fixes it this quarter. That is the slowest lever there is.
  2. The search index is where most of the winnable ground sits. It refreshes continually, it is what the engine consults when it needs something current, and it is reachable by ordinary technical means.
  3. A live fetch is the strongest signal you can observe — somebody asked a question and the engine went and got your page to help answer it. It is also the agent most often confused with the training crawler, because the names look alike and the jobs are opposite.

The trap is treating these as one thing called "AI". They have different latencies, different levers, and only one of them shows up in your server log as it happens.

Can You Actually See Any of It?

Partly, and the boundary is worth being precise about.

  • What you can see: which agents fetched your pages, when, and which URLs. That is your access log, and it is real evidence of being read.
  • What you cannot see: whether any answer named you. ChatGPT does not publish, per answer, which pages it consulted in a form you can audit from outside.

Which leaves one useful surface. Google's AI Overview lists its sources. It is a different engine from ChatGPT with its own retrieval, so it is a proxy and not a substitute — but it is the only place where the relationship between ranking and being cited is visible from the outside at all.

So we measured it.

We Checked 23 AI Overviews and Every Source They Cited

Twenty-three searches, all in the SEO, local-SEO and AI-search space. For each one we took the AI Overview's cited domains and the organic results on the same page, and asked two questions in opposite directions.

The sample and the method are the honest part, so both first:

  • 23 searches, each returning an AI Overview with at least one linked source.
  • 147 unique cited domains across them; a domain cited twice in one answer counts once.
  • Comparison is domain-level against the same search's organic results, top 20.
  • Searches are commercial and informational queries in one category, not a random sample of the web. Read these as a category reading, not a universal rate.
txt
citations from the top 10 organic     90/147   61%
citations ranked 11-20 only           25/147   17%
citations nowhere in the top 20       32/147   22%
top-10 domains that WERE cited        90/223   40%
searches where the #1 result was cited 14/23   61%

The first three lines are the direction everybody publishes. The last two are the direction that matters to you, and they say something different.

Ranking Does Not Get You Cited

Across those 23 searches there were 223 domains sitting in the top 10 organic results. 90 of them were cited in the AI Overview. The other 133 were not.

Put plainly: six out of ten pages ranking on page one were left out of the answer printed above them.

And it goes further up the page than that. On 9 of the 23 searches, the #1 organic result was not cited at all. The single best-ranked page for the query, ignored by the summary sitting directly above it.

That reframes what a good ranking buys you:

  • Ranking is a qualifier, not a ticket. 61% of citations came from the top 10, so being there matters enormously for your odds.
  • But it is not sufficient. Getting into the top 10 gave a domain roughly a 40% chance of being cited, in this sample.
  • The gap is a selection step. Something chooses among the pages it already has, and rank order is not that something.

If you have been reporting a page-one ranking as AI-search progress, this is the number that says how much progress. It is real, and it is not the finish line.

There are Only About Six Slots

Before asking who gets picked, it is worth knowing how many can be.

txt
citations per AI Overview     min 4 · median 6 · max 10 · mean 6.4
top-10 organic results         10, every time

Six sources, drawn from a field that includes ten page-one results plus everything else on the open web. The arithmetic is unforgiving on its own — even if the answer layer only ever picked from page one, four of the ten would still be left out.

It does not only pick from page one, which is why the real figure is 40% and not 60%.

One thing the sample did not show, and it matters: every one of the 23 answers cited at least one of its own top-10 results. Page one is never ignored wholesale. On 4 of the 23, though, most of the citations came from outside it — so the balance swings, query by query, in a way rank order does not predict.

Who Actually Gets Cited

Across the 23 answers there were 120 distinct domains cited. Two things about that set are worth more than the rate:

DomainCited onShare of the 23
youtube.com9 searches39%
reddit.com7 searches30%
semrush.com3 searches13%
everything else1-2 searches each—

The two most-cited domains in a set of business-software and SEO searches are a video platform and a forum. Neither is a blog, and neither is competing on the traditional page-one contest in the way the other 118 are.

And the tail is almost entirely flat: 106 of the 120 domains were cited exactly once. There is no standing list of trusted sources being reused across queries. The selection is made per question, mostly from sites that appear in one answer and never again.

Two things follow for anyone planning content:

  • Format is a variable, not a detail. If a video and a forum thread outrank 118 publishers for citation frequency in this category, the shape of the answer is doing work that the domain's authority is not.
  • A single citation is a normal outcome, not a breakthrough. Being cited once is what 88% of cited domains experienced. Treat one appearance as a data point, not a trend.

And You Can be Cited Without Ranking at All

The inverse is just as true. 32 of the 147 citations — 22% — came from domains that appeared nowhere in the top 20 organic results for that search.

Better than one citation in five went to a site that, by traditional measurement, was not competing on that query at all.

Two readings, both worth holding:

  • The optimistic one: the answer layer is a second door. A page can be selected on the strength of what it says, without first winning the ranking contest.
  • The cautious one: your rank-tracking dashboard cannot see that door. A competitor gaining citations from position 40 produces no movement in any report you currently run.

Read is Not Cited, and Eight of Eight Agree

All of the above is about a surface that publishes its sources. Your server log is not that surface.

We put the question to eight different AI systems: does retrieving a page mean citing it? Eight of eight said no. They are describing their own behaviour, so it is self-report rather than measurement — but eight independent systems agreeing on the distinction, against their own commercial interest in looking useful, is worth something.

So a log entry supports exactly three claims:

  1. You were reachable. The request succeeded, nothing blocked it.
  2. You were read. Bytes were served to that agent, at that time.
  3. Something wanted this URL. Which URL was chosen is genuinely informative.

It supports none of these:

  • that an answer named you
  • that a person saw your name
  • that anything followed from it

This is the honest limit of the whole exercise, and it is why a bot arriving is not the same as being mentioned. A log that looks busy can sit underneath a category where you are never named, and nothing in the file will tell you.

Free tool
See which AI engines are reading your site, and which ones are only claiming to.
Check AI visibility

Why Our 61% is Not the 38% You May Have Read

An Ahrefs study circulated widely in r/SEO put the share of AI Overview citations coming from the top 10 organic results at roughly 38%. We measured 61% on our 23. Both can be right, and the difference is instructive.

  • Sample composition. Ours is 23 commercial and informational queries inside one category. A broad sample spanning news, shopping, health and local intent will behave differently, because the answer layer draws on different source types per intent.
  • Sample size. 23 searches is small. Treat our figure as a reading of this category at this date, not a competing universal rate.
  • What both agree on. Whichever number you take, a large minority of citations comes from outside page one. That is the finding that survives both samples.

Where they genuinely disagree, prefer the larger sample. Where they agree, act on it.

Somebody Else Measured the Same Gap on Ad Spend

The strongest independent check on "ranking does not get you cited" comes from a different direction entirely: paying for placement.

SE Ranking analysed 50,032 commercial keywords in Google AI Mode. 14,733 of them returned a text ad. They then checked whether the advertiser's own domain appeared among the sources the AI cited for that same query. As the practitioner who wrote it up put it:

"So in 88% of cases, a brand bought the ad slot and was not among the sources the AI used to construct the answer sitting directly above that ad."

Their obvious objection was selection bias — maybe lower-authority sites buy ads and higher-authority sites get cited. They tested it, comparing advertising domains against non-advertising ones matched on domain trust, backlinks, referring domains and organic visibility. Advertisers showed no citation advantage.

Different surface, different method, much larger sample, same shape as our 40%: presence on the page does not buy presence in the answer.

What Appears to Get a Page Picked

Our run measures the outcome, not the mechanism. For the mechanism, the most useful analysis we have seen is one circulated in r/bigseo covering 400,000 URLs across 10,000 queries, and it is framed on exactly the distinction this piece is about — what it takes "to go from a URL retrieved (ChatGPT considers you to answer that question) to cited (your URL appears on the summary)."

After clustering 70+ content and domain features, its reported weights:

FactorWeightWhat it affects
Content–answer fit55%citation — how closely the page matches the answer style the model wants
On-page structure14%citation — how easy the page is to parse and quote
Domain authority12%retrieval, not citation
Query relevance12%retrieval
Content consensus7%citation — agreement with other sources

That third row is the whole argument in one line. Authority gets you considered. It is not what gets you used. Our 40% is what that looks like from the outside.

The figures are theirs, not ours, and the underlying analysis is a vendor's — read it before betting on the exact percentages. The direction, though, matches what we measured independently.

A smaller, more human check points the same way. Someone who hand-read 20 ChatGPT citations looking for patterns reported:

"A lot of the cited pages weren't necessarily the most visually impressive or the most heavily optimized from a traditional SEO perspective. What they did well was provide clear, useful, and trustworthy information."

Their list of recurring signals — answer arrives fast with no preamble, clear heading structure, original data, mentions on sites other than your own — is not a ranking checklist. It is a description of a page that is easy to quote.

What the People Running These Tests are Actually Finding

Three things worth knowing before you plan around any of this, all from practitioners publishing their own runs.

  • Single-run citation checks are close to noise. A tester who ran 12 buyer-intent questions five times each against Perplexity and Claude — 120 calls in a day — cited research finding only ~32–43% agreement between identical prompts run minutes apart. If you check a query once and see yourself, you have not learned much.
  • Being in the category does not mean being cited in it. In that same test, 12 of 16 AI-visibility specialist vendors were never cited once across 120 calls, while established SEO platforms were cited at roughly 2.6× their rate. The author builds one of the tools that scored zero, and led with it.
  • Referral traffic does move, sometimes. One team reported ChatGPT referrers going from 20 sessions to 120 over eight weeks after adding FAQ schema and an llms.txt. One team, no control, self-reported — but it is the shape of evidence most people actually have, and worth reading as such.

None of these are ours, all of them are linked, and every one is a person publishing a method rather than a vendor publishing a claim.

What Actually Follows From This

Four things, in the order they are worth doing.

  1. Rank anyway. 61% of citations came from the top 10. Ranking is upstream of citation and remains the highest-leverage input, even though it guarantees nothing. Anybody selling you AI visibility as a replacement for search visibility has the arrow backwards.
  2. Stop reading rank as citation. Two different events, measured on two different surfaces. If you report one as the other you will be wrong roughly 60% of the time, in this sample.
  3. Check whether you are reachable at all before optimising anything else. A page the crawlers cannot read is not in the running for the 40%, and that is a technical problem with a technical fix.
  4. Measure the answer surface separately. The AI Overview shows its sources; check the queries you care about by hand, or track them. Your log tells you who read you. Neither one substitutes for the other.

Conclusion

ChatGPT answers from three places — training data you cannot change now, a search index you can, and a live fetch that happens because somebody asked. Only the last two are worth optimising for this quarter.

You cannot audit that retrieval from outside. You can audit Google's AI Overview, and in 23 of them, 60% of the page-one domains were not cited and the #1 result was skipped nine times. Meanwhile 22% of citations went to sites nowhere near the top 20.

Being read is not being cited. Your log proves the first, cannot see the second, and the distance between them is the whole problem.


AI Citation and Retrieval: Frequently Asked Questions

Where does ChatGPT get its information from?
From three places: the training corpus it absorbed before release, its own search index built by crawling the web, and a live fetch of a specific page made in response to a question. The training corpus is fixed until the next training run, the index refreshes continually, and the live fetch happens in seconds — so a page published this month can only reach you through the second or third.
Does ChatGPT cite the pages it reads?
Not necessarily. Retrieving a page and citing it are separate events, and every AI system we asked said so directly. A page can be fetched, assessed and left out of the answer entirely, which is why a busy access log is not evidence that you were named.
If I rank 1 on Google will I be in the AI Overview?
Often, but far from always. Across 23 searches we checked, the 1 organic result was cited in the AI Overview on 14 of them — so on 9 it was not. Top-10 domains overall were cited about 40% of the time. Ranking improves your odds substantially without settling the question.
Can a page be cited by AI without ranking on Google?
Yes. In our sample, 22% of AI Overview citations went to domains that did not appear anywhere in the top 20 organic results for that search. The answer layer selects on more than rank order, which means a competitor can gain citations without ever showing up in your rank tracking.
How can I tell if AI is citing my site?
For Google's AI Overview, run the queries you care about and read the sources it lists. For ChatGPT and similar assistants there is no published per-answer source list you can audit from outside, so the practical options are asking the questions yourself and recording what comes back, or using a tool that does it at volume.
Does my server log show whether I was mentioned in an answer?
No. It shows that an agent requested a URL and that you served it. That is genuine evidence of being reachable and read, and it is worth having, but the answer that agent produced afterwards never touches your server and is not recorded anywhere you can see.
Should I focus on SEO or on AI visibility?
SEO first, because it is upstream. 61% of the citations we counted came from the top 10 organic results, so ranking materially improves your chances of being selected. Treat AI visibility as a second measurement on top of that work, not as a separate programme that replaces it. --- Meta Title: Where Does ChatGPT Get Its Information? Meta Description: ChatGPT answers from training data, a search index, or a live fetch. We checked 23 AI Overviews: only 40% of page-one domains were cited in the answer.

Find the exact growth leak in your business — in 2 minutes.

Paste your URL. Our AI agent crawls your site, diagnoses what's broken, and ships a step-by-step fix plan. Free, no signup.

Run free audit