How ChatGPT Finds Websites and Decides What to Cite
ChatGPT finds websites two ways: from what it learned during training, and — when browsing is active — from live web searches it runs mid-answer. Which one applies depends on the question, and each rewards a different kind of preparation.
For anything recent or specific, ChatGPT browses, fetches a shortlist of pages with its GPTBot crawler, and cites the ones it can quote. For broad, established topics, it often answers from training data and names the brands it already knows.
The two paths, and what each rewards
The same brand can win either way, but the levers differ:
| Path | When it fires | What wins |
|---|---|---|
| Live browsing | Recent, niche or specific queries | A fetchable page with a clean, quotable answer |
| Training data | Broad, well-established topics | A brand the model already recognises as an entity |
Being fetchable by GPTBot
If your robots.txt disallows GPTBot, ChatGPT's browser cannot read the page, and the live path is closed to you. Check it first — it is the one setting that makes every other effort irrelevant. Our guide to getting cited by AI covers the exact directives.
Being recognised without a fetch
The training-data path is won long before any prompt, by being a resolvable entity: a Wikipedia or Wikidata presence, a consistent name across the web, and complete Organization schema. That is entity authority, and it is why a household name gets recommended off training data while a stronger page from an unknown brand does not.
Frequently asked questions
Does ChatGPT always browse the web?+
No. It browses when the question needs current or specific information and browsing is enabled; otherwise it answers from training data. Preparing for both paths — a fetchable, quotable page and a recognised brand — is the only way to cover every query.
