How ChatGPT Finds Websites and Decides What to Cite
ChatGPT finds websites two ways: from what it learned during training, and — when browsing is active — from live web searches it runs mid-answer. Which one applies depends on the question, and each rewards a different kind of preparation.
For anything recent or specific, ChatGPT browses, fetches a shortlist of pages with its GPTBot crawler, and cites the ones it can quote. For broad, established topics, it often answers from training data and names the brands it already knows.
The two paths, and what each rewards
The same brand can win either way, but the levers differ:
| Path | When it fires | What wins |
|---|---|---|
| Live browsing | Recent, niche or specific queries | A fetchable page with a clean, quotable answer |
| Training data | Broad, well-established topics | A brand the model already recognises as an entity |
Being fetchable by GPTBot
If your robots.txt disallows GPTBot, ChatGPT's browser cannot read the page, and the live path is closed to you. Check it first — it is the one setting that makes every other effort irrelevant. Our guide to getting cited by AI covers the exact directives.
Being recognised without a fetch
The training-data path is won long before any prompt, by being a resolvable entity: a Wikipedia or Wikidata presence, a consistent name across the web, and complete Organization schema. That is entity authority, and it is why a household name gets recommended off training data while a stronger page from an unknown brand does not.
The robots.txt rules that decide it
ChatGPT reaches your site through several distinct user agents, and they do different jobs. Blocking them as a group is the mistake that removes sites from live answers, usually without anyone intending it. The ChatGPT visibility checker sets out what can honestly be established about any of this from your HTML.
OAI-SearchBot fetches pages to build the search index behind live answers. ChatGPT-User fetches a page when a user's question causes the assistant to open it. GPTBot collects data used for model training. The first two are how you appear in answers today; the third is about training use, and blocking it is a legitimate position that carries no penalty in live retrieval.
The failure we see most often is a blanket rule inherited from a security plugin or CDN preset, written years ago, that disallows anything whose name matches a bot pattern. Nobody decided it. Read your robots.txt and decide deliberately.
| User agent | What it does | Blocking it means |
|---|---|---|
| OAI-SearchBot | Indexes pages for live answers | You cannot appear in ChatGPT search |
| ChatGPT-User | Fetches a page on demand | Your page cannot be opened mid-answer |
| GPTBot | Collects training data | No training use; live answers unaffected |
| PerplexityBot | Indexes for Perplexity | Absent from Perplexity answers |
| Google-Extended | Gemini training and grounding | No Gemini training use |
What to check in your own logs
Your server access logs answer, for free, the question most AI visibility tools charge to guess at: are these crawlers actually fetching your pages?
Filter your logs for the user agents — OAI-SearchBot, ChatGPT-User, GPTBot, PerplexityBot, ClaudeBot, Google-Extended — and look at three things. Which URLs are being fetched, because it is rarely the ones you assume. What status codes they receive, since a run of 403s usually means a WAF is blocking them regardless of what robots.txt says. And whether fetch frequency changes after you publish or update, which tells you whether your sitemap and dates are being trusted.
This is the cheapest diagnostic available and almost nobody runs it. A site convinced it has a content problem quite often has a 403 problem instead.
Frequently asked questions
Does ChatGPT read my sitemap?+
Crawlers commonly use sitemaps for discovery, and a clean sitemap with honest lastmod values makes it easier for any of them to find and refresh your pages. It is not a guarantee of fetching, but its absence makes discovery slower.
Why is ChatGPT fetching pages I do not care about?+
Because those are the ones it can reach and parse. Orphaned or client-rendered priority pages get skipped in favour of whatever is well linked and server-rendered — which is itself the finding worth acting on.
Will blocking GPTBot remove me from ChatGPT?+
No. GPTBot is the training crawler. Live answers are served by OAI-SearchBot and ChatGPT-User, so you can decline training use and still appear in answers — provided you have not blocked those two as well.
Does ChatGPT always browse the web?+
No. It browses when the question needs current or specific information and browsing is enabled; otherwise it answers from training data. Preparing for both paths — a fetchable, quotable page and a recognised brand — is the only way to cover every query.
Sources
Published · Last reviewed .
