How ChatGPT Finds Websites and Decides What to Cite

ChatGPT finds websites two ways: from what it learned during training, and — when browsing is active — from live web searches it runs mid-answer. Which one applies depends on the question, and each rewards a different kind of preparation.

For anything recent or specific, ChatGPT browses, fetches a shortlist of pages with its GPTBot crawler, and cites the ones it can quote. For broad, established topics, it often answers from training data and names the brands it already knows.

The two paths, and what each rewards

The same brand can win either way, but the levers differ:

PathWhen it firesWhat wins
Live browsingRecent, niche or specific queriesA fetchable page with a clean, quotable answer
Training dataBroad, well-established topicsA brand the model already recognises as an entity

Being fetchable by GPTBot

If your robots.txt disallows GPTBot, ChatGPT's browser cannot read the page, and the live path is closed to you. Check it first — it is the one setting that makes every other effort irrelevant. Our guide to getting cited by AI covers the exact directives.

Being recognised without a fetch

The training-data path is won long before any prompt, by being a resolvable entity: a Wikipedia or Wikidata presence, a consistent name across the web, and complete Organization schema. That is entity authority, and it is why a household name gets recommended off training data while a stronger page from an unknown brand does not.

The robots.txt rules that decide it

ChatGPT reaches your site through several distinct user agents, and they do different jobs. Blocking them as a group is the mistake that removes sites from live answers, usually without anyone intending it. The ChatGPT visibility checker sets out what can honestly be established about any of this from your HTML.

OAI-SearchBot fetches pages to build the search index behind live answers. ChatGPT-User fetches a page when a user's question causes the assistant to open it. GPTBot collects data used for model training. The first two are how you appear in answers today; the third is about training use, and blocking it is a legitimate position that carries no penalty in live retrieval.

The failure we see most often is a blanket rule inherited from a security plugin or CDN preset, written years ago, that disallows anything whose name matches a bot pattern. Nobody decided it. Read your robots.txt and decide deliberately.

User agentWhat it doesBlocking it means
OAI-SearchBotIndexes pages for live answersYou cannot appear in ChatGPT search
ChatGPT-UserFetches a page on demandYour page cannot be opened mid-answer
GPTBotCollects training dataNo training use; live answers unaffected
PerplexityBotIndexes for PerplexityAbsent from Perplexity answers
Google-ExtendedGemini training and groundingNo Gemini training use

What to check in your own logs

Your server access logs answer, for free, the question most AI visibility tools charge to guess at: are these crawlers actually fetching your pages?

Filter your logs for the user agents — OAI-SearchBot, ChatGPT-User, GPTBot, PerplexityBot, ClaudeBot, Google-Extended — and look at three things. Which URLs are being fetched, because it is rarely the ones you assume. What status codes they receive, since a run of 403s usually means a WAF is blocking them regardless of what robots.txt says. And whether fetch frequency changes after you publish or update, which tells you whether your sitemap and dates are being trusted.

This is the cheapest diagnostic available and almost nobody runs it. A site convinced it has a content problem quite often has a 403 problem instead.

Frequently asked questions

Does ChatGPT read my sitemap?+

Crawlers commonly use sitemaps for discovery, and a clean sitemap with honest lastmod values makes it easier for any of them to find and refresh your pages. It is not a guarantee of fetching, but its absence makes discovery slower.

Why is ChatGPT fetching pages I do not care about?+

Because those are the ones it can reach and parse. Orphaned or client-rendered priority pages get skipped in favour of whatever is well linked and server-rendered — which is itself the finding worth acting on.

Will blocking GPTBot remove me from ChatGPT?+

No. GPTBot is the training crawler. Live answers are served by OAI-SearchBot and ChatGPT-User, so you can decline training use and still appear in answers — provided you have not blocked those two as well.

Does ChatGPT always browse the web?+

No. It browses when the question needs current or specific information and browsing is enabled; otherwise it answers from training data. Preparing for both paths — a fetchable, quotable page and a recognised brand — is the only way to cover every query.

Sources

Published · Last reviewed .