Robots.txt & Sitemap Generator
Control what crawlers may fetch — and decide explicitly whether AI crawlers are welcome (for GEO, they should be: no crawl, no citation).
How to use this
- 1Enter your website address.
- 2Choose whether to let AI assistants (like ChatGPT) read your site — we recommend yes.
- 3Copy the robots.txt text and save it as a file named robots.txt at the top level of your website.
- 4Copy the sitemap and save it as sitemap.xml — it’s the map search engines use to find all your pages.
robots.txt
robots.txt — serve at your domain root
User-agent: * Disallow: /admin Disallow: /api User-agent: GPTBot Allow: / User-agent: ClaudeBot Allow: / User-agent: PerplexityBot Allow: / User-agent: Google-Extended Allow: / User-agent: CCBot Allow: /
XML sitemap
List your canonical URLs only. Serve the file at /sitemap.xml and declare it in robots.txt above.
Paste URLs — the sitemap appears here.
The rules that matter most in 2026
A robots.txt file now has to make two separate decisions that used to be one: what traditional search crawlers may fetch, and what AI crawlers may fetch. Treating them as a single group is the most common and most costly mistake in this file.
AI crawlers split further. Live search crawlers — OAI-SearchBot, PerplexityBot, Claude-SearchBot — fetch pages to answer questions right now, so blocking them removes you from AI answers today. Training crawlers such as GPTBot, ClaudeBot, CCBot and Google-Extended collect data for model training, and declining them is a legitimate content position that does not affect live retrieval.
Two failure modes account for most of the damage we see. A blanket rule inherited from a security plugin or CDN preset that disallows anything matching a bot pattern, catching the search crawlers nobody meant to block. And a robots.txt that allows everything while a web application firewall returns 403 to the same user agents, so the file says one thing and the server does another.
Check your access logs after any change: robots.txt is a request, and the only proof that a crawler is getting through is a 200 in your own logs.
Sitemaps deserve the same scepticism. A sitemap listing URLs that redirect, 404 or carry a canonical pointing elsewhere trains crawlers to trust it less, and honest lastmod values matter more than complete ones — a file claiming every page changed today is information-free.
Sources
- Google Search Central — Introduction to robots.txt
- Google Search Central — Overview of Google crawlers
- schema.org — Getting started
Last reviewed .
