AI Readiness Checklist: 12 Fixes That Get You Cited

AI readiness is how prepared a page is to be found, read and cited by AI assistants. This checklist covers the twelve highest-leverage fixes, ordered so each one builds on the last.

Retrievability: let AI in

An engine that cannot fetch your page cannot cite it. Start at the gate.

  1. Allow AI search crawlers in robots.txt — OAI-SearchBot, ChatGPT-User, Claude-SearchBot, PerplexityBot. Blocking them removes you from live answers.
  2. Decide deliberately about training crawlers (GPTBot, ClaudeBot, Google-Extended). Blocking is a legitimate IP stance, but it slows how fast models learn your brand.
  3. Publish an llms.txt at your domain root listing your most citable pages with one-line descriptions.
  4. Remove accidental noindex directives — a noindex page is invisible to both Google and AI retrieval.

Extractability: give AI something to lift

Once the engine is in, structure decides whether anything is quotable.

  1. Add Organization schema site-wide with name, url, logo, contactPoint and sameAs.
  2. Add Article schema on content pages with headline, author, datePublished and dateModified.
  3. Add FAQPage schema wherever you genuinely answer questions.
  4. Open every H2 section with a one-to-two sentence direct answer, under 30 words.
  5. Convert steps into numbered lists and comparisons into real tables.

Trust: give AI a reason to pick you

Between two extractable pages, engines cite the one with stronger provenance.

  1. Show a visible author byline and mirror it in schema.
  2. Display published and updated dates in <time> elements.
  3. Cite authoritative sources — standards bodies, government data, major references — with descriptive anchor text.

How to work through the list without wasting effort

The order of this checklist is the whole point, because the fixes are not independent. Retrievability gates everything: schema, direct answers and author bylines are all worthless on a page an engine cannot fetch or parse. Work top to bottom and stop when you hit something you cannot fix yet, rather than skipping ahead to the easy items.

In practice most sites find the first section takes an afternoon and the second takes a week. The third — trust and entity authority — is the slowest, because it depends on records outside your own site, and it is also the one most teams never start. If you only do one thing beyond the technical fixes, claim and align your public profiles so your brand name resolves consistently.

Re-run an audit after each section rather than at the end. The signals are structural, so a fix either registers or it did not work, and finding that out immediately is much cheaper than discovering it after three weeks of writing. The four faults we find on nearly every site are the ones worth re-checking first.

Frequently asked questions

Which fix on this list matters most?+

Server-side rendering, if you are missing it. Every other item on the list assumes the crawler can read your content, and a client-rendered page gives a non-executing crawler an empty shell. Nothing else you do registers until that is resolved.

Should I block AI training crawlers?+

That is a business decision, not a technical one, and both answers are defensible. Blocking GPTBot or ClaudeBot protects your content from training use; it also slows how quickly models learn your brand exists. What is rarely defensible is blocking the live search crawlers — OAI-SearchBot, PerplexityBot, Claude-SearchBot — since those are what fetch pages to answer questions right now.

Is llms.txt worth adding?+

It costs ten minutes and no engine is documented as requiring it, so treat it as cheap optionality rather than a ranking factor. The genuine benefit is the exercise: deciding which twenty pages you would most want quoted usually reveals that several of them are not ready.

How long does it take to see AI citations after these fixes?+

Live-retrieval engines like Perplexity can cite you within days of a fix because they fetch pages at question time. Training-data effects take months. Fix retrievability and extractability first — they pay off fastest.

Sources

Published · Last reviewed .