AI search optimisation checklist
Twenty-six checks for being cited by AI assistants, tickable on this page, plus an Excel tracker for measuring it. Most of this is ordinary good practice weighted differently rather than a new discipline — the genuinely new parts are deciding which AI crawlers to allow, writing passages that survive extraction, and accepting that there is no Search Console for this, so manual prompt checks are the measurement.
- CHECKLIST
- XLSX
- FREE
- Best for
- Anyone who wants to be the source an assistant cites, not just a ranking page
- Includes
- 26 tickable checks + Excel citation tracker + PDF
- Time to use
- 2 hours, then monthly checks
Free to download and use in your own client work. No email address required.
Problems this solves
Recurring complaints about this task, and what this resource does about each one. If it does not solve your version of the problem, that is worth knowing before you download anything.
- Nobody can tell you how to show up in ChatGPT or Perplexity
- This says plainly that eighty per cent of it is SEO fundamentals weighted differently, names the twenty per cent that is genuinely new, and separates what the providers document from what they do not — the mechanism is public, the weighting is not, including to me.
- You have no idea whether you are being cited
- The tracker is the measurement: your buyers’ actual prompts, checked across assistants on a schedule, recording cited / mentioned / absent and which competitor was cited instead. There is no Search Console for this.
- Every recommendation contradicts the last one
- A comparison table of what transfers from classic SEO and what actively works against you, so you can tell a real change in practice from a repackaged one.
- You are told to be cited on Reddit and LinkedIn and cannot tell if that is real
- The page covers third-party corroboration honestly: being discussed where assistants crawl does appear to help, it is not something you can reliably manufacture, and treating it as a tactic is how people get banned from the communities in question.
You will see this called AEO, GEO, LLM SEO or AI search optimisation. The acronym does not matter and it will change again. What matters is a narrower question:
When someone asks an assistant a question you could answer, are you the source it cites?
Eighty per cent of this is SEO you already know
Let me be honest about that before selling you a checklist. Crawlability, genuinely unique substance, a named accountable author, accurate structured data — these were already the things that worked, and they matter more here, not less.
What is genuinely different is worth knowing precisely:
| Aspect | Classic search | AI search |
|---|---|---|
| What wins | The best page for a query, judged partly on links | The most extractable, best-sourced passage that answers the specific question |
| Unit of competition | The page, against nine other pages | The passage, against passages from several sites blended into one answer |
| What a win looks like | A ranking position and a click | A citation, and often no click at all |
| Measurement | Impressions, position and clicks in Search Console | Manual prompt checks and referral traffic. There is no Search Console for this yet |
| What does not transfer | Keyword density, exact-match phrasing, thin pages at volume | All of it works against you — filler dilutes the passages you want extracted |
| What transfers completely | Crawlability, unique substance, a named accountable author, accurate structured data | The same four, weighted more heavily |
The last two rows are the important ones. Thin pages at volume actively hurt you here — filler dilutes the passages you want extracted — while the fundamentals transfer completely.
The checklist
Can the assistants reach you at all? 5
Is the answer extractable? 6
Is it worth citing? 5
Machine-readable signals 5
Measuring it 5
Deciding which crawlers to allow
This is the part that is genuinely new, and it is routinely confused. Blocking a training crawler does not stop you being cited in search answers. They are separate decisions.
| User agent | Who runs it | What it does | Advice |
|---|---|---|---|
| OAI-SearchBot | OpenAI | Surfacing pages in ChatGPT search results | Allow. This is how you get cited. |
| GPTBot | OpenAI | Collecting data used for model training | Owner’s decision. Blocking it does not stop you being cited by OAI-SearchBot. |
| PerplexityBot | Perplexity | Indexing for Perplexity answers, which cite sources prominently | Allow. |
| ClaudeBot | Anthropic | Crawling for Claude | Allow unless you have a reason not to. |
| Google-Extended | Controls Gemini training use — NOT Search indexing | Blocking it does not affect Google Search rankings. Decide separately. | |
| Bingbot | Microsoft | Bing index, which also feeds Copilot | Allow. Blocking it costs you Copilot visibility too. |
Google-Extended does not affect Google Search
Blocking Google-Extended controls whether your content is used for Gemini model training. It has no
effect on Google Search indexing or rankings. Plenty of sites have blocked it believing otherwise, and
plenty of others have left it open without ever making a decision. Either is defensible; not knowing which
you have done is not.
Writing a passage that survives extraction
The unit of competition is the passage, not the page. Four things follow from that:
Answer the question in the first 40–80 words, before context or preamble. An assistant extracting an answer takes the passage that contains one, and a page that warms up for three paragraphs offers nothing to take.
One page, one question. A page covering five topics gets cited for none of them, because no single passage in it fully answers anything.
Put facts in sentences and tables, not in images. A number inside a chart is invisible. A number in a table row with a real header can be read on its own.
Say what you do not know. This one surprises people: hedged, caveated pages get cited more, not less. “This works for X but fails at Y” is more useful to a system trying to answer accurately than a confident claim it cannot verify.
Being mentioned somewhere else
You will be told that getting discussed on Reddit, LinkedIn or YouTube is the fastest route to being cited. There is something real here and it is worth being careful about.
What appears to be true: assistants draw on community and discussion sources heavily, so a recommendation of you somewhere else can surface where your own page does not. Being the answer on a page you do not own is still being the answer.
What is not a tactic: posting about yourself. Communities like Reddit are unusually good at spotting it, and the outcome is a ban and a deleted account rather than a citation. Paying for placement in “best X” listicles has the same problem in a slower way — those pages get discounted as the pattern becomes obvious.
What actually works, and is slow: being genuinely useful in places where your buyers already are, under your own name, without linking to yourself most of the time. That is a reputation strategy with an SEO side-effect, not an SEO tactic — and it is worth being honest that it takes months and cannot be outsourced cheaply.
Do not build a strategy on one anecdote
There are widely-shared posts claiming a single Reddit thread produced an enormous result. Some are true and none are repeatable on demand — the ones you see are the survivors, and the same approach applied deliberately at scale is what gets accounts banned. Treat third-party mentions as an outcome of doing good work publicly, not as a channel you can buy.
Measuring it, honestly
There is no Search Console for AI search. Anyone offering you a precise “AI visibility score” is offering a modelled estimate.
What works is unglamorous: a list of prompts a buyer would actually type, checked across two or three assistants on a fixed schedule, recorded. The tracker is where that goes.
| Column | What to put in it |
|---|---|
| Prompt | The words a buyer would actually type, not a keyword. |
| Intent | Comparison, how-to, recommendation, definition. Dropdown. |
| Assistant | ChatGPT, Perplexity, Claude, Google AI Overview. Dropdown. |
| Date checked | Answers change week to week, so an undated check is worthless. |
| Cited? | Yes / Mentioned / No. Dropdown. "Mentioned" means named without a link. |
| Your URL used | Which page it pulled. Often not the one you would expect. |
| Competitors cited | Who was cited instead. The most useful column on the sheet. |
| Gap | What their cited page has that yours does not. |
| Action | What you changed. Then re-check on a stated date. |
The competitors cited column is the most useful thing on the sheet. It tells you what a citable page looks like for your topic, which is far more actionable than your own citation rate.
Treat the citation rate as a direction, not a metric
Answers vary between sessions and change week to week, so a single check proves very little. The same prompt list re-run monthly shows movement; one check shows weather. And watch referral traffic from chatgpt.com and perplexity.ai in analytics — small numbers, but they are real and they are yours.
This site is the worked example
Everything on this page is applied here, which is the fairest way to judge it. Every resource opens with a
40–80 word direct answer that the build fails without. robots.txt names the assistant crawlers
explicitly. There is one Person entity with the same @id sitewide and sameAs links to profiles that
corroborate it. Facts sit in tables with real headers, and the pages say where the approach fails.
Whether it works is measurable by you: ask an assistant something this site should be able to answer, and see who it cites. How the resources are made sets out the rest of the method.
Common mistakes
- Blocking training crawlers and assuming you have also opted out of being cited, or vice versa.
- Publishing thin pages at volume, which dilutes the passages you want extracted.
- Adding FAQ markup for questions the page does not actually answer.
- Burying the answer under three paragraphs of context.
- Buying a tool that reports an “AI visibility score” as though it were measured.
- Judging the work on sessions, when a citation often produces no click at all.