There is a lot of advice on this subject and very little of it separates the parts that are proven from the parts that are hopeful. This is the version we actually run for clients, in order, with the uncertainty left in.
To be cited by an AI assistant you need three things: the assistant's crawler must be able to reach your page, the page must state its answer in a form that can be lifted out cleanly, and independent sources must corroborate that you exist. Most sites fail on the first, which silently cancels the other two. Test crawler access first, then structure, then off-site evidence.
Almost every argument about generative engine optimisation goes wrong because the people having it are describing different mechanisms without saying so. There are three ways a sentence from your website can end up in a machine-generated answer, and they behave nothing like each other.
The assistant decides it needs current information, issues a query to a search index, retrieves a handful of documents and writes an answer grounded in them. This is how ChatGPT search, Perplexity, Claude's web search and Google's AI Overviews mostly work. The index is built by a crawler that visited you earlier, so there is a lag between publishing and being retrievable. This route is the one you can influence most directly, and it is where the majority of business-relevant citations come from today.
Someone pastes your URL into a chat window, or asks a question that makes the assistant open a specific page right now. A different agent does this work, with a different user agent, in real time. There is no index and no lag. If your edge returns a 403 to that agent, the user sees the assistant fail to read a page they can open in their own browser, which is a bad look you will never be told about.
The model recalls something about you from its training data. This is the route people imagine when they talk about being known by AI, and it is the one you have the least control over. Training runs happen on the vendor's schedule, the corpus is largely fixed at that point, and nothing you publish this quarter affects a model that finished training last year. Memory also degrades into confident approximation, which is why assistants get company details subtly wrong even when the company is well known.
The practical consequence is that a strategy aimed at route three is mostly a strategy aimed at nothing you can measure. Aim at routes one and two, and treat any improvement in route three as a slow, unattributable bonus. Our AI search optimisation work is built around that ordering.
Three mechanisms, three different sets of levers. Confusing them is the most common reason a GEO programme produces no measurable change.
| Route | Who does the fetching | What you control | What you do not |
|---|---|---|---|
| Search index | An indexing crawler such as OAI-SearchBot, PerplexityBot or Googlebot | Access, structure, freshness, internal linking, entity clarity | Query rewriting, how many sources the model chooses to cite, ranking inside the index |
| Live user fetch | A user-triggered agent such as ChatGPT-User, Perplexity-User or Claude-User | Access, page speed, whether the answer is above the fold in the HTML, rendering without JavaScript | Which page the user or the assistant decides to open in the first place |
| Pretrained memory | A training crawler such as GPTBot, ClaudeBot or Meta-ExternalAgent | Whether you permit training use at all, and how clearly you are described where it crawled | The training schedule, the corpus mix, and whether the model recalls you correctly |
Route names are ours. The vendors do not use a shared vocabulary for these, which is part of the problem.
Nine items, in this order. The order matters more than the list, because the first two can silently invalidate the other seven.
Request three or four of your own pages with each assistant's user agent and record the HTTP status. You are looking for 200. A 403, a 429, a 503 or an interstitial challenge page all mean you are invisible. Do this before anything else, because every later item on this list depends on it and none of them will tell you it has failed.
Robots.txt is a request that well-behaved crawlers honour. A web application firewall is a wall. Allowlist the search and user-triggered agents at the edge by user agent and by the operators' published IP ranges, because user-agent strings are trivially spoofed and an allowlist based on them alone is a hole. This is involved enough that it has its own article.
Organization, WebPage, Article and FAQPage markup with stable identifiers, and the same company details repeated on the pages and profiles a retrieval system can cross-check. The goal is not a rich result. The goal is that a machine reading your page can tell which real-world organisation it is about.
One self-contained paragraph that answers the page's core question in plain language, with the direct answer in the first sentence. No preamble, no scene-setting, no dependency on the sentence before it. This is the block that gets lifted, and it is the single highest-yield piece of writing on the page.
What does this cost. How long does it take. What is the difference between these two things. Then answer immediately underneath, in the first sentence, before the qualification. Retrieval systems match on question phrasing far more than on keyword density.
Structured comparisons survive extraction intact in a way that a paragraph does not, and they are quoted disproportionately as a result. A table of five options with four attributes each is more useful to a retrieval system than two thousand words describing the same five options.
Move the timestamp when the content moves and not otherwise. Bulk-refreshing dates across an estate to look current is easy to detect, easy to discount, and it destroys your own ability to tell which pages are genuinely stale.
Trade publications, standards bodies, partner sites, directories, conference programmes, open datasets, review platforms, company registers. Assistants weight independent corroboration heavily because a claim on your own site is, from a retrieval system's point of view, just you saying it about yourself.
Thirty to eighty questions your buyers would actually type, run across each engine on the same schedule, recorded consistently. Share over time is the metric. Any single answer is noise.
The unit of extraction is the passage, not the page. A retrieval system splits your content into chunks, embeds them, and pulls back the handful that match the query. Whatever surrounds a chunk is not guaranteed to come with it. That single fact should change how you write.
It means the anaphora habit that good prose relies on, referring back to the thing you just described, works against you. A paragraph that opens with "This means that" is a paragraph that means nothing when it arrives on its own. Name the subject again. Repeat the noun. It reads slightly heavier to a human and it reads correctly to a machine, and the tradeoff is worth making in the sections that carry the answers.
It also means the answer has to arrive before the reasoning. Most business writing is structured as context, then argument, then conclusion, because that is how the writer discovered it. Invert it. State the conclusion, then support it. If a reader stops after the first sentence of a section, they should still have the answer.
Every page on this site opens with one. It is short enough to be quoted whole, long enough to be substantive, self-contained enough to survive being ripped out of context, and written so the first sentence alone would satisfy someone in a hurry. Below 40 words it tends to lose the qualification that makes it accurate. Above about 70 it stops being quotable and starts being summarised, and a summary of your words is worth less to you than your words.
Hedged writing is unquotable. "Solutions can vary depending on your requirements" contains no information and no retrieval system will ever surface it. "A consent management platform deployment typically takes three to six weeks, and the slow part is auditing the tags, not configuring the banner" is quotable because it commits to something. Committing is uncomfortable and it is most of the job. Our glossary exists for the same reason: definitions are the most extractable content format there is.
This is not a ranking of importance. It is a list of the shapes that survive being cut out of your page and pasted into somebody else's answer.
Schema.org markup is widely oversold as an AI visibility technique. It is not a ranking factor in any assistant that has said anything public on the subject, and no amount of JSON-LD will get a thin page cited. What it does do is remove ambiguity, and ambiguity is a reason for a retrieval system to skip you.
The concept that matters is entity resolution. When a machine reads a page about Green Arrow Consultancy, it has to decide whether that is the same organisation as the one in the Companies House register under number 12491770, the one on the ICO's register under ZA822868, and the one with the LinkedIn company page. If those records agree with each other, the entity resolves and the system can attribute claims to a real organisation. If the name is spelled three ways across four sources and the address differs, it cannot, and a cautious system prefers a source it can pin down.
So the practical work is duller than the marketing suggests. Publish Organization markup once, sitewide, with a stable identifier. Include the registration numbers, the VAT number, the registered address and the sameAs links to the external records that corroborate them. Make sure the details on your contact page match the details in the markup, and that both match the register. Then add page-level markup that describes what the page actually is: Article on an article, FAQPage where there is a genuine FAQ, Service on a service page.
Marking up FAQ schema for questions that do not appear on the page, or Review schema for reviews that do not exist, is the fastest way to have your structured data ignored across the board. It is also, in the case of invented reviews, a consumer protection problem before it is a technical one. We have audited estates where an agency had bulk-applied review markup with no reviews behind it, and the remediation was more expensive than the original build.
Of everything on the checklist, the item that most reliably separates the businesses that get cited from the ones that do not is whether anyone else on the internet has written about them. This is uncomfortable because it is the item you control least and the one that takes longest.
The reasoning is straightforward once you think about the retrieval problem. An assistant assembling an answer about a category of supplier has to choose two or three names. Everything on a supplier's own site is self-description. A trade publication, a standards body, a partner's case study, a conference programme, a public register, an open dataset or a comparison site is not. Those sources are what turns a claim into a corroborated fact, and corroborated facts are what a system reaches for when it has to commit to naming someone.
In practice this means the highest-yield activity is often not writing another page. It is getting listed in the directories your sector actually uses, publishing data or research that other people cite, appearing on the register that governs your work, contributing to the specification everyone in your field reads, or being the named partner in someone else's write-up. None of it is fast and none of it is a growth hack.
It also means the old public relations discipline is more relevant to AI visibility than most of the technical tactics being sold under the GEO label. If you have a communications team, they are doing generative engine optimisation whether or not anyone has told them so.
Six numbers, recorded on a schedule. Run each prompt several times per engine, because these systems are non-deterministic and a single response tells you almost nothing.
This field is eighteen months old as a commercial discipline and a good deal of what is confidently asserted in it has never been tested. Here is where we think the uncertainty actually sits.
Unproven: that llms.txt lifts citations. The file is cheap to publish and there is no evidence that any major assistant treats it as a retrieval or ranking input. We publish one and we do not count it as a strategy. There is a longer piece on that.
Unproven: that there is a stable ranking algorithm to reverse-engineer. Retrieval behaviour changes without announcement, models are swapped, and the same prompt returns different sources on different days. Anyone selling a fixed set of ranking factors for AI answers is selling a snapshot of noise.
Contested: how much organic ranking carries over. Published overlap figures between organic search results and AI citations vary enormously by study and by engine, and the honest summary is that Google's AI surfaces lean heavily on what already ranks while ChatGPT's citation set overlaps considerably less. We go into that in detail elsewhere.
Reasonably well supported: access, structure and corroboration. Crawler access is not a theory, it is a status code. Extractability is observable in what gets quoted. The weighting of independent sources is visible in the citation sets themselves. These three are where we would put the budget.
If you want the work done rather than described, that is what our AI search optimisation service is, and the broader AI consulting practice is where it sits when the answer turns out to be a system rather than a set of pages.
Questions we get asked in the first meeting, answered without the sales layer.
If the problem was crawler access, changes can show up within days on the live-fetch surfaces and within a few weeks on the indexed ones, because the index has to be rebuilt before you exist inside it. If the problem is that nobody outside your own domain has written about you, that is a six to twelve month job and no amount of on-page work shortens it. Pretrained memory is slower still and effectively outside your control, since it only changes when a model is retrained.
Not in the sense of buying a citation. Advertising is arriving in some assistants, but sponsored placement is labelled separately from the cited sources and does not make you a source. Anyone selling guaranteed placement inside an organic AI answer is selling something they do not control.
You need to write more decisively, which usually improves the page for people too. The specific changes are answering in the first paragraph rather than the fifth, making each section self-contained so a passage still makes sense when it is lifted out, and stating claims plainly instead of hedging them into vapour. None of that requires a separate version of the page and you should not build one.
No single markup change causes a citation. Structured data helps in a narrower way: it makes the entity on the page unambiguous, so a retrieval system can tell that the Green Arrow Consultancy on this page is the same company as the one in the Companies House record and the LinkedIn profile. Ambiguity is a reason to skip a source. Removing ambiguity is worth doing even though it is not a ranking lever.
Start with whichever one your buyers actually use, which you can find by asking them rather than by reading a market share chart. In practice the underlying work is shared: crawler access, clean structure and independent corroboration help everywhere. Where the engines diverge is in what they retrieve from, which is why measurement needs to be per engine even when the work is not.
Sometimes. If the wrong claim comes from a page you control, fix the page and the live-fetch and search surfaces will follow. If it comes from a third-party page, the correction has to happen there. If it comes from pretrained memory, there is no removal process worth relying on and the practical remedy is to publish correct, well-structured material that retrieval will prefer over the model's recollection.
Only if the extra content answers questions that nobody has answered well. Volume on its own dilutes. We have seen estates with hundreds of near-duplicate pages get cited less than a competitor with thirty specific ones, because near-duplicates compete with each other for the same retrieval slot and none of them wins clearly.
Build a fixed prompt set of thirty to eighty questions your buyers would ask, run them across the engines on the same schedule, and record whether you were named, which URL was cited, and who else appeared. Track the share over time rather than any single run, because these systems are non-deterministic and one answer proves nothing.
Checking that the crawlers can reach you. It is the only item on the list that can silently zero out everything else, it takes an afternoon to test, and a surprising share of well-run corporate estates fail it without anyone knowing.
We will run the crawler access test against your estate, check what an assistant can actually extract from your key pages, and tell you which of the nine items are already fine.