Services · AI search

Generative engine optimisation: being cited, not just ranked

Your buyer now asks an assistant before they open a results page, and the assistant names two or three companies instead of ten blue links. AI search optimisation is the work of making sure your name is one of them, which starts with the unglamorous discovery that your firewall has been returning 403 to the crawlers that matter.

ChatGPTClaudePerplexityGemini
Quick answer

Generative engine optimisation is the work of making your site easy for AI assistants to reach, understand and quote. It has three levers: crawler access, so search and user-triggered agents are not blocked; entity clarity, so an engine can tell exactly who you are; and answer-first content it can lift a passage from. Progress is measured as share of citations across a fixed prompt set, per engine.

The shift

What actually changed

For twenty five years the shape of discovery was stable. A person typed a query, received a page of links, and chose one. Everything in search marketing followed from that: position on the page, click through rate, the long tail, the whole apparatus.

The new behaviour is different in a way that matters more than its size. Somebody types a question into an assistant and receives a paragraph that names two or three companies, with citations underneath that most people do not open. There is no page ten. There is barely a page one. If you are not in the paragraph, you were not considered, and you will never know it happened, because no impression was logged anywhere you can see.

This does not replace search. Traditional search volume remains enormous and Google's own results now carry generated summaries above the links, which is the same dynamic wearing a familiar interface. What has changed is the number of slots. Ten links became three named sources, and the competition for those three is not decided by the signals that decided the ten.

Why this lands hardest on considered purchases

The queries assistants are used for skew towards research: comparison, shortlisting, understanding a regulation, working out which kind of supplier one needs. That is precisely the top of a business to business funnel. A consultancy, a software vendor or a manufacturer with a long sales cycle is more exposed to this shift than a shop selling a commodity, because the assistant is doing the shortlisting that a buyer used to do across six browser tabs.

The uncomfortable part

Most of the sites we audit are not losing this on content quality. They are losing it because a machine could not fetch the page, or fetched it and could not tell which organisation the page belonged to. The work that moves the needle first is technical, cheap and slightly boring, which is why so few agencies lead with it.

Mechanics

How answer engines pick their sources

Three pipelines, three sets of levers, three different speeds of response. Almost every disappointing GEO programme we are asked to review has been optimising for one of them and reporting on another.

Retrieval from a search index
The assistant runs one or more searches against an index it controls or licences, then reads the top results and writes an answer from them. ChatGPT's search, Perplexity and Google's generated results all work broadly this way. The lever here is classic and familiar: be indexable, be in the index, rank well for the underlying queries the assistant generates, which are usually more literal and more question shaped than the ones humans type.
Live fetching by a user-triggered agent
Someone pastes your URL, or asks a question whose answer needs the current version of a page, and the assistant fetches it there and then. This is a different user agent from the indexer, it usually respects a different robots.txt token, and it renders limited or no JavaScript. The lever is access and server-side rendering: if the fetch returns a challenge page, a login wall or an empty shell that fills in on the client, the assistant reports that it could not read your site.
Pretrained memory
The model already knows things about your category from its training data, and will answer from that without fetching anything. This is the slowest and least controllable pipeline. You cannot edit it, you cannot request a recrawl, and it only refreshes when the vendor trains or updates a model. The lever is long-run and indirect: be described consistently across the wider web, in places that end up in training corpora, using the same name, the same category words and the same facts.
Access

The AI crawlers, and what blocking each one costs

The strategic point is one line long: training crawlers and answer crawlers are different crawlers, and the cost of blocking them is not remotely the same.

User agentOperatorWhat it doesIf you block it
GPTBotOpenAICollects content used to train modelsNo effect on citations. A rights and commercial decision, nothing more.
OAI-SearchBotOpenAIBuilds the index behind ChatGPT searchYou are removed from ChatGPT's search results and its citations.
ChatGPT-UserOpenAILive fetch triggered by a user's requestChatGPT cannot open your page when a user asks it to. Pasted links fail.
ClaudeBotAnthropicCollects content used to train modelsNo effect on citations. Same decision shape as GPTBot.
Claude-SearchBotAnthropicIndexes pages so Claude can cite them in search resultsYou disappear from Claude's cited sources.
Claude-UserAnthropicLive fetch when a user's request needs your pageClaude cannot read your site on demand for a user.
PerplexityBotPerplexityIndexes pages for citation in Perplexity answersYou are not available as a Perplexity source.
Perplexity-UserPerplexityLive fetch on behalf of a person who askedThe user-initiated fetch fails, though user-triggered agents are the least predictable class here.
GooglebotGoogleIndexes for Google Search, including generated summariesYou leave Google entirely. Very few organisations should consider this.
Google-ExtendedGoogleA robots.txt token, not a crawler: controls Gemini training and groundingExcluded from Gemini training and grounding. Ordinary Search indexing is unaffected.
ApplebotAppleIndexes for Siri, Spotlight and Apple search featuresYou are absent from Apple's search surfaces.
Applebot-ExtendedAppleA token controlling use of your content in Apple's foundation modelsExcluded from that training. Applebot indexing continues as normal.
Meta-ExternalAgentMetaCrawls for Meta's AI products and trainingRemoved from Meta AI's source pool.
AmazonbotAmazonCrawls to answer user queries in Amazon services including AlexaNot available to Alexa and related answer surfaces.
BytespiderByteDanceCollects content for ByteDance model trainingTraining exclusion. Frequently blocked for crawl-rate reasons rather than policy ones.

User agent names change and operators add new ones. Treat this as a starting list to verify against each vendor's published documentation and against your own edge logs, not as a permanent record.

Diagnosis

The WAF problem, and how to check it tonight

Here is the finding that keeps repeating across audits. The content is good. The schema is present. The site is fast. And the reason it is never cited is that a web application firewall, a CDN bot manager or an over-broad security rule has been returning 403 to AI user agents for the past eighteen months, and nobody knew, because a blocked crawler does not appear in any report anyone reads. There is no message in Search Console. There is no drop in analytics, because the request never became a session. The signal is an absence, and absences do not raise alerts.

Bot management products ship with rule groups that treat unfamiliar automated agents as hostile by default. That is entirely reasonable engineering for scrapers and credential stuffers, and it catches OAI-SearchBot and Claude-SearchBot in the same net. Meanwhile the marketing team, who own the visibility problem, have no access to the console where the rule lives and would not think to look there.

Three checks, about twenty minutes

First, read your robots.txt line by line, including any wildcard user agent block, and confirm you know why every disallow is there. Broad rules written years ago for scrapers routinely catch today's answer engines.

Second, get someone with access to the edge logs to filter the last thirty days by user agent for OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Googlebot and Applebot, and group by HTTP status. You are not looking at volume. You are looking for 403, 429, 503 and any redirect chain that ends somewhere unhelpful.

Third, request a handful of your important pages with one of those user agent strings set, from an address outside your network, and read the actual bytes that come back. A CAPTCHA page, a cookie wall that hides the body content, or an empty shell awaiting client-side rendering are all functionally identical to a block.

Then write the policy on purpose

Once you can see what is happening, the decision is straightforward and it is a business decision, not a security one. Allow the search and user-triggered agents. Decide on training crawlers deliberately, knowing that the decision has no effect on citations either way. Put those choices in robots.txt, then make sure the WAF agrees with robots.txt, because in our experience the two documents disagree more often than they match. This overlaps heavily with the work described under website management, since the people who can change the firewall are rarely the people who noticed the problem.

Structured data

Entity clarity and structured data

Schema.org markup is not decoration and it is not a ranking trick. It is the difference between a page that mentions a company and a page an engine can confidently attribute to one.

Editorial

Writing pages an engine can quote

An assistant does not cite pages. It lifts passages. That single fact changes how a page should be built more than any other consideration in this discipline.

Answer first, in forty to seventy words

Every page should open with a short, self-contained answer to the question the page exists for, written so that it survives being cut out and pasted into someone else's paragraph. No throat-clearing introduction, no company history, no scene setting. The first sentence states the answer; the next two or three qualify it. If that block cannot stand alone, it will not be lifted. There is one at the top of this page and you can judge for yourself whether it works.

Headings shaped like questions

Assistants generate their own internal queries, and those queries are more literal and more conversational than the phrases humans type into a search box. Headings that match that shape are easier to retrieve against: what it costs, how long it takes, what happens if you block it, when it is the wrong tool. Mix them with statement headings so the page still reads like prose written by a person.

Lists and comparison tables are quoted out of proportion

Structured formats are disproportionately represented in generated answers, because they are easy to parse, easy to attribute and easy to summarise without distortion. A comparison table with clear column headers does work that four paragraphs of careful prose will not. The crawler table further up this page exists for that reason as well as for yours.

Definitions, plainly stated

Write the sentence that starts X is. Assistants answer a great many definitional questions, and a clean definition block, marked up and placed near the top of the relevant section, is among the most reliably extracted units on the web. Glossaries earn their keep here in a way they did not five years ago.

Freshness, and why stale pages quietly lose

Engines favour recency when a question is time-sensitive, and a great deal more of the buying process is time-sensitive than people assume. Keep dateModified accurate in your schema and visible on the page, and make it honest: bumping the date without editing the content is a habit engines are getting better at seeing through. Real reviews on a real cycle, with the change reflected in the text, is the version that works. Pages that have not moved in three years lose to pages that were revised last month, even when the old page is better.

llms.txt: worth doing, not a strategy

llms.txt is a proposed convention: a markdown file at the root of your domain listing your important pages with a line of context each, so a model reading it gets a curated map rather than a scrape of your navigation. It costs an afternoon. Adoption is growing, particularly among documentation-heavy sites. The honest position is that public evidence it directly lifts citations remains thin, and no major engine has committed to consuming it. Publish one, because it is cheap and genuinely useful to anyone reading your site with a machine. Do not let anyone charge you for it as though it were the programme. Ours is at /llms.txt.

Measurement

Measuring something that has no rankings

You cannot get rank data out of an assistant, and any dashboard claiming to show your position in ChatGPT is showing you a modelled number.

What you want to knowHow it is actually measuredWhat it will not tell you
Are we named in AI answers?A fixed prompt set of 50 to 150 buyer questions, run monthly against each engine, with cited sources recorded.Anything about volume. This is a sample, not a census.
Are we gaining or losing ground?Share of citations over time on the same prompt set, tracked per engine because engines diverge sharply.Why a change happened. Correlation is all this method offers.
Which of our pages get used?Recording the specific URL cited, not just the domain, so you learn which page shapes get quoted.Whether a competing page was better or merely more reachable.
Who are we appearing beside?Logging every other organisation named in each answer. Often the most useful output of the whole exercise.Their strategy. Only that the engine considers them comparable to you.
Is any of it producing traffic?Referral traffic from assistant domains in Google Analytics 4, segmented and watched as a trend.Full attribution. Many assistant visits arrive with no referrer and land in direct.
Can engines reach us at all?Edge log monitoring by AI user agent and status code, with an alert on any sustained non-200.Whether being reachable led to being used. That is what the prompt set is for.

Answers are non-deterministic, so the same prompt run twice can differ. Read the trend across a large prompt set over months. Do not read a single run, and do not let anyone report one to you as a result.

Method

How a GEO engagement runs

Access first, entity second, content third. Doing them in any other order means spending on content that a crawler still cannot fetch.

  1. 01

    Baseline before touching anything

    We agree the prompt set with you, run it across ChatGPT, Claude, Perplexity and Google's generated results, and record who is cited today. Without this, everything afterwards is anecdote.

  2. 02

    Audit machine access

    Edge logs by user agent and status code, robots.txt line by line, WAF and bot management rules, rendering behaviour without JavaScript, and a live fetch test from outside your network. This phase finds the problem in most engagements.

  3. 03

    Fix access, and verify the fix

    Search and user-triggered agents allowed, training policy set deliberately, robots.txt and firewall brought into agreement, then re-tested rather than assumed. Access work is the fastest-acting change available.

  4. 04

    Make the entity unambiguous

    Organization, ProfessionalService, WebPage and BreadcrumbList implemented consistently, sameAs links pointed at independent registries, and naming reconciled across the site and your external profiles.

  5. 05

    Restructure the pages that matter

    Answer-first blocks, question-shaped headings, comparison tables, definition blocks, honest dateModified, and internal linking that makes the relationship between pages legible to a machine.

  6. 06

    Re-run, report, repeat

    The prompt set runs monthly. You get citation share by engine, the pages being quoted, the competitors sharing your answers, and a short list of what to change next. Programmes run in quarters, because that is the speed at which this actually moves.

Candour

What we do not know yet

This discipline is roughly two years old. The vocabulary is not settled, the same practice is sold as generative engine optimisation, answer engine optimisation and AI search optimisation depending on who is writing, and the underlying systems change without release notes. Anyone speaking about it with total confidence is telling you something about themselves.

So here is what is well established. Crawler access is causal and provable: a blocked agent cannot cite you, and removing the block restores the possibility. Entity resolution is well understood from a decade of knowledge graph work and the mechanics carry over. Extractable structure genuinely helps, which is observable in the passages engines actually quote.

And here is what is not. The relative weight of on-page changes against everything an engine already believes is not publicly known. Whether llms.txt influences any major engine today is unproven. How much third-party mentions matter compared with your own site is contested. Nobody outside the labs can separate retrieval effects from training effects with any rigour, and the published studies in this area are small, quickly outdated and often produced by firms selling the remedy.

Which leads to the only promise worth making. We can make you reachable, resolvable and quotable, and we can measure whether citations move. We cannot guarantee an engine will recommend you, because no such guarantee exists to sell. If a supplier offers you one, ask them to put the mechanism in writing.

Check our work on us

This site is built to the standard described on this page, which felt like the minimum before charging for it. View source and read the JSON-LD graph. Open /llms.txt. Note the extractable answer block at the top of every page, the question-shaped headings, the breadcrumbs, the markdown alternate on each URL and the reviewed date at the foot of the article. Then look at how our crawler policy separates training from search. Five minutes, and you will know whether we practise this or only sell it. The wider technical programme sits alongside ethical analytics and applied AI, because measurement and machine readability are the same problem viewed from two ends.

Questions

Frequently asked questions

Terms used on this page are defined in the glossary, and there is longer writing on this subject in insights.

What is generative engine optimisation?

Generative engine optimisation, or GEO, is the practice of making a website reachable, understandable and quotable by AI assistants, so that when someone asks ChatGPT, Claude, Perplexity, Gemini or Google's AI results about your category, your organisation is one of the sources named. It overlaps with search engine optimisation on fundamentals such as crawlability and clear information architecture, and departs from it on measurement, because there is no ranking to report. You are optimising for inclusion in an answer, not for a position in a list.

Is GEO different from SEO, or just a rebrand?

Mostly the same plumbing, a genuinely different objective. Both need a site that can be fetched, parsed and understood. GEO adds three things SEO does not have to worry about. First, a separate set of crawlers with their own user agents and their own robots.txt tokens. Second, entity resolution, because an engine has to be confident that the company in the document is the same company as the one in the question. Third, passage-level extractability, because the unit that gets used is a paragraph, not a page. Anyone claiming the two disciplines are unrelated is overselling. Anyone claiming they are identical has not looked at their server logs.

How do I know whether AI crawlers can reach my site?

Three checks, about twenty minutes. Read robots.txt line by line, including wildcard blocks. Filter your edge logs by AI user agent and group by HTTP status rather than by volume. Then request a page with one of those user agent strings set, from outside your network, and read what actually comes back. A 403, a challenge page or an empty shell awaiting client-side rendering all mean the same thing to an engine.

Does blocking GPTBot hurt my visibility in ChatGPT?

No, and this is the distinction that trips most teams up. GPTBot collects data used in training. OAI-SearchBot builds the index behind ChatGPT's search, and ChatGPT-User is the fetch that happens live when someone's question sends the assistant to a page. Disallowing GPTBot is a defensible commercial or rights decision with no effect on whether you are cited. Disallowing OAI-SearchBot or ChatGPT-User removes you from the answers entirely. The same split applies to Anthropic's ClaudeBot against Claude-SearchBot and Claude-User, and to Google-Extended, which is a training control that does not affect Googlebot's indexing.

What is llms.txt and do I need one?

It is a proposed convention: a markdown file at your domain root listing your important pages with a line of context each. Adoption is growing. Public evidence that it lifts citations on its own is still thin, and no major engine has committed to consuming it. Publish one, because it costs an afternoon and is useful to anyone reading your site with a machine. Do not buy it as a strategy.

How do you measure whether GEO is working?

With a prompt set, because there is no rank to pull. Fifty to a hundred and fifty questions a real buyer would ask, run monthly against each engine, with every cited source recorded. The output is citation share by engine over time, the pages of yours being quoted, and the competitors sharing your answers. Alongside it, assistant referral traffic in Google Analytics 4. Neither number is a ranking.

How long before we see anything change?

Access problems change fastest. If a firewall rule was blocking a search crawler and you remove it, inclusion can follow within days to a few weeks as that engine recrawls. Entity and structured data work shows up over a few weeks to a couple of months. Content restructuring is slower still, and pretrained model memory is slowest of all, because it only updates when the vendor trains or refreshes. Any timeline quoted with more precision than that is invented.

Can you guarantee we will be recommended by ChatGPT or Claude?

No, and neither can anybody else. There is no paid inclusion, no submission endpoint and no ranking signal to buy. Answers vary between runs of the same prompt, they vary by region and by the account asking, and the systems change without notice. What can be done is remove the reasons an engine cannot use you, make it obvious what you do and who you are, and measure citation share honestly over time. A guarantee of placement in an AI answer is a sales tactic, not a capability.

Should we block AI crawlers to protect our content?

It is a real decision with real trade-offs, and it should be made deliberately rather than by a default in a bot management console. Publishers with licensable archives often have good reason to disallow training crawlers. Almost nobody selling a professional service does. The pattern we keep finding is a site that intended to block training, wrote a broad rule, and removed itself from every AI answer surface at the same time. Separate the two questions, then write robots.txt and firewall policy that reflect the answers.

Do you do this for your own site?

This site is built to the standard described on this page. The schema graph is emitted on every page and you can read it in view source. There is an llms.txt at the root and a markdown alternate for each page. Every page opens with a short extractable answer, the crawler policy separates training from search and user-triggered fetching, and we run our own prompt set monthly. Checking that claim takes about five minutes and we would rather you did.

Written and reviewed by the Green Arrow Consultancy team, led by Darren Tyler, founder and chief executive.

Green Arrow Consultancy Ltd, Cardiff, Wales. Company number 12491770. ICO registration ZA822868. Member of the International Association of Privacy Professionals. Last reviewed .

Find out whether AI assistants can even reach you

We will run a prompt set for your category, check your edge logs and robots.txt against every major AI user agent, and tell you plainly what is blocking you. Most of the value is in the first week.