# Generative Engine Optimisation Agency | Green Arrow Consultancy

> Generative engine optimisation and AI search optimisation: AI crawler access, entity schema, answer-first content and citation tracking, engine by engine.

Source: https://greenarrow.app/services/ai-search-optimisation/
Last updated: 2026-09-04
Publisher: Green Arrow Consultancy Ltd, Cardiff, Wales, United Kingdom

---

Services · AI search

# Generative engine optimisation: being cited, not just ranked

Your buyer now asks an assistant before they open a results page, and the assistant names two or three companies instead of ten blue links. AI search optimisation is the work of making sure your name is one of them, which starts with the unglamorous discovery that your firewall has been returning 403 to the crawlers that matter.

ChatGPT Claude Perplexity Gemini 
 Get an AI visibility audit → Read our llms.txt 
 
 
 
 robots.txt WAF rules Schema 
 
 reachable · resolvable · quotable

“Who should we talk to about AI governance in the UK?”

Answer names 3 firms, one of them yours cited

Quick answer

Generative engine optimisation is the work of making your site easy for AI assistants to reach, understand and quote. It has three levers: crawler access, so search and user-triggered agents are not blocked; entity clarity, so an engine can tell exactly who you are; and answer-first content it can lift a passage from. Progress is measured as share of citations across a fixed prompt set, per engine.

## Key points

- Answer engines source content three different ways, and each pipeline has a different lever. Treating them as one thing is the commonest strategic error.

- Blocking a training crawler costs you nothing in citations. Blocking a search or user-triggered agent removes you from AI answers completely.

- A firewall or bot manager silently returning 403 to AI agents is the single most common cause of invisibility, and it is checkable in minutes.

- Structured data is not decoration here. It is how an engine resolves the words on your page to a real organisation it can name with confidence.

- There is no rank to report, so measurement means citation share across a fixed prompt set, run monthly, per engine.

- This discipline is about two years old. The causal evidence is thinner than the market implies, and guaranteed placement in an AI answer does not exist.

## On this page

- What actually changed

- How answer engines pick their sources

- The AI crawlers, and what blocking each one costs

- The WAF problem, and how to check it tonight

- Entity clarity and structured data

- Writing pages an engine can quote

- Measuring something that has no rankings

- How a GEO engagement runs

- What we do not know yet

- Frequently asked questions

The shift

## What actually changed

For twenty five years the shape of discovery was stable. A person typed a query, received a page of links, and chose one. Everything in search marketing followed from that: position on the page, click through rate, the long tail, the whole apparatus.

The new behaviour is different in a way that matters more than its size. Somebody types a question into an assistant and receives a paragraph that names two or three companies, with citations underneath that most people do not open. There is no page ten. There is barely a page one. If you are not in the paragraph, you were not considered, and you will never know it happened, because no impression was logged anywhere you can see.

This does not replace search. Traditional search volume remains enormous and Google's own results now carry generated summaries above the links, which is the same dynamic wearing a familiar interface. What has changed is the number of slots. Ten links became three named sources, and the competition for those three is not decided by the signals that decided the ten.

### Why this lands hardest on considered purchases

The queries assistants are used for skew towards research: comparison, shortlisting, understanding a regulation, working out which kind of supplier one needs. That is precisely the top of a business to business funnel. A consultancy, a software vendor or a manufacturer with a long sales cycle is more exposed to this shift than a shop selling a commodity, because the assistant is doing the shortlisting that a buyer used to do across six browser tabs.

### The uncomfortable part

Most of the sites we audit are not losing this on content quality. They are losing it because a machine could not fetch the page, or fetched it and could not tell which organisation the page belonged to. The work that moves the needle first is technical, cheap and slightly boring, which is why so few agencies lead with it.

Mechanics

## How answer engines pick their sources

Three pipelines, three sets of levers, three different speeds of response. Almost every disappointing GEO programme we are asked to review has been optimising for one of them and reporting on another.

**Retrieval from a search index**

The assistant runs one or more searches against an index it controls or licences, then reads the top results and writes an answer from them. ChatGPT's search, Perplexity and Google's generated results all work broadly this way. The lever here is classic and familiar: be indexable, be in the index, rank well for the underlying queries the assistant generates, which are usually more literal and more question shaped than the ones humans type.

**Live fetching by a user-triggered agent**

Someone pastes your URL, or asks a question whose answer needs the current version of a page, and the assistant fetches it there and then. This is a different user agent from the indexer, it usually respects a different robots.txt token, and it renders limited or no JavaScript. The lever is access and server-side rendering: if the fetch returns a challenge page, a login wall or an empty shell that fills in on the client, the assistant reports that it could not read your site.

**Pretrained memory**

The model already knows things about your category from its training data, and will answer from that without fetching anything. This is the slowest and least controllable pipeline. You cannot edit it, you cannot request a recrawl, and it only refreshes when the vendor trains or updates a model. The lever is long-run and indirect: be described consistently across the wider web, in places that end up in training corpora, using the same name, the same category words and the same facts.

Access

## The AI crawlers, and what blocking each one costs

The strategic point is one line long: training crawlers and answer crawlers are different crawlers, and the cost of blocking them is not remotely the same.

| User agent | Operator | What it does | If you block it |
|---|---|---|---|
| GPTBot | OpenAI | Collects content used to train models | No effect on citations. A rights and commercial decision, nothing more. |
| OAI-SearchBot | OpenAI | Builds the index behind ChatGPT search | You are removed from ChatGPT's search results and its citations. |
| ChatGPT-User | OpenAI | Live fetch triggered by a user's request | ChatGPT cannot open your page when a user asks it to. Pasted links fail. |
| ClaudeBot | Anthropic | Collects content used to train models | No effect on citations. Same decision shape as GPTBot. |
| Claude-SearchBot | Anthropic | Indexes pages so Claude can cite them in search results | You disappear from Claude's cited sources. |
| Claude-User | Anthropic | Live fetch when a user's request needs your page | Claude cannot read your site on demand for a user. |
| PerplexityBot | Perplexity | Indexes pages for citation in Perplexity answers | You are not available as a Perplexity source. |
| Perplexity-User | Perplexity | Live fetch on behalf of a person who asked | The user-initiated fetch fails, though user-triggered agents are the least predictable class here. |
| Googlebot | Google | Indexes for Google Search, including generated summaries | You leave Google entirely. Very few organisations should consider this. |
| Google-Extended | Google | A robots.txt token, not a crawler: controls Gemini training and grounding | Excluded from Gemini training and grounding. Ordinary Search indexing is unaffected. |
| Applebot | Apple | Indexes for Siri, Spotlight and Apple search features | You are absent from Apple's search surfaces. |
| Applebot-Extended | Apple | A token controlling use of your content in Apple's foundation models | Excluded from that training. Applebot indexing continues as normal. |
| Meta-ExternalAgent | Meta | Crawls for Meta's AI products and training | Removed from Meta AI's source pool. |
| Amazonbot | Amazon | Crawls to answer user queries in Amazon services including Alexa | Not available to Alexa and related answer surfaces. |
| Bytespider | ByteDance | Collects content for ByteDance model training | Training exclusion. Frequently blocked for crawl-rate reasons rather than policy ones. |

User agent names change and operators add new ones. Treat this as a starting list to verify against each vendor's published documentation and against your own edge logs, not as a permanent record.

Diagnosis

## The WAF problem, and how to check it tonight

Here is the finding that keeps repeating across audits. The content is good. The schema is present. The site is fast. And the reason it is never cited is that a web application firewall, a CDN bot manager or an over-broad security rule has been returning 403 to AI user agents for the past eighteen months, and nobody knew, because a blocked crawler does not appear in any report anyone reads. There is no message in Search Console. There is no drop in analytics, because the request never became a session. The signal is an absence, and absences do not raise alerts.

Bot management products ship with rule groups that treat unfamiliar automated agents as hostile by default. That is entirely reasonable engineering for scrapers and credential stuffers, and it catches OAI-SearchBot and Claude-SearchBot in the same net. Meanwhile the marketing team, who own the visibility problem, have no access to the console where the rule lives and would not think to look there.

### Three checks, about twenty minutes

First, read your robots.txt line by line, including any wildcard user agent block, and confirm you know why every disallow is there. Broad rules written years ago for scrapers routinely catch today's answer engines.

Second, get someone with access to the edge logs to filter the last thirty days by user agent for OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Googlebot and Applebot, and group by HTTP status. You are not looking at volume. You are looking for 403, 429, 503 and any redirect chain that ends somewhere unhelpful.

Third, request a handful of your important pages with one of those user agent strings set, from an address outside your network, and read the actual bytes that come back. A CAPTCHA page, a cookie wall that hides the body content, or an empty shell awaiting client-side rendering are all functionally identical to a block.

### Then write the policy on purpose

Once you can see what is happening, the decision is straightforward and it is a business decision, not a security one. Allow the search and user-triggered agents. Decide on training crawlers deliberately, knowing that the decision has no effect on citations either way. Put those choices in robots.txt, then make sure the WAF agrees with robots.txt, because in our experience the two documents disagree more often than they match. This overlaps heavily with the work described under website management, since the people who can change the firewall are rarely the people who noticed the problem.

Structured data

## Entity clarity and structured data

Schema.org markup is not decoration and it is not a ranking trick. It is the difference between a page that mentions a company and a page an engine can confidently attribute to one.

- Organization and ProfessionalService on every page. Name, legal name, registered address, telephone, logo, founding date and a stable @id. This is the node everything else attaches to, and it should be identical site-wide rather than re-invented per template.

- sameAs links to independent registries. Companies House, an ICO registration, a VAT identifier, LinkedIn, Crunchbase, industry bodies. Self-description is weak evidence. A link to a record somebody else maintains is how an engine gets from a string to a resolved entity it is willing to name.

- WebPage and BreadcrumbList on each page. Tells an engine what this page is about, where it sits in the site, and what it is a part of. Breadcrumbs are also the cheapest way to communicate information architecture to a machine.

- FAQPage where the questions are real. Question and answer pairs are extracted particularly readily, because the shape already matches how assistants structure output. Mark up genuine questions, not keyword strings with a question mark.

- Article and Person on anything authored. Named authors with credentials, an organisational affiliation and a reviewed date. Anonymous content carries less weight with engines that are visibly trying to work out who is speaking.

- Consistent naming everywhere. One legal name, one trading name, one category description, repeated across your site, your profiles and your listings. Entity resolution fails on drift, and drift is what happens when five teams write the boilerplate independently.

- Machine-readable alternates. A markdown version of each page and a clean HTML fallback. Content that only exists after client-side rendering is content some fetchers will never see.

- Facts stated in text, not only in images. Prices, specifications, opening hours, coverage areas and contact details written as text on the page. An engine cannot quote a diagram, and alt text is not a substitute for a sentence.

Editorial

## Writing pages an engine can quote

An assistant does not cite pages. It lifts passages. That single fact changes how a page should be built more than any other consideration in this discipline.

### Answer first, in forty to seventy words

Every page should open with a short, self-contained answer to the question the page exists for, written so that it survives being cut out and pasted into someone else's paragraph. No throat-clearing introduction, no company history, no scene setting. The first sentence states the answer; the next two or three qualify it. If that block cannot stand alone, it will not be lifted. There is one at the top of this page and you can judge for yourself whether it works.

### Headings shaped like questions

Assistants generate their own internal queries, and those queries are more literal and more conversational than the phrases humans type into a search box. Headings that match that shape are easier to retrieve against: what it costs, how long it takes, what happens if you block it, when it is the wrong tool. Mix them with statement headings so the page still reads like prose written by a person.

### Lists and comparison tables are quoted out of proportion

Structured formats are disproportionately represented in generated answers, because they are easy to parse, easy to attribute and easy to summarise without distortion. A comparison table with clear column headers does work that four paragraphs of careful prose will not. The crawler table further up this page exists for that reason as well as for yours.

### Definitions, plainly stated

Write the sentence that starts X is. Assistants answer a great many definitional questions, and a clean definition block, marked up and placed near the top of the relevant section, is among the most reliably extracted units on the web. Glossaries earn their keep here in a way they did not five years ago.

### Freshness, and why stale pages quietly lose

Engines favour recency when a question is time-sensitive, and a great deal more of the buying process is time-sensitive than people assume. Keep dateModified accurate in your schema and visible on the page, and make it honest: bumping the date without editing the content is a habit engines are getting better at seeing through. Real reviews on a real cycle, with the change reflected in the text, is the version that works. Pages that have not moved in three years lose to pages that were revised last month, even when the old page is better.

### llms.txt: worth doing, not a strategy

llms.txt is a proposed convention: a markdown file at the root of your domain listing your important pages with a line of context each, so a model reading it gets a curated map rather than a scrape of your navigation. It costs an afternoon. Adoption is growing, particularly among documentation-heavy sites. The honest position is that public evidence it directly lifts citations remains thin, and no major engine has committed to consuming it. Publish one, because it is cheap and genuinely useful to anyone reading your site with a machine. Do not let anyone charge you for it as though it were the programme. Ours is at /llms.txt.

Measurement

## Measuring something that has no rankings

You cannot get rank data out of an assistant, and any dashboard claiming to show your position in ChatGPT is showing you a modelled number.

| What you want to know | How it is actually measured | What it will not tell you |
|---|---|---|
| Are we named in AI answers? | A fixed prompt set of 50 to 150 buyer questions, run monthly against each engine, with cited sources recorded. | Anything about volume. This is a sample, not a census. |
| Are we gaining or losing ground? | Share of citations over time on the same prompt set, tracked per engine because engines diverge sharply. | Why a change happened. Correlation is all this method offers. |
| Which of our pages get used? | Recording the specific URL cited, not just the domain, so you learn which page shapes get quoted. | Whether a competing page was better or merely more reachable. |
| Who are we appearing beside? | Logging every other organisation named in each answer. Often the most useful output of the whole exercise. | Their strategy. Only that the engine considers them comparable to you. |
| Is any of it producing traffic? | Referral traffic from assistant domains in Google Analytics 4, segmented and watched as a trend. | Full attribution. Many assistant visits arrive with no referrer and land in direct. |
| Can engines reach us at all? | Edge log monitoring by AI user agent and status code, with an alert on any sustained non-200. | Whether being reachable led to being used. That is what the prompt set is for. |

Answers are non-deterministic, so the same prompt run twice can differ. Read the trend across a large prompt set over months. Do not read a single run, and do not let anyone report one to you as a result.

Method

## How a GEO engagement runs

Access first, entity second, content third. Doing them in any other order means spending on content that a crawler still cannot fetch.

- 01

### Baseline before touching anything

We agree the prompt set with you, run it across ChatGPT, Claude, Perplexity and Google's generated results, and record who is cited today. Without this, everything afterwards is anecdote.

- 02

### Audit machine access

Edge logs by user agent and status code, robots.txt line by line, WAF and bot management rules, rendering behaviour without JavaScript, and a live fetch test from outside your network. This phase finds the problem in most engagements.

- 03

### Fix access, and verify the fix

Search and user-triggered agents allowed, training policy set deliberately, robots.txt and firewall brought into agreement, then re-tested rather than assumed. Access work is the fastest-acting change available.

- 04

### Make the entity unambiguous

Organization, ProfessionalService, WebPage and BreadcrumbList implemented consistently, sameAs links pointed at independent registries, and naming reconciled across the site and your external profiles.

- 05

### Restructure the pages that matter

Answer-first blocks, question-shaped headings, comparison tables, definition blocks, honest dateModified, and internal linking that makes the relationship between pages legible to a machine.

- 06

### Re-run, report, repeat

The prompt set runs monthly. You get citation share by engine, the pages being quoted, the competitors sharing your answers, and a short list of what to change next. Programmes run in quarters, because that is the speed at which this actually moves.

Candour

## What we do not know yet

This discipline is roughly two years old. The vocabulary is not settled, the same practice is sold as generative engine optimisation, answer engine optimisation and AI search optimisation depending on who is writing, and the underlying systems change without release notes. Anyone speaking about it with total confidence is telling you something about themselves.

So here is what is well established. Crawler access is causal and provable: a blocked agent cannot cite you, and removing the block restores the possibility. Entity resolution is well understood from a decade of knowledge graph work and the mechanics carry over. Extractable structure genuinely helps, which is observable in the passages engines actually quote.

And here is what is not. The relative weight of on-page changes against everything an engine already believes is not publicly known. Whether llms.txt influences any major engine today is unproven. How much third-party mentions matter compared with your own site is contested. Nobody outside the labs can separate retrieval effects from training effects with any rigour, and the published studies in this area are small, quickly outdated and often produced by firms selling the remedy.

Which leads to the only promise worth making. We can make you reachable, resolvable and quotable, and we can measure whether citations move. We cannot guarantee an engine will recommend you, because no such guarantee exists to sell. If a supplier offers you one, ask them to put the mechanism in writing.

### Check our work on us

This site is built to the standard described on this page, which felt like the minimum before charging for it. View source and read the JSON-LD graph. Open /llms.txt. Note the extractable answer block at the top of every page, the question-shaped headings, the breadcrumbs, the markdown alternate on each URL and the reviewed date at the foot of the article. Then look at how our crawler policy separates training from search. Five minutes, and you will know whether we practise this or only sell it. The wider technical programme sits alongside ethical analytics and applied AI, because measurement and machine readability are the same problem viewed from two ends.

Questions

## Frequently asked questions

Terms used on this page are defined in the glossary, and there is longer writing on this subject in insights.

### What is generative engine optimisation?

Generative engine optimisation, or GEO, is the practice of making a website reachable, understandable and quotable by AI assistants, so that when someone asks ChatGPT, Claude, Perplexity, Gemini or Google's AI results about your category, your organisation is one of the sources named. It overlaps with search engine optimisation on fundamentals such as crawlability and clear information architecture, and departs from it on measurement, because there is no ranking to report. You are optimising for inclusion in an answer, not for a position in a list.

### Is GEO different from SEO, or just a rebrand?

Mostly the same plumbing, a genuinely different objective. Both need a site that can be fetched, parsed and understood. GEO adds three things SEO does not have to worry about. First, a separate set of crawlers with their own user agents and their own robots.txt tokens. Second, entity resolution, because an engine has to be confident that the company in the document is the same company as the one in the question. Third, passage-level extractability, because the unit that gets used is a paragraph, not a page. Anyone claiming the two disciplines are unrelated is overselling. Anyone claiming they are identical has not looked at their server logs.

### How do I know whether AI crawlers can reach my site?

Three checks, about twenty minutes. Read robots.txt line by line, including wildcard blocks. Filter your edge logs by AI user agent and group by HTTP status rather than by volume. Then request a page with one of those user agent strings set, from outside your network, and read what actually comes back. A 403, a challenge page or an empty shell awaiting client-side rendering all mean the same thing to an engine.

### Does blocking GPTBot hurt my visibility in ChatGPT?

No, and this is the distinction that trips most teams up. GPTBot collects data used in training. OAI-SearchBot builds the index behind ChatGPT's search, and ChatGPT-User is the fetch that happens live when someone's question sends the assistant to a page. Disallowing GPTBot is a defensible commercial or rights decision with no effect on whether you are cited. Disallowing OAI-SearchBot or ChatGPT-User removes you from the answers entirely. The same split applies to Anthropic's ClaudeBot against Claude-SearchBot and Claude-User, and to Google-Extended, which is a training control that does not affect Googlebot's indexing.

### What is llms.txt and do I need one?

It is a proposed convention: a markdown file at your domain root listing your important pages with a line of context each. Adoption is growing. Public evidence that it lifts citations on its own is still thin, and no major engine has committed to consuming it. Publish one, because it costs an afternoon and is useful to anyone reading your site with a machine. Do not buy it as a strategy.

### How do you measure whether GEO is working?

With a prompt set, because there is no rank to pull. Fifty to a hundred and fifty questions a real buyer would ask, run monthly against each engine, with every cited source recorded. The output is citation share by engine over time, the pages of yours being quoted, and the competitors sharing your answers. Alongside it, assistant referral traffic in Google Analytics 4. Neither number is a ranking.

### How long before we see anything change?

Access problems change fastest. If a firewall rule was blocking a search crawler and you remove it, inclusion can follow within days to a few weeks as that engine recrawls. Entity and structured data work shows up over a few weeks to a couple of months. Content restructuring is slower still, and pretrained model memory is slowest of all, because it only updates when the vendor trains or refreshes. Any timeline quoted with more precision than that is invented.

### Can you guarantee we will be recommended by ChatGPT or Claude?

No, and neither can anybody else. There is no paid inclusion, no submission endpoint and no ranking signal to buy. Answers vary between runs of the same prompt, they vary by region and by the account asking, and the systems change without notice. What can be done is remove the reasons an engine cannot use you, make it obvious what you do and who you are, and measure citation share honestly over time. A guarantee of placement in an AI answer is a sales tactic, not a capability.

### Should we block AI crawlers to protect our content?

It is a real decision with real trade-offs, and it should be made deliberately rather than by a default in a bot management console. Publishers with licensable archives often have good reason to disallow training crawlers. Almost nobody selling a professional service does. The pattern we keep finding is a site that intended to block training, wrote a broad rule, and removed itself from every AI answer surface at the same time. Separate the two questions, then write robots.txt and firewall policy that reflect the answers.

### Do you do this for your own site?

This site is built to the standard described on this page. The schema graph is emitted on every page and you can read it in view source. There is an llms.txt at the root and a markdown alternate for each page. Every page opens with a short extractable answer, the crawler policy separates training from search and user-triggered fetching, and we run our own prompt set monthly. Checking that claim takes about five minutes and we would rather you did.

Written and reviewed by the Green Arrow Consultancy team, led by Darren Tyler, founder and chief executive.

Green Arrow Consultancy Ltd, Cardiff, Wales. Company number 12491770. ICO registration ZA822868. Member of the International Association of Privacy Professionals. Last reviewed 04 September 2026.

Keep reading

## Related

### Ethical Analytics

Where assistant referral traffic shows up, and how to measure it without breaking consent.

Read this →

### Website Management

The firewall, CDN and hosting layer where most AI visibility problems are actually created.

Read this →

### AI Consulting

How retrieval systems choose sources, from the side of the people who build them.

Read this →

## Find out whether AI assistants can even reach you

We will run a prompt set for your category, check your edge logs and robots.txt against every major AI user agent, and tell you plainly what is blocking you. Most of the value is in the first week.

Request an AI visibility audit → 
 See all services
