# How to Get Cited by ChatGPT and AI Search | Green Arrow Consultancy

> A practical guide to getting cited by ChatGPT, Gemini and Perplexity: crawler access, structured data, answer-first writing and citation measurement.

Source: https://greenarrow.app/insights/get-cited-by-ai-search/
Last updated: 2026-09-04
Publisher: Green Arrow Consultancy Ltd, Cardiff, Wales, United Kingdom

---

Insights · Generative engine optimisation

# How to get your business cited by ChatGPT, Gemini and Perplexity

There is a lot of advice on this subject and very little of it separates the parts that are proven from the parts that are hopeful. This is the version we actually run for clients, in order, with the uncertainty left in.

Updated 4 September 2026 About 11 minutes By Darren Tyler 
 Talk to us about AI search → Read the rest of the insights 
 
 
 
 
 
 Quick answer

To be cited by an AI assistant you need three things: the assistant's crawler must be able to reach your page, the page must state its answer in a form that can be lifted out cleanly, and independent sources must corroborate that you exist. Most sites fail on the first, which silently cancels the other two. Test crawler access first, then structure, then off-site evidence.

## Key points

- A page can reach an AI answer by three different routes, and the work that helps on one route often does nothing on another.

- Crawler access is the only failure that silently cancels every other improvement you make.

- The unit of extraction is the passage, not the page, so each section has to make sense on its own.

- Comparison tables and ordered lists are quoted far more often than the same facts written as prose.

- Independent mentions off your own domain move the needle more than another page on it.

- Measure citation share against a fixed prompt set, not against a single impressive screenshot.

## On this page

- Three routes into an AI answer

- What you can influence on each route

- The ordered checklist

- Writing an answer a machine can lift

- Formats that get quoted more than prose

- Structured data and the resolvable entity

- Why independent mentions carry the weight

- How to measure citation share

- What is still unproven

- Frequently asked questions

Mechanics

## Three routes into an AI answer

Almost every argument about generative engine optimisation goes wrong because the people having it are describing different mechanisms without saying so. There are three ways a sentence from your website can end up in a machine-generated answer, and they behave nothing like each other.

### Route one: a search index the assistant queries

The assistant decides it needs current information, issues a query to a search index, retrieves a handful of documents and writes an answer grounded in them. This is how ChatGPT search, Perplexity, Claude's web search and Google's AI Overviews mostly work. The index is built by a crawler that visited you earlier, so there is a lag between publishing and being retrievable. This route is the one you can influence most directly, and it is where the majority of business-relevant citations come from today.

### Route two: a live fetch triggered by a user

Someone pastes your URL into a chat window, or asks a question that makes the assistant open a specific page right now. A different agent does this work, with a different user agent, in real time. There is no index and no lag. If your edge returns a 403 to that agent, the user sees the assistant fail to read a page they can open in their own browser, which is a bad look you will never be told about.

### Route three: pretrained memory

The model recalls something about you from its training data. This is the route people imagine when they talk about being known by AI, and it is the one you have the least control over. Training runs happen on the vendor's schedule, the corpus is largely fixed at that point, and nothing you publish this quarter affects a model that finished training last year. Memory also degrades into confident approximation, which is why assistants get company details subtly wrong even when the company is well known.

The practical consequence is that a strategy aimed at route three is mostly a strategy aimed at nothing you can measure. Aim at routes one and two, and treat any improvement in route three as a slow, unattributable bonus. Our AI search optimisation work is built around that ordering.

Scope of control

## What you can influence on each route

Three mechanisms, three different sets of levers. Confusing them is the most common reason a GEO programme produces no measurable change.

| Route | Who does the fetching | What you control | What you do not |
|---|---|---|---|
| Search index | An indexing crawler such as OAI-SearchBot, PerplexityBot or Googlebot | Access, structure, freshness, internal linking, entity clarity | Query rewriting, how many sources the model chooses to cite, ranking inside the index |
| Live user fetch | A user-triggered agent such as ChatGPT-User, Perplexity-User or Claude-User | Access, page speed, whether the answer is above the fold in the HTML, rendering without JavaScript | Which page the user or the assistant decides to open in the first place |
| Pretrained memory | A training crawler such as GPTBot, ClaudeBot or Meta-ExternalAgent | Whether you permit training use at all, and how clearly you are described where it crawled | The training schedule, the corpus mix, and whether the model recalls you correctly |

Route names are ours. The vendors do not use a shared vocabulary for these, which is part of the problem.

The work

## The ordered checklist

Nine items, in this order. The order matters more than the list, because the first two can silently invalidate the other seven.

- 01

### Confirm the AI crawlers can actually reach you

Request three or four of your own pages with each assistant's user agent and record the HTTP status. You are looking for 200. A 403, a 429, a 503 or an interstitial challenge page all mean you are invisible. Do this before anything else, because every later item on this list depends on it and none of them will tell you it has failed.

- 02

### Fix the firewall, not just robots.txt

Robots.txt is a request that well-behaved crawlers honour. A web application firewall is a wall. Allowlist the search and user-triggered agents at the edge by user agent and by the operators' published IP ranges, because user-agent strings are trivially spoofed and an allowlist based on them alone is a hole. This is involved enough that it has its own article.

- 03

### Publish structured data that resolves to a real entity

Organization, WebPage, Article and FAQPage markup with stable identifiers, and the same company details repeated on the pages and profiles a retrieval system can cross-check. The goal is not a rich result. The goal is that a machine reading your page can tell which real-world organisation it is about.

- 04

### Put a 40 to 70 word answer near the top of every page

One self-contained paragraph that answers the page's core question in plain language, with the direct answer in the first sentence. No preamble, no scene-setting, no dependency on the sentence before it. This is the block that gets lifted, and it is the single highest-yield piece of writing on the page.

- 05

### Shape headings as the questions people ask

What does this cost. How long does it take. What is the difference between these two things. Then answer immediately underneath, in the first sentence, before the qualification. Retrieval systems match on question phrasing far more than on keyword density.

- 06

### Publish comparison tables and ordered lists

Structured comparisons survive extraction intact in a way that a paragraph does not, and they are quoted disproportionately as a result. A table of five options with four attributes each is more useful to a retrieval system than two thousand words describing the same five options.

- 07

### Keep dateModified honest

Move the timestamp when the content moves and not otherwise. Bulk-refreshing dates across an estate to look current is easy to detect, easy to discount, and it destroys your own ability to tell which pages are genuinely stale.

- 08

### Earn mentions on domains you do not own

Trade publications, standards bodies, partner sites, directories, conference programmes, open datasets, review platforms, company registers. Assistants weight independent corroboration heavily because a claim on your own site is, from a retrieval system's point of view, just you saying it about yourself.

- 09

### Measure citation share across a fixed prompt set

Thirty to eighty questions your buyers would actually type, run across each engine on the same schedule, recorded consistently. Share over time is the metric. Any single answer is noise.

Craft

## Writing an answer a machine can lift

The unit of extraction is the passage, not the page. A retrieval system splits your content into chunks, embeds them, and pulls back the handful that match the query. Whatever surrounds a chunk is not guaranteed to come with it. That single fact should change how you write.

It means the anaphora habit that good prose relies on, referring back to the thing you just described, works against you. A paragraph that opens with "This means that" is a paragraph that means nothing when it arrives on its own. Name the subject again. Repeat the noun. It reads slightly heavier to a human and it reads correctly to a machine, and the tradeoff is worth making in the sections that carry the answers.

It also means the answer has to arrive before the reasoning. Most business writing is structured as context, then argument, then conclusion, because that is how the writer discovered it. Invert it. State the conclusion, then support it. If a reader stops after the first sentence of a section, they should still have the answer.

### The 40 to 70 word block

Every page on this site opens with one. It is short enough to be quoted whole, long enough to be substantive, self-contained enough to survive being ripped out of context, and written so the first sentence alone would satisfy someone in a hurry. Below 40 words it tends to lose the qualification that makes it accurate. Above about 70 it stops being quotable and starts being summarised, and a summary of your words is worth less to you than your words.

### Say the specific thing

Hedged writing is unquotable. "Solutions can vary depending on your requirements" contains no information and no retrieval system will ever surface it. "A consent management platform deployment typically takes three to six weeks, and the slow part is auditing the tags, not configuring the banner" is quotable because it commits to something. Committing is uncomfortable and it is most of the job. Our glossary exists for the same reason: definitions are the most extractable content format there is.

Format

## Formats that get quoted more than prose

This is not a ranking of importance. It is a list of the shapes that survive being cut out of your page and pasted into somebody else's answer.

- Comparison tables. Two or more named options against a consistent set of attributes. The most reliably quoted format we have measured, because the structure carries the meaning even after extraction.

- Numbered procedures. An ordered sequence with a verb at the start of each step. Assistants reproduce these almost verbatim when the question is procedural.

- Definition lists. Term, then a one or two sentence definition. This is why glossaries punch above their weight in AI answers relative to their traffic.

- Question-shaped headings with immediate answers. The heading matches the query, the first sentence beneath it is the answer, and the pair extracts cleanly as a unit.

- Explicit criteria and thresholds. Numbers, ranges, conditions and cut-offs. "Under 50 employees" is retrievable. "Smaller organisations" is not.

- Honest negative statements. What a thing does not do, who it is not for, when not to use it. Rare enough on commercial sites that it stands out to a system looking for balanced sources.

- Short FAQ answers. Sixty to a hundred and twenty words, one question, no cross-references. FAQ blocks are extracted as discrete units and behave well.

- Dated statements of fact. "As of September 2026" attached to anything time-sensitive. It lets a retrieval system judge freshness rather than guess at it.

Structure

## Structured data and the resolvable entity

Schema.org markup is widely oversold as an AI visibility technique. It is not a ranking factor in any assistant that has said anything public on the subject, and no amount of JSON-LD will get a thin page cited. What it does do is remove ambiguity, and ambiguity is a reason for a retrieval system to skip you.

The concept that matters is entity resolution. When a machine reads a page about Green Arrow Consultancy, it has to decide whether that is the same organisation as the one in the Companies House register under number 12491770, the one on the ICO's register under ZA822868, and the one with the LinkedIn company page. If those records agree with each other, the entity resolves and the system can attribute claims to a real organisation. If the name is spelled three ways across four sources and the address differs, it cannot, and a cautious system prefers a source it can pin down.

So the practical work is duller than the marketing suggests. Publish Organization markup once, sitewide, with a stable identifier. Include the registration numbers, the VAT number, the registered address and the sameAs links to the external records that corroborate them. Make sure the details on your contact page match the details in the markup, and that both match the register. Then add page-level markup that describes what the page actually is: Article on an article, FAQPage where there is a genuine FAQ, Service on a service page.

### Do not mark up what is not there

Marking up FAQ schema for questions that do not appear on the page, or Review schema for reviews that do not exist, is the fastest way to have your structured data ignored across the board. It is also, in the case of invented reviews, a consumer protection problem before it is a technical one. We have audited estates where an agency had bulk-applied review markup with no reviews behind it, and the remediation was more expensive than the original build.

Corroboration

## Why independent mentions carry the weight

Of everything on the checklist, the item that most reliably separates the businesses that get cited from the ones that do not is whether anyone else on the internet has written about them. This is uncomfortable because it is the item you control least and the one that takes longest.

The reasoning is straightforward once you think about the retrieval problem. An assistant assembling an answer about a category of supplier has to choose two or three names. Everything on a supplier's own site is self-description. A trade publication, a standards body, a partner's case study, a conference programme, a public register, an open dataset or a comparison site is not. Those sources are what turns a claim into a corroborated fact, and corroborated facts are what a system reaches for when it has to commit to naming someone.

In practice this means the highest-yield activity is often not writing another page. It is getting listed in the directories your sector actually uses, publishing data or research that other people cite, appearing on the register that governs your work, contributing to the specification everyone in your field reads, or being the named partner in someone else's write-up. None of it is fast and none of it is a growth hack.

It also means the old public relations discipline is more relevant to AI visibility than most of the technical tactics being sold under the GEO label. If you have a communications team, they are doing generative engine optimisation whether or not anyone has told them so.

Measurement

## How to measure citation share

Six numbers, recorded on a schedule. Run each prompt several times per engine, because these systems are non-deterministic and a single response tells you almost nothing.

**Fixed prompt set**

Thirty to eighty questions written the way a buyer would type them, agreed once and then left alone. Changing the prompts changes the numbers, so the set has to be stable for the trend to mean anything.

**Citation rate**

The share of runs in which your domain is one of the cited sources, per engine. This is the primary metric and the only one that maps directly to whether a buyer sees your name.

**Mention rate**

The share of runs in which your brand is named in the answer text even without a link. It moves before citation rate does and is worth tracking separately.

**Share of voice**

Your citation rate against the set of competitors that appear in the same answers. Useful because absolute rates drift as the engines change their retrieval behaviour, and a relative measure is more stable.

**Cited URL distribution**

Which of your pages get cited, and how concentrated that is. A single page carrying every citation is a fragile position and tells you where the gaps are.

**Answer accuracy**

Whether what the assistant says about you is correct. A high citation rate attached to a wrong description of your services is a problem, not a win, and nothing else in the measurement stack catches it.

Honesty

## What is still unproven

This field is eighteen months old as a commercial discipline and a good deal of what is confidently asserted in it has never been tested. Here is where we think the uncertainty actually sits.

Unproven: that llms.txt lifts citations. The file is cheap to publish and there is no evidence that any major assistant treats it as a retrieval or ranking input. We publish one and we do not count it as a strategy. There is a longer piece on that.

Unproven: that there is a stable ranking algorithm to reverse-engineer. Retrieval behaviour changes without announcement, models are swapped, and the same prompt returns different sources on different days. Anyone selling a fixed set of ranking factors for AI answers is selling a snapshot of noise.

Contested: how much organic ranking carries over. Published overlap figures between organic search results and AI citations vary enormously by study and by engine, and the honest summary is that Google's AI surfaces lean heavily on what already ranks while ChatGPT's citation set overlaps considerably less. We go into that in detail elsewhere.

Reasonably well supported: access, structure and corroboration. Crawler access is not a theory, it is a status code. Extractability is observable in what gets quoted. The weighting of independent sources is visible in the citation sets themselves. These three are where we would put the budget.

If you want the work done rather than described, that is what our AI search optimisation service is, and the broader AI consulting practice is where it sits when the answer turns out to be a system rather than a set of pages.

Questions

## Frequently asked questions

Questions we get asked in the first meeting, answered without the sales layer.

### How long does it take to start appearing in AI answers?

If the problem was crawler access, changes can show up within days on the live-fetch surfaces and within a few weeks on the indexed ones, because the index has to be rebuilt before you exist inside it. If the problem is that nobody outside your own domain has written about you, that is a six to twelve month job and no amount of on-page work shortens it. Pretrained memory is slower still and effectively outside your control, since it only changes when a model is retrained.

### Is there a way to pay to be cited?

Not in the sense of buying a citation. Advertising is arriving in some assistants, but sponsored placement is labelled separately from the cited sources and does not make you a source. Anyone selling guaranteed placement inside an organic AI answer is selling something they do not control.

### Do I need to write differently for AI than for people?

You need to write more decisively, which usually improves the page for people too. The specific changes are answering in the first paragraph rather than the fifth, making each section self-contained so a passage still makes sense when it is lifted out, and stating claims plainly instead of hedging them into vapour. None of that requires a separate version of the page and you should not build one.

### Does schema markup make an assistant cite me?

No single markup change causes a citation. Structured data helps in a narrower way: it makes the entity on the page unambiguous, so a retrieval system can tell that the Green Arrow Consultancy on this page is the same company as the one in the Companies House record and the LinkedIn profile. Ambiguity is a reason to skip a source. Removing ambiguity is worth doing even though it is not a ranking lever.

### Which engine should we optimise for first?

Start with whichever one your buyers actually use, which you can find by asking them rather than by reading a market share chart. In practice the underlying work is shared: crawler access, clean structure and independent corroboration help everywhere. Where the engines diverge is in what they retrieve from, which is why measurement needs to be per engine even when the work is not.

### Can we get removed from an AI answer that says something wrong about us?

Sometimes. If the wrong claim comes from a page you control, fix the page and the live-fetch and search surfaces will follow. If it comes from a third-party page, the correction has to happen there. If it comes from pretrained memory, there is no removal process worth relying on and the practical remedy is to publish correct, well-structured material that retrieval will prefer over the model's recollection.

### Does publishing more content increase citations?

Only if the extra content answers questions that nobody has answered well. Volume on its own dilutes. We have seen estates with hundreds of near-duplicate pages get cited less than a competitor with thirty specific ones, because near-duplicates compete with each other for the same retrieval slot and none of them wins clearly.

### How do we measure this without a rank tracker?

Build a fixed prompt set of thirty to eighty questions your buyers would ask, run them across the engines on the same schedule, and record whether you were named, which URL was cited, and who else appeared. Track the share over time rather than any single run, because these systems are non-deterministic and one answer proves nothing.

### What is the single highest-value change for most sites?

Checking that the crawlers can reach you. It is the only item on the list that can silently zero out everything else, it takes an afternoon to test, and a surprising share of well-run corporate estates fail it without anyone knowing.

Written and reviewed by the Green Arrow Consultancy team, led by Darren Tyler, founder and chief executive.

Green Arrow Consultancy Ltd, Cardiff, Wales. Company number 12491770. ICO registration ZA822868. Member of the International Association of Privacy Professionals. Last reviewed 04 September 2026.

Keep reading

## Related

### Your firewall is probably blocking the AI crawlers you want

Step one of the checklist, in full, with the commands to test it yourself.

Read this →

### GEO and SEO are not the same job

What carries over from search optimisation and what genuinely does not.

Read this →

### AI Search Optimisation

The service that runs this checklist end to end, with measurement attached.

Read this →

## Find out whether you are reachable at all

We will run the crawler access test against your estate, check what an assistant can actually extract from your key pages, and tell you which of the nine items are already fine.

Request an AI visibility check → 
 See the service
