Definitions for the words that turn up in AI proposals, privacy reviews, accessibility audits and AI search briefings. Written to be quoted, argued with and pasted into a policy. Where a term is contested or a vendor has stretched it, the entry says so rather than picking a side quietly.
This glossary defines the terms that come up when a business works on AI, AI search visibility, privacy, security, accessibility and web measurement. Each entry is written to be quoted: a plain definition first, then the part that changes what you do about it. Where a term is contested or a vendor has muddied it, the entry says so.
A glossary is a strange thing to spend time on until you sit in a meeting where two people use the word consent to mean two different things and nobody notices for twenty minutes. Most of the expensive mistakes we are called in to fix started as a vocabulary problem. A board approved an AI programme believing fine-tuning meant the model would learn their product catalogue. A marketing team blocked every AI crawler in one line of robots.txt and then asked why the brand had gone quiet inside ChatGPT. A supplier claimed level AA conformance and meant that a script had been installed.
So the entries here are written to a rule. The first sentence is a definition that stands on its own and can be quoted without the rest. What follows is the part that changes a decision: the common misreading, the thing a vendor will not volunteer, or the reason the term matters more than it sounds. Where the industry has not agreed on a meaning, the entry says that instead of pretending otherwise.
The six sections are the disciplines that now overlap in almost every project we take on. An AI assistant is a privacy question, a security question and an accessibility question before it is a model question. A consent banner is a performance problem as well as a legal one. AI search visibility depends on crawl access rules that live in the same file as your SEO configuration. Splitting these into separate glossaries would be tidier and less true to how the work runs. Our services are organised along the same seams.
Take the entries as a set of questions. If a proposal says the system will be trained on your data, ask whether they mean fine-tuning or retrieval, because the cost, the update cycle and the privacy analysis are different in each case. If an agency offers AI search optimisation, ask how they will measure citation share and how often they will sample it. If a platform promises compliance, ask which lawful basis applies to each purpose and where the record of processing lives.
We have been doing this since 2012, first as a web development and website management firm and then across privacy, consent, analytics, accessibility and applied AI for global consumer brands. The definitions come out of that work rather than out of a marketing brief. If you would rather see the ideas running than read about them, seven of our production systems are generalised into live demonstrations under Neuro.
The vocabulary of building with language models. If a supplier uses one of these words loosely, ask them which definition they mean.
Terms for the surface that replaced ten blue links for a growing share of questions. This is the fastest-moving section on the page.
Three different jobs, routinely governed by one careless line in robots.txt. Check the file and your firewall rules together, because a rule at the edge beats anything robots.txt says.
| Agent or token | Operated by | What it is for | What blocking it costs you |
|---|---|---|---|
| GPTBot | OpenAI | Collects public web content used in model training | Exclusion from future training corpora. No effect on ChatGPT search results. |
| OAI-SearchBot | OpenAI | Builds the index ChatGPT consults when it browses | You disappear from ChatGPT's live search results. |
| ChatGPT-User | OpenAI | Fetches one page because a user asked about it | A user who pastes your URL into a conversation gets an error. |
| ClaudeBot | Anthropic | Crawls public content for Anthropic | Exclusion from Anthropic's crawl. Check the separate agents before setting one rule. |
| PerplexityBot | Perplexity | Indexes pages for the Perplexity answer engine | Loss of one of the few surfaces that cites visibly and sends referral traffic. |
| Google-Extended | A robots.txt token, not a crawler; governs generative model use | No effect on Search ranking, crawling or AI Overviews. | |
| Googlebot | The single crawl behind Search and AI Overviews | You leave Google Search entirely. There is no separate AI Overviews opt-out. | |
| Bingbot | Microsoft | Powers Bing and Microsoft Copilot | You leave Bing and the Copilot surfaces built on it. |
| Applebot and Applebot-Extended | Apple | Siri and Spotlight indexing; the token governs generative training | Loss of Apple search surfaces, or of training use, depending which you set. |
| CCBot | Common Crawl | Builds a public archive that many training sets draw on | A broad indirect training opt-out. Past archives are not withdrawn. |
Verify current agent strings against each vendor's published documentation before you change a live file. We keep this checked as part of AI search optimisation engagements.
UK and EU data protection vocabulary, written for the people who have to operate it rather than for the people who cite it.
The failure modes that are specific to systems built on language models, plus the older controls that still do most of the work.
Standards, roles and artefacts that appear in accessibility requirements, procurement documents and audit reports.
The web performance, publishing and measurement terms that decide whether any of the above is actually reachable.
Eight distinctions that account for a disproportionate share of the confusion we are asked to unpick.
Broader questions about how we work are answered in the full FAQ.
People who have to make a decision and need the vocabulary to be precise about it: a marketing director asked to explain generative engine optimisation to a board, a general counsel reviewing an AI feature, a developer inheriting a consent platform, a procurement lead reading a supplier's accessibility report. The definitions assume intelligence and no prior specialism. Where a term is genuinely contested, the entry says so instead of picking a side quietly.
Because they are. Generative engine optimisation and answer engine optimisation are used interchangeably by some practitioners and distinguished by others. Hallucination is disliked by researchers who prefer confabulation. llms.txt is a proposal with no confirmed adoption by a major AI vendor. Presenting an unsettled term as settled makes a glossary less useful, not more, so we mark the disagreement and explain what turns on it.
SEO aims to place a page in a ranked list of links. Answer engine optimisation aims to have a passage lifted as the direct answer. Generative engine optimisation aims to be the source an AI assistant draws on and cites when it composes an answer from several places. The underlying technical work overlaps heavily, because all three need a crawlable, fast, well-structured site. What differs is the shape of the content and how success is measured. We cover this in AI search optimisation.
Not with one rule. Training crawlers, search crawlers and user-triggered fetchers are separate agents with separate consequences. Blocking GPTBot keeps content out of future OpenAI training sets. Blocking OAI-SearchBot removes you from ChatGPT's live results. Blocking ChatGPT-User means a user who pastes your URL into a conversation gets nothing back. Most organisations want some of those outcomes and not others, and the table on this page sets out which is which.
It costs an hour and it is harmless, so publish one if you like. Be honest about what it is: a community proposal for a Markdown index of your important pages, not a standard, and no major AI vendor has publicly confirmed that it reads one. Anyone selling llms.txt as the fix for AI visibility is selling the cheap part of a harder job. Crawl access, page structure and content quality are what actually decide whether you get cited.
No. The banner is an interface. Compliance depends on whether non-essential scripts are genuinely blocked until consent is given, whether refusing is as easy as accepting, whether the choice is recorded in a way you can produce as evidence, and whether the underlying processing has a lawful basis and a retention rule. Most of the failures we find are in the tag layer behind the banner, not the banner itself. See consent management.
No, and the entry in this glossary says so plainly. Overlays are widely criticised by disabled users and accessibility practitioners. They do not repair the underlying markup, they can interfere with a user's own assistive technology, and installing one has not protected organisations from complaints or litigation. The work that does help is unglamorous: semantic markup, keyboard operability, contrast, focus visibility and real testing with assistive technology. That is what our accessibility service does.
The AI search terms move fastest, because crawler names, robots.txt tokens and product surfaces change on vendor timetables rather than ours. We revise entries when the underlying fact changes rather than on a fixed schedule, and the review date appears at the foot of the page. If you find an entry that has gone stale, tell us and we will correct it.
Yes. Quote them in a policy, a training deck, an internal wiki or a supplier brief, with a link back to this page. We would rather the vocabulary spread than sit behind a form. If you need a definition adapted to your own regulatory context, that is a conversation rather than a copy and paste, and we are happy to have it.
The AI search terms move on vendor timetables rather than ours. If an entry is out of date, or you need one adapted for your own regulatory context, tell us and we will look at it.