Reference · Glossary

An AI and privacy glossary for people who have to decide something

Definitions for the words that turn up in AI proposals, privacy reviews, accessibility audits and AI search briefings. Written to be quoted, argued with and pasted into a policy. Where a term is contested or a vendor has stretched it, the entry says so rather than picking a side quietly.

94 termsSix sectionsReviewed 4 September 2026
Quick answer

This glossary defines the terms that come up when a business works on AI, AI search visibility, privacy, security, accessibility and web measurement. Each entry is written to be quoted: a plain definition first, then the part that changes what you do about it. Where a term is contested or a vendor has muddied it, the entry says so.

Orientation

How to read this glossary

A glossary is a strange thing to spend time on until you sit in a meeting where two people use the word consent to mean two different things and nobody notices for twenty minutes. Most of the expensive mistakes we are called in to fix started as a vocabulary problem. A board approved an AI programme believing fine-tuning meant the model would learn their product catalogue. A marketing team blocked every AI crawler in one line of robots.txt and then asked why the brand had gone quiet inside ChatGPT. A supplier claimed level AA conformance and meant that a script had been installed.

So the entries here are written to a rule. The first sentence is a definition that stands on its own and can be quoted without the rest. What follows is the part that changes a decision: the common misreading, the thing a vendor will not volunteer, or the reason the term matters more than it sounds. Where the industry has not agreed on a meaning, the entry says that instead of pretending otherwise.

What is in scope, and why these six areas

The six sections are the disciplines that now overlap in almost every project we take on. An AI assistant is a privacy question, a security question and an accessibility question before it is a model question. A consent banner is a performance problem as well as a legal one. AI search visibility depends on crawl access rules that live in the same file as your SEO configuration. Splitting these into separate glossaries would be tidier and less true to how the work runs. Our services are organised along the same seams.

How to use it with a supplier

Take the entries as a set of questions. If a proposal says the system will be trained on your data, ask whether they mean fine-tuning or retrieval, because the cost, the update cycle and the privacy analysis are different in each case. If an agency offers AI search optimisation, ask how they will measure citation share and how often they will sample it. If a platform promises compliance, ask which lawful basis applies to each purpose and where the record of processing lives.

We have been doing this since 2012, first as a web development and website management firm and then across privacy, consent, analytics, accessibility and applied AI for global consumer brands. The definitions come out of that work rather than out of a marketing brief. If you would rather see the ideas running than read about them, seven of our production systems are generalised into live demonstrations under Neuro.

Foundations

AI and machine learning

The vocabulary of building with language models. If a supplier uses one of these words loosely, ask them which definition they mean.

Large language model (LLM)
A large language model is a model trained on very large volumes of text to predict the next token in a sequence, which in practice produces a system that can write, summarise, classify and reason over language. Claude, GPT and Gemini are large language models. The label says nothing about accuracy. An LLM produces plausible text, so anything it asserts has to be grounded in a source or checked by a person before a business relies on it.
Retrieval-augmented generation (RAG)
Retrieval-augmented generation is an architecture in which the system searches your own content for relevant passages and then asks the model to answer using only those passages. It keeps answers current, because updating a document updates the answer, and it makes citation possible. Most business questions that look like they need a custom model actually need retrieval. We cover the choice in AI consulting.
Embedding
An embedding is a numeric representation of text, an image or audio as a list of several hundred or several thousand numbers, arranged so that things with similar meaning sit close together. Embeddings are what make semantic search work: a question and a passage do not need to share words, only meaning. They are produced by a separate, smaller model from the one that writes the answer.
Vector database
A vector database is a store built to hold embeddings and return the nearest matches to a query quickly. Pinecone, Weaviate, Qdrant, pgvector and the vector features inside Elasticsearch and PostgreSQL all do this. A vector database is a component, not an architecture. Retrieval quality usually depends far more on how documents were split and labelled than on which store holds them.
Context window
The context window is the amount of text a model can consider at once, counted in tokens and covering the system prompt, the conversation so far, any retrieved documents and the answer being written. Larger windows allow more evidence per question, cost more, and can dilute attention. A long context window is not a substitute for retrieval, because you still have to choose what goes into it.
Chunking
Chunking is the way a document is split before it is embedded and indexed. Chunk too large and retrieval returns noise around the useful sentence. Chunk too small and the passage loses the context that made it meaningful. Splitting on headings, and keeping the section title attached to every chunk, usually beats splitting on a fixed character count. Chunking decisions affect answer quality more than the choice of vector store does.
Reranking
Reranking is a second pass over retrieval results that uses a slower, more accurate model to reorder the shortlist before it reaches the answering model. Vector search is good at recall and mediocre at precision, so reranking often produces a larger quality gain than any prompt change. It adds latency and cost, which is why it is applied to a shortlist rather than to the whole corpus.
Prompt engineering
Prompt engineering is the practice of writing the instructions, examples and formatting rules that shape a model's output. It is real engineering when prompts are versioned, tested against an evaluation set and reviewed on change. It is folklore when it is a paragraph somebody tweaked until a demo worked. Prompts belong under the same change control as code, because they are code.
System prompt
A system prompt is the standing instruction given to a model before any user message, setting its role, scope, tone, refusal rules and output format. It is where most guardrail behaviour lives in a business assistant. A system prompt is not a security boundary: a determined user or a poisoned document can try to argue around it, which is why prompt injection testing matters.
Fine-tuning
Fine-tuning is further training of an existing model on your own examples so that it adopts a format, a style or a narrow classification behaviour. It changes how the model writes, not what it knows about today, so it is a poor way to keep facts current. Fine-tuning is proposed far more often than it is needed, and retrieval solves most of the problems it gets sold for.
Distillation
Distillation is training a smaller model to reproduce the behaviour of a larger one, giving most of the quality at a fraction of the cost and latency. It is widely used to make production systems affordable at volume. Terms of service for hosted models often restrict using their outputs to train competing models, so distillation is a licensing question as much as a technical one.
Open-weight model
An open-weight model is one whose trained parameters are published, so it can be downloaded and run on infrastructure you control. The Llama, Mistral and Qwen families are the common examples. Open weights are not the same as open source, because training data and training code are usually not released. They matter when data residency, unit cost at very high volume or offline operation rules out a hosted API.
Inference
Inference is a single run of a trained model to produce an output, as opposed to training. It is what you pay for per request in production, measured in tokens in and tokens out, and it is where latency lives. Most cost control in an AI system is inference engineering: routing simple work to smaller models, caching repeated work, and not retrieving more context than the answer needs.
Token
A token is the unit a model reads and writes, roughly three quarters of an English word, produced by splitting text into common fragments. Pricing, context limits and rate limits are all counted in tokens. Non-English text and source code often use more tokens per character, which is why multilingual and technical workloads cost more than a word count would suggest.
Temperature
Temperature is a setting that controls how much randomness is allowed when a model picks each next token. Low values produce repeatable, conservative output and suit extraction, classification and factual answering. Higher values produce more variation and suit drafting and ideation. Temperature does not make a model more accurate at any setting, it only changes how much the output varies between runs.
Multimodal model
A multimodal model accepts more than one kind of input, typically text plus images, and sometimes audio, video or documents as rendered pages. Practical uses include reading a photograph of a part label, extracting a table from a scanned invoice, or drafting alt text for an image. Accuracy varies sharply by task, so multimodal features need their own evaluation set rather than inheriting the text one.
Hallucination
A hallucination is a confident, fluent, wrong answer produced by a model completing a pattern rather than consulting a source. The model is not lying, and fluency is not evidence. The practical controls are grounding answers in retrieved sources, requiring citations, and testing that the system will say it does not know. Some researchers dislike the word and prefer confabulation, on the grounds that hallucination implies perception the model does not have.
Grounding
Grounding is constraining a model to answer from supplied evidence rather than from what it absorbed during training. In practice that means retrieving passages, placing them in the context window, and instructing the model to use nothing else. Grounding is the single largest quality lever in an enterprise assistant, and it is what makes an answer checkable by a human being in one click.
Citation
A citation is a reference in an AI answer back to the source it came from, ideally a specific document and section rather than a whole website. For a reader it is the difference between an assertion and a checkable claim. For a business publishing content, being the cited source inside an assistant's answer is the visibility that now stands in for a ranked link. See AI search optimisation.
AI agent
An AI agent is a system that plans a sequence of steps and takes actions towards a goal, rather than returning a single answer. An agent might read a ticket, look up an account, draft a reply and schedule a follow-up. Agency raises the stakes, because every action needs a scope, a log and a way to reverse it. We build these under AI agents and automation.
Tool use (function calling)
Tool use is the mechanism by which a model calls a defined function, such as a search, a database query or an API write, and receives the result back so it can continue. It is what turns a chat interface into a system that can actually do something. It is also where most agent security work concentrates, because a tool is a real permission wearing a friendly name.
Structured output
Structured output means forcing a model to return data in a fixed shape, usually JSON matching a schema, rather than free prose. It is what makes model output safe to pass into another system, because the receiving code can validate it before acting. Where a workflow depends on a decision, structured output plus validation is far more reliable than parsing sentences with regular expressions.
Evaluation set
An evaluation set is a fixed collection of real questions with agreed correct answers, used to score a system on every change. It is the deliverable nobody asks for and the one that decides whether you can safely upgrade a model. Without it, the observation that the assistant feels worse this month is an argument. With it, the observation is a number. The people who own the subject should write it.
Guardrail
A guardrail is a control that constrains what a system will accept or produce: topic limits, refusal rules, output validation, content filters, rate limits and permission checks. Guardrails belong at several layers rather than only in the prompt, because a prompt is text and text can be argued with. Test them adversarially. See AI security.
Human in the loop
Human in the loop describes a design in which a person reviews, approves or corrects the system's output before it takes effect. It is a governance requirement for higher-risk uses under the EU AI Act and sound engineering everywhere else. Meaningful oversight requires the reviewer to have the time, the information and the authority to say no, which is where most implementations quietly fail.
Model card
A model card is a short structured document describing a model or an AI system: what it is for, what it was trained or grounded on, its known limitations, its evaluation results, and the uses it should not be put to. It began as a publication practice for released models and is now a practical internal artefact that auditors ask for. It is a standard output of AI governance work.
Reference table

Which AI crawler does what

Three different jobs, routinely governed by one careless line in robots.txt. Check the file and your firewall rules together, because a rule at the edge beats anything robots.txt says.

Agent or tokenOperated byWhat it is forWhat blocking it costs you
GPTBotOpenAICollects public web content used in model trainingExclusion from future training corpora. No effect on ChatGPT search results.
OAI-SearchBotOpenAIBuilds the index ChatGPT consults when it browsesYou disappear from ChatGPT's live search results.
ChatGPT-UserOpenAIFetches one page because a user asked about itA user who pastes your URL into a conversation gets an error.
ClaudeBotAnthropicCrawls public content for AnthropicExclusion from Anthropic's crawl. Check the separate agents before setting one rule.
PerplexityBotPerplexityIndexes pages for the Perplexity answer engineLoss of one of the few surfaces that cites visibly and sends referral traffic.
Google-ExtendedGoogleA robots.txt token, not a crawler; governs generative model useNo effect on Search ranking, crawling or AI Overviews.
GooglebotGoogleThe single crawl behind Search and AI OverviewsYou leave Google Search entirely. There is no separate AI Overviews opt-out.
BingbotMicrosoftPowers Bing and Microsoft CopilotYou leave Bing and the Copilot surfaces built on it.
Applebot and Applebot-ExtendedAppleSiri and Spotlight indexing; the token governs generative trainingLoss of Apple search surfaces, or of training use, depending which you set.
CCBotCommon CrawlBuilds a public archive that many training sets draw onA broad indirect training opt-out. Past archives are not withdrawn.

Verify current agent strings against each vendor's published documentation before you change a live file. We keep this checked as part of AI search optimisation engagements.

Data protection

Privacy and data protection

UK and EU data protection vocabulary, written for the people who have to operate it rather than for the people who cite it.

Personal data
Personal data is any information relating to an identified or identifiable living person under UK GDPR and the EU GDPR. It is broader than most teams assume: an online identifier, a cookie ID, an IP address or a device fingerprint can each be personal data with no name attached. If the data can be linked back to a person by any means reasonably likely to be used, treat it as personal data and stop arguing about it.
Special category data
Special category data is data revealing racial or ethnic origin, political opinions, religious or philosophical beliefs, trade union membership, genetic or biometric data used for identification, health data, or data concerning sex life or sexual orientation. It needs a lawful basis and a separate Article 9 condition. Inferred categories count, which matters the moment a model starts classifying free text about people.
Controller
A controller is the organisation that decides why and how personal data is processed. Controllers carry the primary obligations: lawful basis, transparency, responding to individuals, and accountability for what their processors do. Most brands are the controller for their own website data even when an agency operates the site day to day. See privacy consulting.
Processor
A processor is an organisation that processes personal data on a controller's documented instructions, such as a hosting provider, an analytics vendor or an agency running a website. Processors need a written contract containing specified terms and cannot decide new purposes for the data. The labels follow the facts of who decides what, not the words the contract happens to use.
Lawful basis
A lawful basis is one of the six legal grounds under UK and EU GDPR that must exist before personal data is processed: consent, contract, legal obligation, vital interests, public task, or legitimate interests. You choose one per purpose, before processing starts, and you record the choice. Switching basis later because the first one became inconvenient is not permitted.
Legitimate interests
Legitimate interests is a lawful basis available where processing is necessary for interests pursued by the controller or a third party and those interests are not overridden by the individual's rights. It requires a documented balancing assessment. It is not a catch-all, and it cannot be used to place non-essential cookies or similar tracking technologies, which need consent under the separate ePrivacy rules.
Data protection impact assessment (DPIA)
A data protection impact assessment is a structured analysis of a processing activity likely to result in high risk to individuals, carried out before that processing begins. Large-scale profiling, systematic monitoring and the use of innovative technology are common triggers, which is why most new AI features need one. It documents necessity, proportionality, the risks and the measures that reduce them.
Record of processing (ROPA)
A record of processing is the inventory required by Article 30 of the GDPR, describing what personal data an organisation processes, why, on what basis, who it is shared with, where it goes and how long it is kept. Regulators ask for it first. Keeping it accurate is unglamorous, and it is the fastest way to discover that a system nobody remembers is still collecting data every day.
Data subject request
A data subject request is a request from an individual exercising a right: access to their data, correction, erasure, restriction, portability, or objection to processing. Deadlines are short, generally one month with limited extension. The hard part is rarely the legal analysis. It is knowing every system that holds the person's data, which is precisely what the record of processing exists to tell you.
International transfer
An international transfer is a movement of personal data outside the UK or the EEA, which requires a transfer mechanism such as an adequacy decision, standard contractual clauses or the UK addendum, plus a transfer risk assessment. AI features routinely create transfers nobody planned, because a model API may be served from another jurisdiction. Ask where inference happens before you ship, not after.
First-party data
First-party data is data an organisation collects directly from its own customers and visitors, in its own systems, with a direct relationship behind it. It is more durable than third-party data as browsers restrict cross-site tracking. It is not automatically compliant: consent, transparency, purpose limitation and retention rules apply to it exactly as they apply to everything else.
Data minimisation
Data minimisation is the principle that you collect only the personal data you actually need for a stated purpose, and no more. It is the cheapest privacy control in existence, because data you never collected cannot be breached, mishandled or requested back. In AI work it usually means stripping identifiers before content reaches a model, and asking why a form field exists at all.
Retention
Retention is how long data is kept before it is deleted or genuinely anonymised, set per purpose and written down. Indefinite retention is not a policy, it is the absence of one. Retention is where AI systems quietly go wrong, because prompts, retrieved passages and logs are a new copy of personal data with a lifecycle of their own that nobody has assigned an owner to.
Privacy by design
Privacy by design means building data protection into the design of a system rather than reviewing it at the end. In UK and EU law it is an obligation rather than a philosophy. In practice it means the data flow, the lawful basis, the retention rule and the deletion path are settled in the first design meeting, while changing them is still cheap. It is the working method behind our web development.
Risk

AI security

The failure modes that are specific to systems built on language models, plus the older controls that still do most of the work.

Prompt injection
Prompt injection is an attack in which text supplied to a model is crafted to override the instructions it was given, causing it to ignore its rules, reveal its system prompt or misuse a tool. It is not a bug awaiting a patch. It is a consequence of instructions and data sharing one channel. Mitigation is architectural: least privilege, output validation, and never treating model output as a command. We test for it in AI security work.
Indirect prompt injection
Indirect prompt injection is prompt injection delivered through content the system retrieves rather than through the user's own message: a poisoned web page, a PDF, a calendar invitation, a support ticket or an email. It is the more dangerous form, because the attacker never needs access to your interface. Any agent that reads untrusted content and can also act should be assumed reachable this way.
Jailbreak
A jailbreak is a prompt designed to make a model produce output its provider's policies or your own rules forbid, usually through role-play, hypothetical framing or layered instructions. Jailbreaks target the model's alignment, while prompt injection targets your application's instructions. Business systems should be tested for both, and should fail safe by refusing rather than by improvising something plausible.
Data exfiltration
Data exfiltration is getting data out of a system that was supposed to keep it in. In AI systems the routes are novel: an assistant answering from documents the asker cannot open, a tool call that writes to an external service, an image URL that carries data in its query string, or logs that quietly retain prompt content. Mapping source-system permissions into retrieval is the primary defence.
Excessive agency
Excessive agency is giving an AI system more capability, permission or autonomy than the task requires, so that a bad output becomes a bad action. An assistant able to delete records because deletion happened to sit in the same API as reading is the classic example. The fix is boring and effective: scoped credentials, an explicit allowlist of actions, confirmation for anything destructive, and complete logging.
Least privilege
Least privilege is granting every component, credential and agent only the access it needs to do its job and nothing more. It is decades old and it still limits the blast radius of an AI incident more than any other single control. In retrieval systems it means the index respects source-system permissions, so a question can only ever be answered from documents the asker is entitled to read.
Insecure output handling
Insecure output handling is passing model output into another system without validating it: rendering it as HTML, executing it as code or as a database query, or acting on it as a command. Model output is untrusted input, exactly like a form field. Treating it as trusted is how a language model becomes a cross-site scripting or injection vector inside an otherwise conventional application.
Red teaming
Red teaming is deliberately attacking your own system to find out how it fails: injection attempts, jailbreaks, requests for out-of-scope advice, attempts to extract system prompts or training data, and probing for actions the agent should refuse. It should be repeatable and scored rather than a one-off afternoon, and the results belong in the evidence pack you show an auditor.
OWASP Top 10 for LLM Applications
The OWASP Top 10 for LLM Applications is a widely used, freely published list of the most significant security risks in applications built on large language models, covering prompt injection, insecure output handling, training data poisoning, supply chain risk, excessive agency and more. It is the most practical starting checklist for a security review of an AI feature, and it is revised as the field moves.
Supply chain risk
Supply chain risk is risk inherited from the models, datasets, libraries, plugins and hosted APIs an AI system depends on. A downloaded model, a package from a public registry or a third-party tool integration can each carry a compromise. The practical controls are pinned versions, provenance checks, a bill of materials for the AI stack, and knowing whose terms govern the data you send.
Inclusion

Accessibility

Standards, roles and artefacts that appear in accessibility requirements, procurement documents and audit reports.

WCAG
The Web Content Accessibility Guidelines are published by the World Wide Web Consortium and are the reference standard for accessible digital content worldwide. WCAG 2.1 is the version most legislation still names, and WCAG 2.2 adds criteria covering focus appearance, dragging alternatives, target size, consistent help and accessible authentication. Its success criteria are testable statements, which is what makes them usable as contract terms.
Conformance level AA
Level AA is the middle of WCAG's three conformance levels and the one nearly all legislation and procurement asks for. Level A is a minimum that leaves real barriers in place, and level AAA is not expected across a whole site. Claiming AA means every applicable success criterion at A and AA is met on the pages in scope, and that claim should be evidenced rather than asserted.
EN 301 549
EN 301 549 is the European accessibility standard for information and communication technology procurement and products. It incorporates WCAG at level AA for web content and adds requirements for software, documents, hardware and support services. It is the standard referenced by EU public sector rules and by the European Accessibility Act, so it is usually the form of the requirement that turns up in a contract.
European Accessibility Act
The European Accessibility Act is EU Directive 2019/882, which extends accessibility requirements to private sector products and services including e-commerce, banking, e-books, transport and telecoms, with obligations applying from 28 June 2025. It reaches organisations outside the EU that sell to EU consumers, which is why a number of UK and US brands ran remediation programmes. See accessibility.
Screen reader
A screen reader is software that converts on-screen content into synthesised speech or refreshable braille and provides keyboard commands for moving through a page by heading, landmark, link or form field. JAWS, NVDA, VoiceOver and TalkBack are the common ones. Screen reader users navigate by structure, which makes headings, labels and landmarks functional controls rather than styling choices.
Alt text
Alt text is the text alternative for an image, announced in place of the image and displayed when the image fails to load. Good alt text conveys the purpose of the image in its context rather than describing every pixel. Decorative images should carry an empty alt attribute so they are skipped entirely. Charts and diagrams usually need a longer description nearby as well.
Reading order
Reading order is the sequence in which content is exposed to assistive technology and to keyboard navigation, which follows the underlying document order rather than the visual layout. CSS can move things visually and leave the reading order wrong, and that mismatch is one of the most common audit findings. It matters in documents as much as on the web, particularly in PowerPoint and PDF.
Focus indicator
A focus indicator is the visible outline showing which element the keyboard is currently on. Removing it, which a stylesheet can do in a single line, makes a site unusable for anyone who is not using a mouse. WCAG 2.2 raised the bar here with criteria on focus appearance and on focus not being obscured by sticky headers, chat launchers or consent banners.
Accessibility overlay
An accessibility overlay is a third-party script that claims to make a site conformant automatically. Overlays are widely criticised by disabled users and accessibility practitioners. They do not repair the underlying code, they can interfere with a person's own assistive technology, and installing one has not protected organisations from complaints or litigation. We do not deploy them, and we say so when asked.
VPAT and Accessibility Conformance Report
A Voluntary Product Accessibility Template is a standard document in which a supplier reports how a product meets accessibility standards such as WCAG, EN 301 549 or Section 508, criterion by criterion. The completed document is properly called an Accessibility Conformance Report. Buyers increasingly ask for one, and an honest report with noted gaps is worth considerably more than an optimistic one.
Accessibility statement
An accessibility statement is a published page describing how accessible a service is, which standard it aims at, what is known to be broken, what is planned, and how to report a barrier. It is a legal requirement for UK and EU public sector bodies and good practice for everyone else. A statement claiming full conformance with no known issues is usually evidence that nobody tested.
Delivery

Web and measurement

The web performance, publishing and measurement terms that decide whether any of the above is actually reachable.

Core Web Vitals
Core Web Vitals are Google's field measurements for page experience: Largest Contentful Paint for loading, Interaction to Next Paint for responsiveness, and Cumulative Layout Shift for visual stability. They are measured on real users through the Chrome User Experience Report rather than in a lab, so a good Lighthouse score does not guarantee good vitals. They are a ranking input and a useful proxy for whether a site feels broken.
Largest Contentful Paint (LCP)
Largest Contentful Paint measures the time until the largest visible element in the viewport, usually a hero image or a heading block, has finished rendering. The commonly used threshold is 2.5 seconds at the 75th percentile of real users. It is normally fixed by server response time, image weight and format, and render-blocking resources, in roughly that order.
Interaction to Next Paint (INP)
Interaction to Next Paint measures how long a page takes to respond visibly to a user's interaction, assessed across a whole visit and reported at the worst end rather than the first click. It replaced First Input Delay in March 2024 because responsiveness later in a session was the part users actually felt. Heavy third-party scripts, including some consent and tag manager setups, are a frequent cause of poor scores.
Cumulative Layout Shift (CLS)
Cumulative Layout Shift measures how much content moves around unexpectedly while a page loads, scored upward from zero with 0.1 or below as the target. The usual causes are images without declared dimensions, fonts that swap in late, and banners injected above existing content. Consent banners offend regularly, which is an implementation problem to solve rather than a reason to remove the banner.
Canonical tag
A canonical tag is a link element telling search engines which URL is the preferred version of a page when the same or very similar content is reachable at several addresses. It consolidates signals rather than blocking access. It is a strong hint rather than a directive, so engines can and sometimes do choose differently, particularly when the canonical points somewhere that does not match the content.
robots.txt
The robots.txt file is a plain-text file at the root of a domain telling compliant crawlers which paths they may request, with rules grouped by user agent. It controls crawling rather than indexing, and it is advisory, so well-behaved bots honour it and others ignore it. It is the main lever for AI crawler policy, which is why so many sites have accidentally excluded the assistants they wanted to appear in.
XML sitemap
An XML sitemap is a file listing the URLs a site wants crawled, with optional last-modified dates, referenced from robots.txt and submitted in a search console. It helps discovery on large or poorly linked sites. It does not improve ranking, and a sitemap full of URLs that redirect, error or are blocked does active harm by wasting the crawl budget a search engine allocates to you.
Headless CMS
A headless content management system stores and edits content but does not render the site, exposing content through an API for a separate front end to consume. It suits multi-channel publishing and gives developers freedom. It also moves responsibility for accessibility, structured data and rendering onto the front-end team, which is a trade rather than a free upgrade. We discuss it in website management.
Server-side tagging
Server-side tagging means running tag logic on a server you control, so the browser sends one request to your own endpoint and the server distributes data onward. It improves page performance, reduces exposure to browser tracking restrictions, and gives you one place to enforce consent and strip fields. It does not make tracking lawful by itself, and it is often mis-sold as a way around consent. See ethical analytics.
Google Analytics 4
Google Analytics 4 is Google's current analytics product, event-based rather than session-based, with consent mode integration, modelled data where consent is absent, and reporting built on explorations rather than fixed reports. It is free at normal volumes and it is not compliant by default. Configuration, retention settings, regional handling and consent signalling decide that, and that work sits in ethical analytics.
hreflang
hreflang is annotation that tells search engines which language and regional version of a page to show to which audience, declared in the page head, in the sitemap or in HTTP headers. Every version must reference every other version including itself, which is why hreflang breaks so often on large multi-country estates. Getting it wrong splits authority across near-duplicate pages.
Precision

Terms that get misused

Eight distinctions that account for a disproportionate share of the confusion we are asked to unpick.

Questions

Frequently asked questions

Broader questions about how we work are answered in the full FAQ.

Who is this glossary for?

People who have to make a decision and need the vocabulary to be precise about it: a marketing director asked to explain generative engine optimisation to a board, a general counsel reviewing an AI feature, a developer inheriting a consent platform, a procurement lead reading a supplier's accessibility report. The definitions assume intelligence and no prior specialism. Where a term is genuinely contested, the entry says so instead of picking a side quietly.

Why do some entries say a term is contested or unsettled?

Because they are. Generative engine optimisation and answer engine optimisation are used interchangeably by some practitioners and distinguished by others. Hallucination is disliked by researchers who prefer confabulation. llms.txt is a proposal with no confirmed adoption by a major AI vendor. Presenting an unsettled term as settled makes a glossary less useful, not more, so we mark the disagreement and explain what turns on it.

What is the difference between GEO, AEO and SEO?

SEO aims to place a page in a ranked list of links. Answer engine optimisation aims to have a passage lifted as the direct answer. Generative engine optimisation aims to be the source an AI assistant draws on and cites when it composes an answer from several places. The underlying technical work overlaps heavily, because all three need a crawlable, fast, well-structured site. What differs is the shape of the content and how success is measured. We cover this in AI search optimisation.

Should we block AI crawlers?

Not with one rule. Training crawlers, search crawlers and user-triggered fetchers are separate agents with separate consequences. Blocking GPTBot keeps content out of future OpenAI training sets. Blocking OAI-SearchBot removes you from ChatGPT's live results. Blocking ChatGPT-User means a user who pastes your URL into a conversation gets nothing back. Most organisations want some of those outcomes and not others, and the table on this page sets out which is which.

Is llms.txt worth publishing?

It costs an hour and it is harmless, so publish one if you like. Be honest about what it is: a community proposal for a Markdown index of your important pages, not a standard, and no major AI vendor has publicly confirmed that it reads one. Anyone selling llms.txt as the fix for AI visibility is selling the cheap part of a harder job. Crawl access, page structure and content quality are what actually decide whether you get cited.

Is a consent banner enough to comply with UK GDPR?

No. The banner is an interface. Compliance depends on whether non-essential scripts are genuinely blocked until consent is given, whether refusing is as easy as accepting, whether the choice is recorded in a way you can produce as evidence, and whether the underlying processing has a lawful basis and a retention rule. Most of the failures we find are in the tag layer behind the banner, not the banner itself. See consent management.

Do accessibility overlays work?

No, and the entry in this glossary says so plainly. Overlays are widely criticised by disabled users and accessibility practitioners. They do not repair the underlying markup, they can interfere with a user's own assistive technology, and installing one has not protected organisations from complaints or litigation. The work that does help is unglamorous: semantic markup, keyboard operability, contrast, focus visibility and real testing with assistive technology. That is what our accessibility service does.

How current is this page?

The AI search terms move fastest, because crawler names, robots.txt tokens and product surfaces change on vendor timetables rather than ours. We revise entries when the underlying fact changes rather than on a fixed schedule, and the review date appears at the foot of the page. If you find an entry that has gone stale, tell us and we will correct it.

Can we quote these definitions?

Yes. Quote them in a policy, a training deck, an internal wiki or a supplier brief, with a link back to this page. We would rather the vocabulary spread than sit behind a form. If you need a definition adapted to your own regulatory context, that is a conversation rather than a copy and paste, and we are happy to have it.

Written and reviewed by the Green Arrow Consultancy team, led by Darren Tyler, founder and chief executive.

Green Arrow Consultancy Ltd, Cardiff, Wales. Company number 12491770. ICO registration ZA822868. Member of the International Association of Privacy Professionals. Last reviewed .

Found a definition that has gone stale?

The AI search terms move on vendor timetables rather than ours. If an entry is out of date, or you need one adapted for your own regulatory context, tell us and we will look at it.