# Accessible AI Interfaces and WCAG 2.2 | Green Arrow Consultancy

> How to build an accessible AI chat interface against WCAG 2.2: live regions, streaming announcements, focus management, a stop control and contrast.

Source: https://greenarrow.app/insights/accessible-ai-interfaces/
Last updated: 2026-09-04
Publisher: Green Arrow Consultancy Ltd, Cardiff, Wales, United Kingdom

---

Insights · Accessibility

# Accessible AI interfaces: what WCAG 2.2 asks of a chatbot

A chat interface is one of the harder accessibility problems on a modern website, and one of the least examined. It updates continuously, it produces content nobody authored, and it is usually built in a hurry. Here is what actually goes wrong, and what to build instead.

WCAG 2.2 AA EN 301 549 Tested with a screen reader 
 Talk to us about an audit → Accessibility services 
 
 
 
 
 
 Quick answer

An accessible AI chat interface announces completed messages through a polite live region, keeps focus where the user put it, offers a real stop control, and makes every element reachable by keyboard including citations. WCAG 2.2 criteria that bite hardest are Status Messages, Focus Visible, Target Size and Consistent Help. Automated scanners catch almost none of it.

## Key points

- The core problem is that a chat transcript is auto-updating content, and most accessibility patterns assume content that sits still.

- Attaching a live region directly to streaming text is the single most common defect, and it makes the interface worse than saying nothing.

- Announce the completed message. Do not announce every token, and do not steal focus from where the user left it.

- Citations, suggested prompts and the small icon buttons around a message are where keyboard access usually breaks.

- An automated scan will pass an interface that is unusable with a screen reader. Test it with an actual one.

## On this page

- Why chat interfaces fail so reliably

- Symptoms, who they affect, and the fix

- Getting announcements right

- Focus, stop controls and time

- How to build it, in order

- The WCAG 2.2 criteria that bite

- Testing it properly

- Frequently asked questions

Diagnosis

## Why chat interfaces fail so reliably

Accessibility work on a conventional page is largely a matter of structure. Headings in order, labels on inputs, alternative text on images, enough contrast, everything reachable by keyboard. The page holds still while you check it.

A chat interface holds still for nobody. Content arrives without the user requesting each piece of it, arrives gradually, arrives at a length nobody predicted, and arrives into a region that already contains everything said before it. Add citations that open panels, suggested prompts that appear and vanish, a thinking indicator that animates, and controls that attach themselves to each message once it settles, and you have a component with more moving parts than most single-page applications had a decade ago.

The second reason is organisational. These interfaces are usually built quickly, often by a team excited about the model rather than the front end, and frequently outside the design system where the organisation's accessible components live. We have audited AI features that were markedly less accessible than the twelve-year-old website they were bolted onto, built by the same company, because the chat widget was treated as a demonstration that got promoted.

The third reason is that the failures are invisible to the people shipping them. Every defect described below passes a sighted mouse test comfortably. The interface looks lovely. You have to put a keyboard and a screen reader in front of it before anything is wrong, and on most projects nobody does that until an audit, which happens after launch.

This is not a small population. Screen reader users, keyboard-only users, people with low vision, people with vestibular conditions triggered by motion, and people with cognitive or attention differences who need to control pace are all affected by the specific defects in this article. And these interfaces are increasingly the front door to support, so an inaccessible one does not merely inconvenience somebody, it removes their route to help.

Audit findings

## Symptoms, who they affect, and the fix

These are the findings that recur across the AI interfaces we audit. The order is roughly how often we see them.

| What the user hits | Who it affects | The fix |
|---|---|---|
| A reply arrives and nothing is announced, so the user waits in silence | Screen reader users | Write the completed message into a polite live region; give the transcript a role that conveys new entries |
| The reply is announced over and over as it streams in | Screen reader users | Keep the streaming element out of the live region; announce once on completion, or at sentence boundaries |
| Focus jumps to the new message, losing the user's place | Screen reader and keyboard users | Leave focus in the composer; provide a deliberate control to move to the latest response |
| Focus disappears entirely when the response renders | Keyboard users | Never remove or re-mount the focused element; if the container must re-render, restore focus explicitly |
| Send only works by pressing Enter in the text box | Keyboard and switch users, and anyone using speech input | A real button element, in the tab order, with an accessible name, alongside the Enter shortcut |
| No way to stop a long response once it starts | Everyone, particularly cognitive and attention needs | A stop control that is reachable and operable while generation is in progress, not after it |
| A session times out mid-conversation and the transcript is gone | Cognitive needs, motor impairments, anyone composing slowly | Warn before expiry, allow the limit to be extended, and preserve the draft and the transcript |
| The thinking indicator animates continuously | Vestibular and attention needs | Respect reduced motion preferences; offer a static state; keep any animation short and non-essential |
| Thinking and streaming text is pale grey on white | Low vision users | Meet text contrast in every state, including in-progress ones; meet non-text contrast for indicators and focus rings |
| Citation chips cannot be reached with the keyboard | Keyboard and screen reader users | Real focusable controls with meaningful accessible names, in the tab order, with focus returned on close |
| Suggested prompt chips are reachable but unlabelled | Screen reader users | Give each one a name that reads as a question, and group them with a label saying what they are |
| Icon buttons on each message are tiny and packed together | Motor impairments, touch users, anyone on a small screen | Meet the minimum target size, add spacing, and give every icon-only control a text alternative |
| The transcript reads as one undifferentiated block | Screen reader users | Structure it so each turn is a distinct item with the speaker conveyed in text, not by colour or alignment |
| The assistant is the only route to help and moves around the site | Cognitive needs, and everyone under pressure | Keep help mechanisms in a consistent relative order across pages, and keep a non-AI route available |

Compiled from accessibility reviews of chat and assistant interfaces. None of these were caught by the client's automated scan.

Announcements

## Getting announcements right

Almost everything difficult about an accessible chat interface concentrates in one decision: how and when the screen reader is told that something has happened.

### Use a live region, and use it once

A live region is a container marked so that assistive technology announces changes to it without moving focus. That is precisely the behaviour you want for an arriving reply, and it is what Status Messages at Level AA is asking for. The mistake is to attach the live region to the element that is being streamed into. Screen readers respond to mutations, and a token-by-token stream produces a very large number of mutations. The result is either a stuttering repetition of partial sentences or, on some combinations, the entire message re-read from the beginning several times.

The pattern that works is to separate the visible stream from the announced content. Render the streaming text visually in a region assistive technology is told to ignore while it is in flight, and write the finished message into the live region once, when generation completes. For long responses where the wait would be uncomfortable, announce at sentence or paragraph boundaries instead, so the user hears coherent units rather than fragments.

### Polite, not assertive

Politeness here is technical as well as social. An assertive region interrupts whatever the screen reader is currently speaking, which in a chat interface frequently means cutting across the user re-reading an earlier answer. Use polite for replies, for status changes and for confirmations. Reserve assertive for things that genuinely cannot wait: a failed request that ends the interaction, or a session about to expire. If everything is urgent, the interface is unusable.

### Announce state, not just content

Users need to know that the request was received, that a response is being generated, that it has finished, and that it was stopped or failed. Sighted users get all of this from an animated indicator. Announce it in words: a short, non-repeating status message when generation starts, and the completed reply when it ends. Do not announce a progress state on a loop.

### Give the transcript a real structure

The transcript should be marked up so that each turn is a discrete item and the speaker is conveyed in text that assistive technology can read. On a great many interfaces the only thing distinguishing the user's words from the assistant's is background colour and alignment, both of which are invisible to a screen reader and unreliable for users who override colours. Say who is speaking, in content, once per turn.

Control

## Focus, stop controls and time

### Do not move focus when a reply arrives

It is tempting to send focus to the new message so the user is taken straight to it. Do not. Focus belongs where the user put it, and moving it under them is disorienting for screen reader users and disruptive for anyone mid-composition. Announce the message through the live region and leave focus in the composer. Then provide a deliberate way to reach the response: a control that moves focus to the latest reply, or a transcript that is navigable by heading or by list item.

The related failure is focus being destroyed rather than moved. Many chat components re-render the whole transcript when a message lands. If the element that had focus is replaced, focus falls back to the document and the user is silently returned to the top of the page. This one is easy to miss and catastrophic to experience. If a re-render is unavoidable, capture what had focus and restore it.

### Make focus visible, and keep it visible

Focus indicators are still routinely removed for aesthetic reasons, and chat interfaces are among the worst offenders because the design tends toward soft, borderless elements. Every focusable element needs an indicator that is clearly perceivable and meets non-text contrast against its background. Check it in the composer, on the send control, on the stop control, on every citation, on every suggested prompt and on the per-message icon buttons. Also check that a sticky composer or a floating panel does not sit on top of the focused element and hide it, which WCAG 2.2 addresses directly with its focus obscuring criterion.

### A real stop control

Streaming output is auto-updating content, it starts without the user authorising each part of it, and it regularly runs beyond five seconds. That places it squarely within Pause, Stop, Hide at Level A. The control has to exist, has to be reachable by keyboard while generation is running, and has to be announced when it appears. A stop button that only becomes focusable after the response has finished is not a stop button.

### Time limits

Conversational interfaces often sit behind a session timeout inherited from an authenticated area. If a limit exists, the user has to be warned before it expires and given a straightforward way to extend it, and the transcript and any draft should survive. Composing a question can take a long time. An interface that discards a carefully written message because a timer expired is failing the people most likely to need it.

### Reduced motion

Typing indicators, message entrance animations and shimmering placeholder blocks all move. Respect the operating system reduced motion preference and provide a static equivalent. Nothing essential should be communicated only by animation, and a user who has asked for less movement should get a calm interface rather than a slightly slower one.

Implementation

## How to build it, in order

The order matters. Retrofitting structure onto a finished chat component is considerably more work than starting with it.

- 01

### Start from your design system, not from the demo

Use the organisation's existing accessible button, input and dialogue components. Most chat interfaces fail because they were built outside the system that already solved these problems.

- 02

### Build the transcript as structured content first

Each turn a discrete item, the speaker conveyed in text, timestamps available but not noisy. Get this right before adding any dynamic behaviour, because everything else depends on it.

- 03

### Add a single polite live region, wired to completion

One region, written once per message, holding the finished text. Keep the streaming element out of it and hidden from assistive technology while it is in flight.

- 04

### Make every control real and reachable

A button element for send, a button element for stop, buttons for citations and suggested prompts. Accessible names that describe the action or the source, not the icon. Everything in the tab order in a sensible sequence.

- 05

### Manage focus explicitly at every transition

On send, on completion, on stop, on error, on opening and closing a citation panel. Write down where focus should be after each one, then implement that rather than letting the framework decide.

- 06

### Meet contrast in every state

Idle, focused, disabled, thinking, streaming, stopped and error. Thinking and streaming states are where designers reach for pale grey, and they are states users spend a lot of time looking at.

- 07

### Respect reduced motion and check target sizes

Static alternatives for every animation. Minimum target sizes and adequate spacing for the small icon controls that accumulate around each message.

- 08

### Test with a keyboard, then with a screen reader, then with a user

In that order, on desktop and on mobile. Fix what you find, then run the automated scan last to catch the mechanical mistakes you missed.

Standards

## The WCAG 2.2 criteria that bite

Where a criterion number is not given here, the requirement is named in words instead. EN 301 549 incorporates WCAG at Level AA, which is the standard the European Accessibility Act points at in practice.

- Status Messages, 4.1.3, Level AA. The arriving reply, the generating state and the error state all convey status and must reach assistive technology without taking focus. This is the criterion that governs the live region.

- Focus Visible, 2.4.7, Level AA. Every focusable element needs a visible indicator. Chat interfaces lose this most often on citations and on icon-only message controls.

- Focus Not Obscured, Level AA in WCAG 2.2. A sticky composer, a floating panel or a suggestion tray must not hide the element that currently has focus. Test by tabbing through a long transcript.

- Focus Appearance, 2.4.13, Level AAA in WCAG 2.2. Sets a minimum size and contrast for the focus indicator itself. Worth meeting even though it is AAA, because soft borderless chat designs tend to produce faint indicators.

- Target Size (Minimum), 2.5.8, Level AA in WCAG 2.2. A floor of 24 by 24 CSS pixels for pointer targets, with spacing exceptions. The cluster of copy, regenerate and feedback icons under each message is the usual failure.

- Consistent Help, 3.2.6, Level A in WCAG 2.2. If the assistant is a help mechanism, it must appear in the same relative order across pages, and a non-AI route to help should remain available.

- Pause, Stop, Hide, 2.2.2, Level A. Covers auto-updating content that starts automatically and runs beyond five seconds. This is the basis for requiring a working stop control on a streaming response.

- Timing Adjustable, 2.2.1, Level A. Where a session or interaction has a time limit, warn before it expires and let the user extend it without losing the conversation.

- Keyboard, 2.1.1, and No Keyboard Trap, 2.1.2, Level A. Everything operable by pointer must be operable by keyboard, and nothing may trap focus. Citation panels and modal source viewers are where traps appear.

- Info and Relationships, 1.3.1, Level A. Turn structure and speaker identity have to be conveyed programmatically, not only by colour, alignment or an avatar image.

- Name, Role, Value, 4.1.2, Level A. Icon-only controls, citation chips and suggestion chips all need an accessible name and the correct role. A styled span with a click handler has neither.

- Contrast (Minimum), 1.4.3, and Non-text Contrast, 1.4.11, Level AA. Text at 4.5 to 1, and interface components and focus indicators at 3 to 1, in every state including thinking, streaming and disabled.

Verification

## Testing it properly

An automated scan is a floor, not a test. Tools such as axe are excellent at what they measure: contrast ratios, missing accessible names, structural errors, invalid attribute use. They will pass a chat interface that announces the same paragraph four times, loses focus on every reply, and offers a stop control that no keyboard user can reach. Those are failures that only exist across time, and a static scan has no way to see them.

### The manual pass, in order

First, unplug the mouse. Send a message, stop a response, open a citation, close it, reach every control on every message, and return to the composer. Note every point where you cannot see where you are, or cannot get somewhere, or cannot get back.

Then turn on a screen reader and do it again. Test with at least two, because behaviour around live regions varies considerably between them. On Windows, NVDA and JAWS. On macOS and iOS, VoiceOver. On Android, TalkBack. Listen for the specific things that go wrong: silence when a reply arrives, repetition while it streams, an unannounced state change, a transcript that gives no indication of who is speaking.

Then test the states that only happen occasionally. An error mid-stream. A stopped generation. An empty response. A very long response. A session expiring. These are the states that get built last and tested never, and they are disproportionately where accessibility breaks.

Finally, if you can, test with someone who uses assistive technology daily. The gap between a developer operating a screen reader for the first time and a person who uses one every day is enormous, and the findings that matter most tend to come from the second group.

### Where this fits with our other work

Green Arrow Consultancy has run accessibility programmes across client website estates for years, and that practice now covers AI interfaces as well as pages and documents. We also build accessibility tooling of our own: one of our production systems is a document accessibility remediation tool that repairs PowerPoint and PDF libraries, generalised into the live demonstration at Neuro Access. Auditing an assistant is described under accessibility services, and building one that passes from the start is part of how we deliver under AI consulting.

One closing point. An accessible chat interface is a better interface for everyone. A stop control, a clear indication of state, a keyboard route to every source citation and a transcript you can navigate are all things a power user wants too. The accessibility work is not a tax on the design, it is the part of the design that survives contact with real users.

Questions

## Frequently asked questions

More terminology in the glossary, and more writing in insights.

### Why do AI chat interfaces fail accessibility so often?

Because they were built as a novelty rather than as a component, usually at speed, and because the accessibility patterns for a streaming, asynchronous, continuously updating region are less familiar than the patterns for a form or a navigation menu. Most teams know how to label an input. Far fewer have had to decide how a screen reader should experience a paragraph that arrives four words at a time over eight seconds. The failure is not indifference, it is an unfamiliar problem shipped on a short deadline.

### Which WCAG 2.2 success criterion covers new chat messages?

Status Messages, criterion 4.1.3 at Level AA, is the one people reach for. It requires that content conveying status, progress or results can be presented to assistive technology without receiving focus. An arriving assistant reply is exactly that. In practice you also have to satisfy Info and Relationships for the transcript structure, Name, Role and Value for the controls, and Focus Order for what happens after a response lands, so treating it as one criterion is how teams end up with a transcript that announces but cannot be navigated.

### Should the live region be polite or assertive?

Polite, in nearly every case. Assertive interrupts whatever the screen reader is currently saying, including the user's own re-reading of an earlier message, and in a chat interface that produces an experience people describe as being shouted over. Reserve assertive for genuine interruptions: an error that stops the interaction, or a session about to expire. Everything else, including the assistant's answer, should queue politely behind whatever the user is doing.

### How do you stop streaming text re-announcing on every token?

Do not put the streaming element inside the live region. Render the visible streaming text in a container that assistive technology ignores, and write to the live region once, when the message is complete. Some teams announce at sentence boundaries instead for long responses, which is a reasonable compromise if the visible stream is slow. The pattern to avoid is a live region attached directly to an element whose text content changes dozens of times a second, because most screen readers will either repeat the whole message or produce an unintelligible stutter.

### Does a chatbot need a stop button?

Yes, and there is a defensible criterion behind it. Pause, Stop, Hide at Level A covers content that moves, blinks or auto-updates, starts automatically and lasts more than five seconds. A streaming response is auto-updating content that the user did not individually authorise token by token, and long responses routinely run past five seconds. Beyond compliance it is simply respectful: a user who realises in the first sentence that the answer is wrong should not have to sit through four hundred more words.

### What is the target size requirement for chat controls?

Target Size (Minimum), criterion 2.5.8 at Level AA in WCAG 2.2, sets a floor of 24 by 24 CSS pixels for pointer targets, with exceptions including targets that have sufficient spacing around them and inline targets within a sentence. This bites on chat interfaces because of the small icon buttons that accumulate around a message: copy, regenerate, thumbs up, thumbs down, cite, expand. They are frequently well under the floor and packed tightly together. The Level AAA version, Target Size (Enhanced), asks for 44 by 44.

### How should citations be built so they are usable?

As real, focusable, labelled controls in the tab order, positioned after the text they support, with an accessible name that says what the source is rather than just a number. A citation rendered as a superscript span with a click handler is invisible to keyboard users and meaningless to a screen reader. If the citation opens a panel, manage focus into the panel and return it to the citation on close. This is one of the most commonly broken elements in AI interfaces, and it is the element most likely to be used by somebody checking whether the answer is true.

### Do automated accessibility tools catch these problems?

Almost none of them. Automated scanners such as axe are good at contrast, missing labels and structural errors, and they will catch a genuinely unlabelled send button. They cannot tell you that the live region announced the same paragraph three times, that focus vanished when the response arrived, that the transcript reads as an undifferentiated wall with no indication of who is speaking, or that the stop control cannot be reached before the response has finished. Those are behavioural failures over time. They need a person, a keyboard and a screen reader.

### Does the European Accessibility Act apply to a chatbot?

If the chatbot is part of a product or service in scope, then the accessibility requirements apply to it like any other part of the interface. The Act points at harmonised standards, and in practice that means EN 301 549, which incorporates WCAG at Level AA. There is no carve-out for interfaces that happen to be powered by a language model. The same logic applies to public sector accessibility duties in the UK and to procurement requirements in the United States.

Written and reviewed by the Green Arrow Consultancy team, led by Darren Tyler, founder and chief executive.

Green Arrow Consultancy Ltd, Cardiff, Wales. Company number 12491770. ICO registration ZA822868. Member of the International Association of Privacy Professionals. Last reviewed 04 September 2026.

Keep reading

## Related

### Website Accessibility

Audits, remediation and accessibility statements across websites, documents and AI interfaces.

Read this →

### Neuro Access

Our document accessibility remediation tool, running live on sample files.

Read this →

### Twelve privacy questions before you ship AI

The other pre-launch review an AI interface should pass before it goes anywhere near customers.

Read this →

## Have your assistant audited by people who build them

We test AI interfaces with a keyboard and a screen reader, report findings against WCAG 2.2 at Level AA, and hand the engineering team fixes rather than a spreadsheet of criteria.

Start a conversation → 
 Accessibility services
