A chat interface is one of the harder accessibility problems on a modern website, and one of the least examined. It updates continuously, it produces content nobody authored, and it is usually built in a hurry. Here is what actually goes wrong, and what to build instead.
An accessible AI chat interface announces completed messages through a polite live region, keeps focus where the user put it, offers a real stop control, and makes every element reachable by keyboard including citations. WCAG 2.2 criteria that bite hardest are Status Messages, Focus Visible, Target Size and Consistent Help. Automated scanners catch almost none of it.
Accessibility work on a conventional page is largely a matter of structure. Headings in order, labels on inputs, alternative text on images, enough contrast, everything reachable by keyboard. The page holds still while you check it.
A chat interface holds still for nobody. Content arrives without the user requesting each piece of it, arrives gradually, arrives at a length nobody predicted, and arrives into a region that already contains everything said before it. Add citations that open panels, suggested prompts that appear and vanish, a thinking indicator that animates, and controls that attach themselves to each message once it settles, and you have a component with more moving parts than most single-page applications had a decade ago.
The second reason is organisational. These interfaces are usually built quickly, often by a team excited about the model rather than the front end, and frequently outside the design system where the organisation's accessible components live. We have audited AI features that were markedly less accessible than the twelve-year-old website they were bolted onto, built by the same company, because the chat widget was treated as a demonstration that got promoted.
The third reason is that the failures are invisible to the people shipping them. Every defect described below passes a sighted mouse test comfortably. The interface looks lovely. You have to put a keyboard and a screen reader in front of it before anything is wrong, and on most projects nobody does that until an audit, which happens after launch.
This is not a small population. Screen reader users, keyboard-only users, people with low vision, people with vestibular conditions triggered by motion, and people with cognitive or attention differences who need to control pace are all affected by the specific defects in this article. And these interfaces are increasingly the front door to support, so an inaccessible one does not merely inconvenience somebody, it removes their route to help.
These are the findings that recur across the AI interfaces we audit. The order is roughly how often we see them.
| What the user hits | Who it affects | The fix |
|---|---|---|
| A reply arrives and nothing is announced, so the user waits in silence | Screen reader users | Write the completed message into a polite live region; give the transcript a role that conveys new entries |
| The reply is announced over and over as it streams in | Screen reader users | Keep the streaming element out of the live region; announce once on completion, or at sentence boundaries |
| Focus jumps to the new message, losing the user's place | Screen reader and keyboard users | Leave focus in the composer; provide a deliberate control to move to the latest response |
| Focus disappears entirely when the response renders | Keyboard users | Never remove or re-mount the focused element; if the container must re-render, restore focus explicitly |
| Send only works by pressing Enter in the text box | Keyboard and switch users, and anyone using speech input | A real button element, in the tab order, with an accessible name, alongside the Enter shortcut |
| No way to stop a long response once it starts | Everyone, particularly cognitive and attention needs | A stop control that is reachable and operable while generation is in progress, not after it |
| A session times out mid-conversation and the transcript is gone | Cognitive needs, motor impairments, anyone composing slowly | Warn before expiry, allow the limit to be extended, and preserve the draft and the transcript |
| The thinking indicator animates continuously | Vestibular and attention needs | Respect reduced motion preferences; offer a static state; keep any animation short and non-essential |
| Thinking and streaming text is pale grey on white | Low vision users | Meet text contrast in every state, including in-progress ones; meet non-text contrast for indicators and focus rings |
| Citation chips cannot be reached with the keyboard | Keyboard and screen reader users | Real focusable controls with meaningful accessible names, in the tab order, with focus returned on close |
| Suggested prompt chips are reachable but unlabelled | Screen reader users | Give each one a name that reads as a question, and group them with a label saying what they are |
| Icon buttons on each message are tiny and packed together | Motor impairments, touch users, anyone on a small screen | Meet the minimum target size, add spacing, and give every icon-only control a text alternative |
| The transcript reads as one undifferentiated block | Screen reader users | Structure it so each turn is a distinct item with the speaker conveyed in text, not by colour or alignment |
| The assistant is the only route to help and moves around the site | Cognitive needs, and everyone under pressure | Keep help mechanisms in a consistent relative order across pages, and keep a non-AI route available |
Compiled from accessibility reviews of chat and assistant interfaces. None of these were caught by the client's automated scan.
Almost everything difficult about an accessible chat interface concentrates in one decision: how and when the screen reader is told that something has happened.
A live region is a container marked so that assistive technology announces changes to it without moving focus. That is precisely the behaviour you want for an arriving reply, and it is what Status Messages at Level AA is asking for. The mistake is to attach the live region to the element that is being streamed into. Screen readers respond to mutations, and a token-by-token stream produces a very large number of mutations. The result is either a stuttering repetition of partial sentences or, on some combinations, the entire message re-read from the beginning several times.
The pattern that works is to separate the visible stream from the announced content. Render the streaming text visually in a region assistive technology is told to ignore while it is in flight, and write the finished message into the live region once, when generation completes. For long responses where the wait would be uncomfortable, announce at sentence or paragraph boundaries instead, so the user hears coherent units rather than fragments.
Politeness here is technical as well as social. An assertive region interrupts whatever the screen reader is currently speaking, which in a chat interface frequently means cutting across the user re-reading an earlier answer. Use polite for replies, for status changes and for confirmations. Reserve assertive for things that genuinely cannot wait: a failed request that ends the interaction, or a session about to expire. If everything is urgent, the interface is unusable.
Users need to know that the request was received, that a response is being generated, that it has finished, and that it was stopped or failed. Sighted users get all of this from an animated indicator. Announce it in words: a short, non-repeating status message when generation starts, and the completed reply when it ends. Do not announce a progress state on a loop.
The transcript should be marked up so that each turn is a discrete item and the speaker is conveyed in text that assistive technology can read. On a great many interfaces the only thing distinguishing the user's words from the assistant's is background colour and alignment, both of which are invisible to a screen reader and unreliable for users who override colours. Say who is speaking, in content, once per turn.
It is tempting to send focus to the new message so the user is taken straight to it. Do not. Focus belongs where the user put it, and moving it under them is disorienting for screen reader users and disruptive for anyone mid-composition. Announce the message through the live region and leave focus in the composer. Then provide a deliberate way to reach the response: a control that moves focus to the latest reply, or a transcript that is navigable by heading or by list item.
The related failure is focus being destroyed rather than moved. Many chat components re-render the whole transcript when a message lands. If the element that had focus is replaced, focus falls back to the document and the user is silently returned to the top of the page. This one is easy to miss and catastrophic to experience. If a re-render is unavoidable, capture what had focus and restore it.
Focus indicators are still routinely removed for aesthetic reasons, and chat interfaces are among the worst offenders because the design tends toward soft, borderless elements. Every focusable element needs an indicator that is clearly perceivable and meets non-text contrast against its background. Check it in the composer, on the send control, on the stop control, on every citation, on every suggested prompt and on the per-message icon buttons. Also check that a sticky composer or a floating panel does not sit on top of the focused element and hide it, which WCAG 2.2 addresses directly with its focus obscuring criterion.
Streaming output is auto-updating content, it starts without the user authorising each part of it, and it regularly runs beyond five seconds. That places it squarely within Pause, Stop, Hide at Level A. The control has to exist, has to be reachable by keyboard while generation is running, and has to be announced when it appears. A stop button that only becomes focusable after the response has finished is not a stop button.
Conversational interfaces often sit behind a session timeout inherited from an authenticated area. If a limit exists, the user has to be warned before it expires and given a straightforward way to extend it, and the transcript and any draft should survive. Composing a question can take a long time. An interface that discards a carefully written message because a timer expired is failing the people most likely to need it.
Typing indicators, message entrance animations and shimmering placeholder blocks all move. Respect the operating system reduced motion preference and provide a static equivalent. Nothing essential should be communicated only by animation, and a user who has asked for less movement should get a calm interface rather than a slightly slower one.
The order matters. Retrofitting structure onto a finished chat component is considerably more work than starting with it.
Use the organisation's existing accessible button, input and dialogue components. Most chat interfaces fail because they were built outside the system that already solved these problems.
Each turn a discrete item, the speaker conveyed in text, timestamps available but not noisy. Get this right before adding any dynamic behaviour, because everything else depends on it.
One region, written once per message, holding the finished text. Keep the streaming element out of it and hidden from assistive technology while it is in flight.
A button element for send, a button element for stop, buttons for citations and suggested prompts. Accessible names that describe the action or the source, not the icon. Everything in the tab order in a sensible sequence.
On send, on completion, on stop, on error, on opening and closing a citation panel. Write down where focus should be after each one, then implement that rather than letting the framework decide.
Idle, focused, disabled, thinking, streaming, stopped and error. Thinking and streaming states are where designers reach for pale grey, and they are states users spend a lot of time looking at.
Static alternatives for every animation. Minimum target sizes and adequate spacing for the small icon controls that accumulate around each message.
In that order, on desktop and on mobile. Fix what you find, then run the automated scan last to catch the mechanical mistakes you missed.
Where a criterion number is not given here, the requirement is named in words instead. EN 301 549 incorporates WCAG at Level AA, which is the standard the European Accessibility Act points at in practice.
An automated scan is a floor, not a test. Tools such as axe are excellent at what they measure: contrast ratios, missing accessible names, structural errors, invalid attribute use. They will pass a chat interface that announces the same paragraph four times, loses focus on every reply, and offers a stop control that no keyboard user can reach. Those are failures that only exist across time, and a static scan has no way to see them.
First, unplug the mouse. Send a message, stop a response, open a citation, close it, reach every control on every message, and return to the composer. Note every point where you cannot see where you are, or cannot get somewhere, or cannot get back.
Then turn on a screen reader and do it again. Test with at least two, because behaviour around live regions varies considerably between them. On Windows, NVDA and JAWS. On macOS and iOS, VoiceOver. On Android, TalkBack. Listen for the specific things that go wrong: silence when a reply arrives, repetition while it streams, an unannounced state change, a transcript that gives no indication of who is speaking.
Then test the states that only happen occasionally. An error mid-stream. A stopped generation. An empty response. A very long response. A session expiring. These are the states that get built last and tested never, and they are disproportionately where accessibility breaks.
Finally, if you can, test with someone who uses assistive technology daily. The gap between a developer operating a screen reader for the first time and a person who uses one every day is enormous, and the findings that matter most tend to come from the second group.
Green Arrow Consultancy has run accessibility programmes across client website estates for years, and that practice now covers AI interfaces as well as pages and documents. We also build accessibility tooling of our own: one of our production systems is a document accessibility remediation tool that repairs PowerPoint and PDF libraries, generalised into the live demonstration at Neuro Access. Auditing an assistant is described under accessibility services, and building one that passes from the start is part of how we deliver under AI consulting.
One closing point. An accessible chat interface is a better interface for everyone. A stop control, a clear indication of state, a keyboard route to every source citation and a transcript you can navigate are all things a power user wants too. The accessibility work is not a tax on the design, it is the part of the design that survives contact with real users.
More terminology in the glossary, and more writing in insights.
Because they were built as a novelty rather than as a component, usually at speed, and because the accessibility patterns for a streaming, asynchronous, continuously updating region are less familiar than the patterns for a form or a navigation menu. Most teams know how to label an input. Far fewer have had to decide how a screen reader should experience a paragraph that arrives four words at a time over eight seconds. The failure is not indifference, it is an unfamiliar problem shipped on a short deadline.
Status Messages, criterion 4.1.3 at Level AA, is the one people reach for. It requires that content conveying status, progress or results can be presented to assistive technology without receiving focus. An arriving assistant reply is exactly that. In practice you also have to satisfy Info and Relationships for the transcript structure, Name, Role and Value for the controls, and Focus Order for what happens after a response lands, so treating it as one criterion is how teams end up with a transcript that announces but cannot be navigated.
Polite, in nearly every case. Assertive interrupts whatever the screen reader is currently saying, including the user's own re-reading of an earlier message, and in a chat interface that produces an experience people describe as being shouted over. Reserve assertive for genuine interruptions: an error that stops the interaction, or a session about to expire. Everything else, including the assistant's answer, should queue politely behind whatever the user is doing.
Do not put the streaming element inside the live region. Render the visible streaming text in a container that assistive technology ignores, and write to the live region once, when the message is complete. Some teams announce at sentence boundaries instead for long responses, which is a reasonable compromise if the visible stream is slow. The pattern to avoid is a live region attached directly to an element whose text content changes dozens of times a second, because most screen readers will either repeat the whole message or produce an unintelligible stutter.
Yes, and there is a defensible criterion behind it. Pause, Stop, Hide at Level A covers content that moves, blinks or auto-updates, starts automatically and lasts more than five seconds. A streaming response is auto-updating content that the user did not individually authorise token by token, and long responses routinely run past five seconds. Beyond compliance it is simply respectful: a user who realises in the first sentence that the answer is wrong should not have to sit through four hundred more words.
Target Size (Minimum), criterion 2.5.8 at Level AA in WCAG 2.2, sets a floor of 24 by 24 CSS pixels for pointer targets, with exceptions including targets that have sufficient spacing around them and inline targets within a sentence. This bites on chat interfaces because of the small icon buttons that accumulate around a message: copy, regenerate, thumbs up, thumbs down, cite, expand. They are frequently well under the floor and packed tightly together. The Level AAA version, Target Size (Enhanced), asks for 44 by 44.
As real, focusable, labelled controls in the tab order, positioned after the text they support, with an accessible name that says what the source is rather than just a number. A citation rendered as a superscript span with a click handler is invisible to keyboard users and meaningless to a screen reader. If the citation opens a panel, manage focus into the panel and return it to the citation on close. This is one of the most commonly broken elements in AI interfaces, and it is the element most likely to be used by somebody checking whether the answer is true.
Almost none of them. Automated scanners such as axe are good at contrast, missing labels and structural errors, and they will catch a genuinely unlabelled send button. They cannot tell you that the live region announced the same paragraph three times, that focus vanished when the response arrived, that the transcript reads as an undifferentiated wall with no indication of who is speaking, or that the stop control cannot be reached before the response has finished. Those are behavioural failures over time. They need a person, a keyboard and a screen reader.
If the chatbot is part of a product or service in scope, then the accessibility requirements apply to it like any other part of the interface. The Act points at harmonised standards, and in practice that means EN 301 549, which incorporates WCAG at Level AA. There is no carve-out for interfaces that happen to be powered by a language model. The same logic applies to public sector accessibility duties in the UK and to procurement requirements in the United States.
We test AI interfaces with a keyboard and a screen reader, report findings against WCAG 2.2 at Level AA, and hand the engineering team fixes rather than a spreadsheet of criteria.