Builders Camp

Tools

How to use AI for accessibility review, and where it stops

Deque's analysis of more than 2,000 audits and nearly 300,000 issues found automated testing completely covered 57 percent of them by volume, and that number is much lower when counted by WCAG success criteria instead. Automate the criteria that turn on presence, run a three-minute keyboard walk for the ones that turn on rendering, and stop calling the result an audit.

What does the 57 percent figure actually mean?

Deque analysed anonymised data from over 2,000 audits, covering more than 13,000 pages and nearly 300,000 identified issues, all first-time assessments run with its axe tooling, and reported that 57 percent of those issues were completely covered by automated testing. That is a real number from a real dataset and it is routinely quoted without its unit.

The unit is issues, not criteria. Counted the other way, by how many WCAG success criteria a machine can decide, the figure has long been estimated far lower, and Deque's own write-up frames the study as a deliberate move away from that older way of counting. Both numbers are true about different questions. The volume number tells you how much of the bug list automation will hand you. The criteria number tells you how much of the standard you have actually checked.

Which one matters depends on what you are doing. Clearing a backlog: use the volume number, automation is doing most of the work. Claiming conformance: use the criteria number, and the answer is that you have not checked most of it.

The criteria that turn on presence, and the ones that turn on meaning

SC 1.1.1 Non-text Content, at Level A, requires that all non-text content presented to the user has a text alternative serving the equivalent purpose. A tool settles the first half of that sentence in milliseconds: the attribute is there or it is not. It cannot settle the second half, because whether an alternative serves the equivalent purpose is a question about what the image was for.

This split runs through most of the standard. SC 1.3.1 Info and Relationships, also Level A, requires that information, structure and relationships conveyed through presentation can be programmatically determined or are available in text. A checker will tell you a heading element exists; only a reader can tell you the heading describes the section beneath it. A model is genuinely useful on the second half, and this is the place where a language model adds something a rule-based checker cannot: ask it whether each alt text would let somebody who cannot see the image understand what the image was doing there, and it produces a list of the generic ones. Then you check the list.

SC 3.3.1 Error Identification, Level A, is the other high-yield one. It requires that when an input error is automatically detected, the item in error is identified and the error described to the user in text. Both halves are visible in source, and both are commonly half-implemented: a red border with no text, or a message that names a field the user never saw.

What no static check can answer

Three criteria are worth knowing by number because they are the ones people assume were covered and were not.

SC 2.4.7 Focus Visible, Level AA, requires that any keyboard operable interface has a mode of operation where the focus indicator is visible. A checker can see that a focus style is declared. Whether the ring is visible against the background it actually lands on, at the contrast the theme produces at runtime, is a rendered question.

SC 2.4.11 Focus Not Obscured (Minimum), Level AA, requires that a component receiving keyboard focus is not entirely hidden by author-created content. This is the sticky header failure: tab down a long form, and the focused field scrolls underneath the header that is pinned to the top of the viewport. Nothing in the source says that will happen.

SC 2.5.8 Target Size (Minimum), Level AA, sets 24 by 24 CSS pixels as the minimum for pointer targets, with exceptions covering spacing, equivalent controls elsewhere on the page, inline targets, user agent control and Essential cases. The measurement is arithmetic once the box is rendered. Whether one of five exceptions applies is a judgement.

The three-minute test that finds more than the scan

Put the mouse down and tab through the screen from the top, completing the task with the keyboard alone.

  • Can you reach every control, including the ones inside menus, dialogs and custom components?
  • Can you see where you are at every stop, against the background the ring actually lands on?
  • Can you finish the task, including dismissing anything that opened, without touching the pointer?

Those three questions cover the focus criteria above, the reachability half of operability, and most of what an automated pass structurally cannot see. Three minutes per screen is a cost small enough that the only reason not to do it is not knowing it was your job.

The honest limitation is that a keyboard walk by a sighted developer is not a screen reader test by somebody who uses one daily, and nothing in this page substitutes for that. What it does is stop the obvious failures reaching the person whose time is expensive.

Where a model helps and where it flatters you

A language model is good at the meaning half: reading alt text and telling you which entries are generic, reading error messages and telling you which ones name an internal field, reading a form and telling you which labels are placeholders pretending to be labels. It is also good at turning a scan result into plain sentences an engineer can act on rather than a rule identifier nobody recognises.

It is bad at telling you it does not know. Asked whether a screen is accessible, it will answer, and the answer will be a confident paragraph assembled from the source it was given and the general shape of accessibility advice. Ask instead for a table: criterion number, verdict, and the specific element or line as evidence, with a required verdict of not checkable from source for anything that depends on rendering. The rows that come back marked not checkable are the list of what you owe a browser.

The pattern for running that pass against a real codebase is covered in running a design review with Claude Code, which also handles the contrast arithmetic. For the layer above the standard, where the question is whether the screen is usable rather than conformant, getting design feedback from AI against the usability heuristics is the companion pass.

Put it where it costs least

Before the design review, not after merge. A missing label found on a branch is a two-minute fix; the same label found in an audit six months later is a ticket competing with a roadmap. Builders Camp's Prototyping with AI bootcamp spends 1 week, 2 live sessions and 9 microlessons on the prototype stage where empty states, onboarding copy and UI concepts get decided, and several of these failures are cheapest to prevent right there. Product Manager Foundations covers the end-to-end process the check sits inside, over 2 weeks with 4 live sessions and 7 microlessons.

Usability Testing for Product Managers covers the part that no scan replaces: recruiting the right participants, writing tasks that reveal behaviour, and turning what you observe into prioritised fixes, in 1 week with 1 live session and 6 microlessons.

See the Usability Testing for Product Managers bootcamp

If you want one number to track rather than a conformance claim you cannot back, count how many of your screens have had a keyboard walk this quarter. It is a worse metric than a real audit and a far better one than a green scan result, because it is the only figure on this page that moves when somebody does the work.

Bootcamps referred in this Guide

Frequently asked questions

How much of accessibility can automation actually catch?

It depends on what you count. Deque analysed anonymised data from over 2,000 audits covering more than 13,000 pages and nearly 300,000 issues, and reported that 57 percent of issues were completely covered by automated testing. That figure counts issues by volume. Counted by WCAG success criteria instead, the long-standing industry estimate is far lower, because the criteria a machine can settle are a minority of the total.

Why do the two numbers differ so much?

Because the failures automation is best at are also the most numerous. Missing alternative text, unlabelled inputs and low-contrast text appear hundreds of times on a large site, so they dominate an issue count while representing a handful of criteria. Deque's own framing was a shift from counting testable criteria to counting the volume of real issues found.

Which WCAG 2.2 criteria does a machine genuinely settle?

The ones that turn on presence rather than meaning. SC 1.1.1 Non-text Content at Level A requires a text alternative that serves the equivalent purpose, and a tool can tell you the attribute exists. Whether the alternative serves the equivalent purpose is a reading comprehension question, which is why the criterion is only half automatable.

What needs a person and a keyboard?

SC 2.4.7 Focus Visible at Level AA requires a visible keyboard focus indicator, and SC 2.4.11 Focus Not Obscured (Minimum), also Level AA, requires that a focused component is not entirely hidden by author-created content. Both describe what happens when a real focus ring lands behind a real sticky header, which is a rendered-state question no static check answers.

Can a model check target sizes?

Partly. SC 2.5.8 Target Size (Minimum) at Level AA sets a minimum of 24 by 24 CSS pixels for pointer targets, with exceptions for spacing, equivalent controls, inline targets, user agent control and Essential cases. The arithmetic is trivial once you know the rendered box. Deciding whether an exception applies is a judgement about your interface.

What is the cheapest manual test worth running every time?

The keyboard walk. Put the mouse down, tab through the screen from the top, and complete the task. You will find focus order problems, focus rings hidden behind sticky elements, dialogs that do not trap focus and controls that cannot be reached at all, and it takes about three minutes per screen.

Does this replace an accessibility audit?

No. An audit involves assistive technology and, in the versions worth paying for, disabled users testing real tasks. An automated pass plus a keyboard walk clears the mechanical failures so an audit's expensive attention goes to the problems only a person finds. Builders Camp's Usability Testing for Product Managers covers running sessions and turning observations into prioritised issues over 1 week with 1 live session and 6 microlessons.

Sources

Written by

Andre Albuquerque

Andre Albuquerque

CEO of Builders Camp, SuperOperator, and other companies. Building products.

CEO of Builders Camp, SuperOperator, and other companies. Building products.

LinkedInMore guides by Andre Albuquerque
Ricardo Luiz

Ricardo Luiz

He is an accomplished Product Director, bringing a wealth of experience in driving innovation, building high-performing teams, and fostering collaborative environments.

He is an accomplished Product Director, bringing a wealth of experience in driving innovation, building high-performing teams, and fostering collaborative environments.

LinkedInMore guides by Ricardo Luiz

Last updated 2026-09-18

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.

See the Usability Testing for Product Managers bootcamp