What automated checks miss
Lighthouse, axe, scanners, overlay widgets and AI testing all read your code, and some things they judge well. Whether a person can actually use the page is a different question, and much of WCAG asks it.
In the tools’ own words
- 57%
What axe-core says it finds automatically, on average. It marks other results “incomplete” where a person has to look.
- 40%
The share of 142 barriers on one test page that the best tool in GOV.UK’s audit fully detected. The weakest found 13%.
- 100
A Lighthouse score is a weighted average of the automated audits it ran. Its manual checks don’t count toward it.
The check passes. A person is still stuck.
Each rule below can be met in the code while the task fails. Only someone using the page finds out.
-
Passes The check: Every image has alt text.
A person finds: The product photo’s alt text is “image”. A screen reader user can’t tell which colour they are buying.
WCAG 1.1.1
-
Passes The check: Every field has a label.
A person finds: Two fields are both labelled “Name”. One wants the name on the card, the other a nickname for saving it.
WCAG 2.4.6
-
Passes The check: Every control can be reached with Tab.
A person finds: Tab goes through the whole footer before it reaches the checkout form.
WCAG 2.4.3
-
Passes The check: The dialog has a dialog role and a name.
A person finds: Tab moves out of the open dialog and into the page behind it.
WCAG 2.4.3
-
Passes The check: No rule failed on the menu.
A person finds: The menu opens on hover and click. A keyboard can’t open it at all.
WCAG 2.1.1
-
Passes The check: The error message is in the page.
A person finds: Nothing announces it. After pressing Pay, a screen reader user hears silence.
WCAG 4.1.3
-
Passes The check: The page sets a responsive viewport.
A person finds: At 400% zoom the price table scrolls both ways and the Buy button is off screen.
WCAG 1.4.10
-
Passes The check: The video has a captions track.
A person finds: The captions were generated automatically and get every product name wrong.
WCAG 1.2.2
What about an AI agent that uses the browser?
An agent can press Tab, open the menu and read the accessibility tree. That makes it a better first pass than a rule. It still can’t stand in for the person, for three reasons.
It isn’t the person.
Using a screen reader, zoom or a switch every day, with your own settings and pace, is an experience the agent doesn’t have. It can’t say whether a page makes sense, or wears out, the person relying on it.
Its results don’t repeat.
Run it twice and the findings can differ. A test result has to come out the same when someone checks it again.
Nobody stands behind it.
A finding a team acts on, or that goes into a conformance report, needs a tester who followed a stated method and can show the evidence. An agent’s output is a lead for that person to confirm.
What each kind of tool is for
Lighthouse, axe and checks in CI
Good for
Catching the problems a rule can find, on every change, before they ship. Keep the gate.
Can’t tell you
A green build means none of the rules it ran failed. It doesn’t mean the page works for the people using it.
Scanners and monitoring dashboards
Good for
Tracking rule violations across many pages over time.
Can’t tell you
The score is the share of rules passed, not a conformance result.
Overlay widgets
Good for
Offering display preferences that some visitors like.
Can’t tell you
They change the page after it loads rather than fixing its code, and can get in the way of the screen reader or zoom a visitor already uses.
AI testing and AI auto-fix
Good for
A fast first pass, and draft fixes for a person to review.
Can’t tell you
It guesses at meaning. Generated alt text and labels can be wrong and still pass the rule, and the results can change from one run to the next.
Audits exported from a scan
Good for
A complete list of rule violations.
Can’t tell you
Which few problems actually stop people, and how to fix them in your code.
Run the checks. Then test the tasks.
Keep automated checks in CI so the problems they can find never come back. Then have a person go through the tasks that matter, such as signing up, searching and checking out, with a keyboard, a screen reader and zoom. That is how I work: automation first, then testing by hand to the Section 508 Trusted Tester process.