HTML summary
A score answers a narrower question than it appears to answer.
Automated tools can identify real, code-detectable accessibility issues quickly and consistently. Their scores remain partial measurements of a particular tool, version, ruleset, configuration, page sample, and moment in time.
An automated score is evidence about an automated test. It is not, by itself, evidence that a website is accessible.
Automated Accessibility Scores Are Not Accessibility Evidence
What the review examined
A transparent synthesis without a manufactured universal percentage.
This is a rapid review by one evidence reviewer. It is transparent, but it is not exhaustive.
- 300
- records returned by the OpenAlex and Crossref API searches
- 278
- unique records after deduplication
- 20
- primary studies included in the synthesis
- 1
- secondary systematic mapping included
A statistical meta-analysis was considered and rejected because the studies measured different constructs with incompatible denominators and generally did not report poolable uncertainty data.
A practical evidence model
Three layers answer three different questions.
Automated checks
What repeatable, code-detectable findings did this tool identify?
Experienced manual review
What requires context, judgment, interaction, and a reproducible method?
Task-based user evidence
What happens when disabled people complete meaningful tasks with their own strategies and technologies?
Before you rely on a score
Ask what the number can actually support.
The full paper explains each question and provides a ten-question checklist for vendors and internal teams.
- 01
What exactly was scanned? A homepage cannot stand in for a form, document, account workflow, or state that was never tested.
- 02
Which tool, version, rules, and configuration produced it? Without provenance, the result cannot be reproduced or compared meaningfully.
- 03
What could the tool decide automatically? Separate definite failures, potential issues, and unresolved manual checks.
- 04
Who resolved the judgment calls? Experienced review and a defined method matter.
- 05
What interaction testing was completed? Keyboard, focus, zoom, reflow, screen-reader use, errors, touch, and dynamic states require interaction.
- 06
What claim is the score being used to support? A bounded scanner result is not a conformance, accessibility, or legal determination.
Claims and limitations
Useful evidence still needs honest boundaries.
The review supports the conclusion that automated accessibility testing is useful but incomplete and must be interpreted in context. It does not establish the accessibility or WCAG conformance of Wingspan Labs, HorizonSong, or a client website; legal compliance; a universal automation coverage percentage; superiority of a scanner or methodology; customer outcomes; or market demand.
The search used broad scholarly APIs rather than subscription databases, one reviewer conducted screening and extraction, study methods varied substantially, and publication bias remains possible.
Complete publication
Read the full evidence review and references.
The PDF includes the complete method, synthesis, practical reporting guidance, claims boundaries, limitations, and cited sources.
Download the PDF- Author
- Kevin Hintzman
- Publisher
- Wingspan LLC
- Published
- Document ID
- AUTOMATED-ACCESSIBILITY-EVIDENCE-REVIEW-20260720-R1

