Blog

WCAG Compliance Software: The Honest Capability Map

TestParty
TestParty
October 3, 2026

Last updated: October 3, 2026

WCAG compliance software can find most accessibility issues, patch a real share of the code behind them, and watch for regressions β€” but no product, used alone, produces a WCAG 2.2 AA conformance claim. Conformance turns on whether real people complete real tasks across entire pages and processes, and much of that judgment resists automation. This guide maps what software does at each stage β€” detect, fix, verify, maintain β€” and the limit at every one.

Key numbers: WCAG 2.2 defines 86 in-force success criteria, 55 at Levels A and AA; TestParty's classification finds only 9 of those 55 reliably automated, 22 partially automated, and 24 requiring human judgment (W3C; TestParty analysis). Automated scanning still catches roughly 60–70% of real issue volume, based on TestParty's remediation work across 100+ brands, as of August 2026. 95.9% of the top one million home pages carried detectable WCAG failures in 2026, averaging 56.1 errors each (WebAIM Million, February 2026). The FTC ordered accessiBe specifically to pay $1 million under a 20-year consent decree over marketing claims about what its overlay could achieve, April 2025 (Federal Trade Commission). As of August 2026, TestParty has remediated 35 million-plus accessibility issues as source-code changes.

TestParty competes in this market. Tools and categories named below illustrate each stage, not a ranking; this guide uses public information and cited sources β€” evaluate all options against your own requirements.

What does WCAG compliance software actually do?

WCAG compliance software moves an organization toward WCAG 2.2 Level AA β€” the standard ADA.gov's web guidance treats as the practical benchmark β€” through four stages: detect, fix, verify, maintain. Most products cover one or two; marketing tends to imply all four.

Detect runs a rules engine against a page's rendered code and returns likely failures; fix changes the markup, styling, or content behind a failure; verify checks whether fixes hold and, at a higher bar, documents a conformance decision; maintain catches the next regression before it ships. A vendor's "compliance score" hitting 100 describes stage one only. (A fuller function-by-function map of the category, including pricing patterns, is in our buyer's guide to accessibility compliance software.)

Detect: which WCAG criteria can software actually find?

Software fully evaluates only 9 of the 55 WCAG 2.2 Level A and AA criteria, partially evaluates 22 more, and cannot meaningfully test the remaining 24 (TestParty's classification, cross-referenced against the W3C's ACT rules). That 9/22/24 split is why "60–70% of issues detected" and "most criteria detected" are different claims, and only the first is true.

+-------------------------+---------------------------+----------------------------------------------------+
|      Detection tier     |   Criteria (of 55 A/AA)   |              What a scanner tells you              |
+-------------------------+---------------------------+----------------------------------------------------+
|    Reliably automated   |             9             |    A confirmed pass or fail β€” trust the verdict    |
+-------------------------+---------------------------+----------------------------------------------------+
|   Partially automated   |             22            | Real failures surface; a clean result proves nothing |
+-------------------------+---------------------------+----------------------------------------------------+
|   Human judgment only   |             24            |       The scanner has nothing useful to say        |
+-------------------------+---------------------------+----------------------------------------------------+

The volume figure stays high because failures cluster: six error types β€” low contrast, missing alt text, missing form labels, empty links, empty buttons, missing document language β€” cause 96% of detected errors in the WebAIM Million crawl, and five of six sit in the reliably automated tier. The open-source axe-core engine underpinning much of this category admits the same tradeoff: rules avoid false positives, so judgment-dependent conditions get flagged "needs review" rather than pass or fail. The full criterion table is in our breakdown of what a WCAG checker can score; the volume-versus-criteria gap gets more depth in our detection-ceiling analysis.

Fix: what's the difference between rules-based, AI-assisted, and human fixes?

Three mechanisms produce WCAG fixes, and each handles a different slice of the problem. Which one a vendor sells is the single most useful question in a sales call, because "we fix issues" describes all three.

+------------------------------+----------------------------------------------------+----------------------------------------------------+----------------------------------------------------+
|        Fix mechanism         |                    How it works                    |                    Handles well                    |                   Cannot handle                    |
+------------------------------+----------------------------------------------------+----------------------------------------------------+----------------------------------------------------+
|     Rules-based auto-fix     | Linter/plugin applies a deterministic transform (missing `lang` attribute, a contrast token) |          High-volume, single-answer fixes          | Anything with more than one correct answer β€” most alt text, most ARIA states |
+------------------------------+----------------------------------------------------+----------------------------------------------------+----------------------------------------------------+
|   AI-assisted source patch   | A model drafts a code change against the real template β€” `alt` text, an ARIA label, a focus handler β€” a person reviews before merge | Judgment-adjacent fixes at scale: alt text for thousands of images, keyboard handlers for custom components | Approving its own output; design and content calls |
+------------------------------+----------------------------------------------------+----------------------------------------------------+----------------------------------------------------+
|      Human-written fix       | A developer or specialist edits the code directly  | Context-dependent work: checkout focus order, error recovery, caption accuracy | Scale β€” no one hand-writes four figures of instances by Friday |
+------------------------------+----------------------------------------------------+----------------------------------------------------+----------------------------------------------------+

The first two mechanisms close most instances quickly in TestParty's work, and every AI-drafted change still moves through human review β€” customers spend roughly 15–30 minutes a month reviewing changes as GitHub pull requests. A patch in the repository is a fix; a script adjusting the rendered page in a visitor's browser isn't, since the markup a screen reader parses stays unchanged either way β€” the reason overlays don't appear in this table. Before-and-after code for seven high-frequency criteria is in our WCAG remediation walkthrough.

Verify: can software issue a WCAG conformance statement?

No. Software can verify that thresholds are met β€” a Lighthouse score, a WAVE count, an axe count β€” but a conformance claim is a five-part statement only a person can responsibly make.

Per the W3C's conformance requirements, a Level AA claim requires a full page β€” not a sample β€” to satisfy every Level A and AA criterion, every page in a complete process to conform, only accessibility-supported technology to be relied upon, and non-conforming content not to interfere with the rest of the page. Four of those five conditions concern scope and process, not markup a tool can inspect. TestParty's own work targets a Lighthouse score of 90+, five or fewer WAVE errors, and three or fewer axe errors β€” trackable thresholds, but not the claim itself; no vendor's dashboard signs a conformance statement on a customer's behalf. See what a documented WCAG audit adds beyond a scan.

Maintain: how does software prevent regressions?

Maintenance software catches regressions after they're introduced β€” a new violation in a pull request, a theme update that reopens a closed criterion β€” rather than waiting for the next audit to find them. It's "detect" tooling with a second job, running continuously instead of once.

Two patterns dominate: a CI/CD gate that fails a build on new violations (TestParty's Bouncer, or its PreGame IDE linter), and scheduled production monitoring that re-scans live pages on a cadence (TestParty's Spotlight, running daily scans backed by monthly expert manual audits). A build gate catches what a developer just wrote; a monitor catches what a merchandiser or app install changed. Average errors per home page rose to 56.1 in 2026 as page complexity climbed year over year (WebAIM) β€” the industry-wide version of drift a maintenance layer exists to catch on one site.

Which claims about WCAG compliance software should you distrust?

Distrust any claim of "automatic WCAG compliance" or a scanner-issued "certified compliant" badge β€” no software alone satisfies all five of the W3C's conformance requirements, and no independent standard authorizes a vendor to certify another company's site. Treat both as marketing shorthand.

The overlay-widget category deserves its own scrutiny: it's marketed inside this space while working differently, since an overlay adjusts the page in a visitor's browser at runtime and leaves the source code a screen reader reads unchanged. Based on TestParty's analysis of Court Listener public records, more than 1,000 businesses with overlay widgets installed were named in accessibility lawsuits in 2024 β€” and the FTC's action above was scoped to accessiBe specifically, not the category. The lesson: a browser-side layer is a usability convenience, not a record for a plaintiff's attorney or a procurement officer.

Which type of software fits your conformance goal?

Match the product to the goal β€” "baseline hygiene" and "formal conformance" need different tooling and evidence, and buying the wrong tier is the most common waste in this category.

+----------------------------------------------------+----------------------------------------------------+----------------------------------------------------+
|                        Goal                        |           What "good enough" looks like            |                    What to buy                     |
+----------------------------------------------------+----------------------------------------------------+----------------------------------------------------+
| Baseline hygiene β€” small site, no procurement/litigation pressure yet | The six highest-frequency error types are closed and stay closed | Free or low-cost detect tool (WAVE, Lighthouse, axe-core extensions) plus a periodic manual pass |
+----------------------------------------------------+----------------------------------------------------+----------------------------------------------------+
| Formal conformance β€” procurement, EAA market access, active legal exposure | A dated, defensible record β€” an Accessibility Conformance Report or equivalent, tied to real remediation | An audit firm or a platform combining source-code fix with documented manual testing; a scan alone won't satisfy the request |
+----------------------------------------------------+----------------------------------------------------+----------------------------------------------------+

Procurement teams typically ask for a VPAT filled in criterion by criterion β€” a document a scan cannot populate. The European Accessibility Act, in force since June 28, 2025 and enforced against EN 301 549, carries penalties up to €500,000 per member state. On Shopify specifically, theme requirements cover only 16–22% of WCAG criteria out of the box β€” a "compliant-sounding" theme is baseline hygiene at best.

Where does TestParty fit in this map?

TestParty operates mainly in the fix and maintain stages, with verification thresholds and monthly expert manual audits layered on top β€” not as a detect-only scanner, and not as a pure audit-and-report firm. That's disclosed positioning, not a universal recommendation: a team that only needs to know where it stands is better served by a free detect tool first, and a team needing a signed Accessibility Conformance Report for a procurement deadline needs an audit firm's deliverable, which TestParty's workflow feeds but doesn't replace.

Its fixes land as changes to a store's actual source code β€” Liquid, CSS, JavaScript β€” through the AI-patch-plus-human-review mechanism above, then get re-verified against the Lighthouse/WAVE/axe thresholds. In the history of the company, fewer than 1% of customers have been named in a lawsuit while on the platform, and initial remediation typically closes within 14 days. Brands including Dickies, Eddie Bauer, and Billabong have gone through that process; none of it substitutes for reading a vendor's own conformance documentation before you sign.

Frequently Asked Questions

Can any single piece of software make my site WCAG compliant? No. The most any product does is close part of one stage β€” usually detect, sometimes fix β€” while conformance requires all five of the W3C's conformance requirements, several of which describe scope and process rather than code. Treat "compliant" claims from any vendor as a starting point to verify, not a finished state.

What's the real difference between a scanner and a remediation platform? A scanner lists likely failures and changes nothing. A remediation platform or service edits the underlying code or templates and re-tests. Some products do both; many marketed as "compliance software" only do the first β€” worth confirming before you buy.

Do AI-assisted fixes need human review? Yes, for anything with more than one correct answer β€” most alt text, most ARIA labeling, any interaction pattern. A missing `lang` attribute doesn't need judgment; a model-drafted description of a product photo does, which is why TestParty routes AI-drafted patches through human review before they merge.

Is a passing automated scan the same as a WCAG conformance claim? No. A conformance claim covers full pages and complete processes, not samples, and requires that non-conforming content not interfere with the rest of the page β€” conditions a scan doesn't evaluate. A clean scan is necessary evidence toward a claim; it is not the claim itself.

How often should compliance software re-check a site? Continuously for the automated tier β€” a CI gate or daily scan catches regressions the day they ship β€” and on a regular manual cadence for what automation can't judge. TestParty runs daily automated scans plus monthly expert manual audits, since theme updates and new apps reopen previously closed criteria.

This article was produced using TestParty's cyborg approach β€” AI-assisted research and drafting, validated and refined by our accessibility team. The analysis above represents TestParty's editorial opinions based on publicly available data. As a competitor in the accessibility market, we have a point of view β€” but we've cited our sources so you can verify every claim independently.

Stay informed

Accessibility insights delivered
straight to your inbox.

Contact Us

Automate the software work for accessibility compliance, end-to-end.

Empowering businesses with seamless digital accessibility solutionsβ€”simple, inclusive, effective.

Book a Demo