Blog

Manual vs Automated Accessibility Testing: What Each Catches, Misses, and Costs

TestParty
TestParty
September 24, 2026

Last updated: September 24, 2026

Automated accessibility testing finds roughly 60–70% of the issues on a typical ecommerce site in minutes; manual testing finds the rest β€” and the rest includes most of what gets you sued. What follows is which method does which job, what each costs, who should perform it, and how credible programs combine both.

Key numbers: Automated scanning detects roughly 60–70% of accessibility issue volume; the remaining 30% requires expert human testing (TestParty remediation data across 100+ brands, as of August 2026). Per W3C, evaluation tools "can not determine accessibility, they can only assist in doing so," and "human judgement is required" (W3C WAI). In 2026, 95.9% of the top million home pages had detectable WCAG failures, averaging 56.1 errors per page (WebAIM Million). Plaintiffs filed 3,117 federal website accessibility suits in 2025, up 27%, and ecommerce is named in 69–77% of digital accessibility suits (Seyfarth Shaw).

TestParty competes in this market. This comparison uses public information and cited sources; evaluate all options against your own requirements.

Which finds more, manual or automated accessibility testing?

Automation finds more issues; manual testing finds more important issues β€” and conflating the two is the most expensive mistake in accessibility programs. A scanner returns hundreds of repeating findings on an untested site; a human tester walks the checkout as a blind shopper does and returns twelve, four of which stop the sale.

+-------------------------+----------------------------------------------------+----------------------------------------------------+
|        Dimension        |                 Automated testing                  |                   Manual testing                   |
+-------------------------+----------------------------------------------------+----------------------------------------------------+
|          Speed          |      Seconds per page; a full site overnight       |    Hours per template; days to weeks per audit     |
+-------------------------+----------------------------------------------------+----------------------------------------------------+
|           Cost          | Free to low monthly subscriptions; barely rises with page count | $100–$250 per primary page on published rate cards (August 2026) |
+-------------------------+----------------------------------------------------+----------------------------------------------------+
|    Coverage β€” volume    | ~60–70% of detected issues (TestParty data, August 2026) | The remaining ~30%, plus confirmation of automated findings |
+-------------------------+----------------------------------------------------+----------------------------------------------------+
|   Coverage β€” criteria   | A minority of the 55 Level A and AA criteria in WCAG 2.2 evaluated end to end | Every applicable criterion, including judgment-based ones |
+-------------------------+----------------------------------------------------+----------------------------------------------------+
|      Repeatability      |  Deterministic; identical input, identical output  | Varies by tester and assistive technology; needs a documented method |
+-------------------------+----------------------------------------------------+----------------------------------------------------+
|     What it catches     | Contrast ratios, missing alt attributes, unlabeled inputs, empty links and buttons, broken ARIA references | Whether alt text means anything, whether focus order makes sense, whether a screen reader user can finish checkout |
+-------------------------+----------------------------------------------------+----------------------------------------------------+
|      What it misses     | Meaning, sequence, equivalence, anything behind an interaction no script triggered | Nothing structurally β€” but you cannot run it on every page, every day |
+-------------------------+----------------------------------------------------+----------------------------------------------------+
|       When to run       | Every pull request, every deploy, daily in production | Pre-launch, on a fixed cadence, at major releases, after a demand letter |
+-------------------------+----------------------------------------------------+----------------------------------------------------+

What does automated testing catch reliably?

Machine-checkable conditions that repeat across templates: contrast ratios, missing alt attributes, unlabeled inputs, empty links and buttons, missing document language, and ARIA references pointing nowhere.

That list is short, but it is where the volume lives. WebAIM's 2026 crawl found six error types account for roughly 96% of all detected errors across a million home pages β€” each a condition a rule engine was built to evaluate, each repeating across templated pages. The open-source axe-core rule library publishes its rules with WCAG mappings and an issue type of either "failure" or "needs review" β€” which makes the boundary auditable.

The leverage is structural: a store's thousands of violations usually resolve to a couple of dozen components, and one template fix clears thousands of instances. In TestParty's audit work a default Dawn theme ships with 30–100 detectable violations and premium themes 100–350, before any app is installed. For the mechanics, see what automated WCAG testing catches and misses.

What can only a human find?

Five failure classes, all of them about meaning rather than markup β€” and all of them capable of surviving a scan that reports zero errors.

  • Meaningful alternative text (1.1.1 Non-text Content). `alt="IMG_4471"` passes every rule in every engine, as does alt text describing a product photo accurately while omitting the colorway the image exists to show. WCAG asks for an equivalent purpose, and equivalence is a judgment.
  • Focus order logic (2.4.3 Focus Order, 2.1.2 No Keyboard Trap, 2.4.11 Focus Not Obscured). A scanner confirms focusable elements exist. Whether focus enters an opened cart drawer β€” and whether it can get back out β€” is answered only by pressing Tab.
  • Screen-reader flow quality (1.3.2 Meaningful Sequence, 4.1.2 Name Role Value, 4.1.3 Status Messages). DOM order, visual order, and reading order are three different things; a cart total that updates silently is valid HTML and a broken experience.
  • Cognitive load and error recovery (3.3.2 Labels or Instructions, 3.3.3 Error Suggestion, 3.2.6 Consistent Help, 3.3.7 Redundant Entry). Whether an error message tells someone how to fix the problem is not something a machine reads off a page.
  • Caption accuracy (1.2.2 Captions). Machine captions get brand and product names wrong at exactly the moments they matter; a transcript that is 95% correct fails in the 5%.

In TestParty's monthly expert audits, the most common finding automation never reports is the variant selector: swatches are focusable and labeled, the scan is clean, and a screen reader announces nothing when the selection changes.

Because plaintiffs test the way users do, not the way scanners do. The barriers pleaded in website accessibility complaints are overwhelmingly experiential β€” a screen reader that cannot complete checkout, a control a keyboard cannot reach.

Volume makes this a business problem: 3,117 federal website accessibility suits were filed in 2025, up 27%, and ecommerce is named in 69–77% of digital accessibility suits (Seyfarth Shaw). Across the demand letters TestParty has reviewed, the recurring allegations cluster into a short list β€” screen-reader incompatibility, keyboard navigation failures, missing or meaningless alternative text, unlabeled form fields, and low contrast.

Read that list carefully. Three of the five are things automation catches β€” which is why a program running scans is already ahead. What survives a clean scan and still gets pleaded is the experiential set: keyboard operation and screen-reader flow. That is the honest case for a hybrid program. Evidence of repair carries weight too: in public TestParty matters, Dorai Home resolved a $74,999 demand for $2,000 after documented remediation, and the Joanna Vargas matter was dismissed at $0.

What does a hybrid testing program actually look like?

Two operating models work, and both pair continuous machine coverage with human testing on a fixed cadence. The difference is where the automation lives.

Model one β€” continuous scanning plus scheduled expert audits. Automated scans run daily across every template and state; expert manual audits run monthly against the criteria automation cannot reach. This is TestParty's model: initial remediation runs on a 14-day cycle, fixes ship as pull requests, and customers typically spend 15–30 minutes a month reviewing them. Our post-remediation targets (Lighthouse 90+, WAVE errors at 5 or fewer, axe errors at 3 or fewer) are thresholds to clear, not certificates β€” they measure the automatable slice only.

Model two β€” pipeline gates plus periodic assistive-technology testing. Checks run on every pull request so regressions never merge, with a keyboard and screen-reader pass at each major release and a formal conformance evaluation annually. This suits teams with mature CI/CD.

Either way, the cadence gap is the failure point. Regressions arrive through routine work β€” a theme update, a new app, a CMS edit β€” which is why continuous accessibility monitoring matters more than audit frequency, and why a formal WCAG conformance audit is an anchor, not the whole program.

What does each method cost?

The two methods scale differently: automation's cost stays nearly flat as pages grow, while manual cost rises linearly with templates and flows tested.

Automated tooling starts at zero β€” WAVE, Lighthouse, axe extensions, Pa11y, and IBM Equal Access are all free, and our roundup of free accessibility checkers covers what each surfaces. Commercial scanning runs from roughly $20 per month for a single site to five figures a year inside enterprise governance bundles β€” the same scan, priced by the buying process (Accessible.org, June 2026).

Manual testing is priced by the hour. Published rate cards list roughly $100–$250 per primary page; audits for small and mid-sized sites run $1,500–$5,000, full manual WCAG 2.2 AA audits for mid-size ecommerce $3,000–$15,000+, and complex enterprise applications $10,000–$25,000+ (published agency rates as of August 2026). Sampling is the cost lever: 12 templates plus every step of cart-to-confirmation buys most of the value of testing 400 pages. See how professional accessibility audits are scoped and run.

Who should perform the manual testing?

Trained auditors working from a documented method, ideally alongside people who use assistive technology daily. Manual testing without either produces a false clean bill β€” worse than no test.

Credentials are the shortcut for evaluating a vendor or a hire: IAAP certifications β€” CPACC for foundational knowledge, WAS for hands-on technical auditing, CPWA for both β€” plus DHS Trusted Tester for Web, validated against the federal Section 508 process. Certified practitioners are scarce; an agency answering by describing a partnership is being straight.

The assistive-technology matrix should follow real usage: JAWS is the primary screen reader for 40.5% of users and NVDA for 37.7% (WebAIM Screen Reader User Survey #10). Your own team can do useful work without hiring anyone β€” NVDA is free on Windows, VoiceOver ships with every Mac, and fifteen minutes completing a purchase with one tells you more than any dashboard.

Which testing should you run right now?

The answer depends less on budget than on what happens next week. Four scenarios cover most cases.

Pre-launch or mid-redesign. Automated first β€” template failures are cheapest to fix before content is poured in β€” then one manual pass on the revenue flows. A keyboard trap in staging costs a ticket; in production it costs a release.

Steady state. Automation continuously, manual on a cadence. Per-deploy scanning catches regressions; monthly or quarterly expert testing catches what scanning cannot. One annual audit on a store that publishes weekly is archaeology.

After a demand letter. Manual, urgently, on the flows named in the complaint β€” then automated everywhere else. Dated evidence of remediation changes a settlement conversation, and the cited barriers are usually ones a scan will not reproduce.

Procurement, VPAT, or EAA scope. Full manual evaluation under a documented methodology, with automation for breadth across the sample. A conformance claim covers all applicable criteria, and no scanner can make it. Shopify merchants get caught here: Theme Store requirements map to roughly 16–22% of WCAG success criteria (TestParty analysis), and nobody in the chain reviews third-party apps for accessibility.

Frequently Asked Questions

Can automated testing alone make a site WCAG compliant? No. Conformance is evaluated criterion by criterion, and a tool that fully evaluates a minority of the 55 Level A and AA criteria cannot establish it. A zero-error scan means the automatable checks passed β€” it says nothing about reading order, caption quality, or whether a keyboard-only shopper can buy.

How much of WCAG can actually be tested automatically? Two different numbers, both accurate. Automation detects roughly 60–70% of issue volume, because failures cluster in a few high-frequency types repeating across templates. It fully evaluates only a minority of success criteria, since the standard asks things like whether a heading describes its section β€” questions markup cannot answer.

How often should manual testing run? Monthly for stores that publish frequently, quarterly at minimum, plus before major releases and after any demand letter. TestParty pairs daily automated scans with monthly expert audits because the two cadences solve different problems: scanning catches regressions fast, human testing catches what scanning cannot see.

Is a clean automated scan worth anything if we get sued? Dated scan records help demonstrate an ongoing remediation effort, materially different from an unmonitored site. A clean report is not a defense on its own, since plaintiffs' testing surfaces keyboard and screen-reader failures no scanner reports. Continuous testing paired with shipped fixes carries the weight.

If we can only fund one method this quarter, which comes first? Automation. It costs the least, covers every page, and clears the high-volume failures that make manual testing expensive by burying real findings in noise. Then spend the manual budget narrowly: one expert pass on cart-to-confirmation buys more risk reduction per dollar than a site-wide audit.

Like everything at TestParty, this article reflects our cyborg philosophy: AI handles the heavy lifting, humans bring the expertise. The data and opinions here are based on publicly available sources as of publication. TestParty is a participant in the accessibility market β€” we believe in transparency, so we encourage you to cross-reference our claims and evaluate all options for your business.

Stay informed

Accessibility insights delivered
straight to your inbox.

Contact Us

Automate the software work for accessibility compliance, end-to-end.

Empowering businesses with seamless digital accessibility solutionsβ€”simple, inclusive, effective.

Book a Demo