Continuous Accessibility Testing: The Full Pipeline View, Stage by Stage
TABLE OF CONTENTS
- What is continuous accessibility testing?
- Where in the pipeline do accessibility checks actually run?
- Why does catching a violation earlier cost less?
- What tooling runs at each stage?
- How do you gate builds without blocking shipping?
- What does continuous automated testing still miss?
- Which regressions does continuous testing actually prevent?
- Where does TestParty's stack fit?
- How do you start? The minimal viable pipeline
- Frequently Asked Questions
Last updated: October 1, 2026
Continuous accessibility testing runs automated WCAG checks at every stage where code changes β in the editor, in pull requests, in CI/CD, on staging, and in production β so violations are caught when they are cheapest to fix rather than at the annual audit. It is the difference between treating accessibility as an event and treating it as a property of your delivery pipeline.
Key numbers: In 2026, 95.9% of the top one million home pages had detectable WCAG failures, averaging 56.1 errors per page β up 10.1% from 51 errors in 2025, reversing six years of gradual improvement (WebAIM Million). Automated testing detects roughly 60β70% of accessibility issue volume; the remaining ~30% requires expert human review (TestParty remediation data across 100+ brands, as of August 2026). The open-source axe-core rule library maps 74 rules to WCAG 2.0β2.2 Level A and AA, plus 27 best-practice rules, against the 55 Level A and AA success criteria in WCAG 2.2. Plaintiffs filed 3,117 federal website accessibility lawsuits in 2025, up 27% year over year, with ecommerce named in 69β77% of digital accessibility suits (Seyfarth Shaw). As of August 2026, TestParty has remediated more than 35 million accessibility issues across customer stores.
What is continuous accessibility testing?
Continuous accessibility testing is the practice of running automated WCAG checks at every stage a codebase changes β authoring, review, build, pre-release, and production β rather than at a scheduled audit. Each stage catches a different class of defect, and each catches it at a different price.
The word doing the work is continuous, in the same sense as continuous integration. A point-in-time audit answers "was this site conformant on the day we looked?" A continuous program answers a more useful question: "did the change we are about to ship make it worse?" A machine answers that in seconds, and it is the question that keeps a remediated site remediated.
Where in the pipeline do accessibility checks actually run?
Five stages, running from cheapest to most expensive: editor linting, pull request checks, CI/CD gates, staging scans, and production monitoring. A mature program runs all five; most teams run one and assume it is enough.
+---------------------------+----------------------------------------------------+----------------------------------------------------+----------------------------------------------------+
| Stage | What runs there | What it catches | Cost to fix at this stage |
+---------------------------+----------------------------------------------------+----------------------------------------------------+----------------------------------------------------+
| Editor / pre-commit | Accessibility linters β `eslint-plugin-jsx-a11y`, template linters, IDE rules, commercial linters such as axe Linter or TestParty's PreGame | Static markup errors: missing `alt`, unlabeled inputs, invalid ARIA attributes, non-interactive elements with handlers | Seconds. The developer is already in the file with the context loaded |
+---------------------------+----------------------------------------------------+----------------------------------------------------+----------------------------------------------------+
| Pull request | axe-core in unit/component tests; a diff-scoped scan posting results as a PR comment | Component-level regressions: a shared button losing its accessible name, a modal without focus management | Minutes, before merge, while the change is still one person's problem |
+---------------------------+----------------------------------------------------+----------------------------------------------------+----------------------------------------------------+
| CI/CD gate | axe-core driven by @axe-core/cli, @axe-core/playwright, or @axe-core/puppeteer; Accessibility Insights scans | Rendered-DOM failures across key routes: contrast, heading structure, landmark issues, ARIA references that resolve to nothing | Hours. A failed build, fixed the same day, before anything reaches users |
+---------------------------+----------------------------------------------------+----------------------------------------------------+----------------------------------------------------+
| Staging / pre-release | Full-site crawl plus scripted interaction states β cart, checkout, filters, modals, logged-in views | Integration failures that only appear when real components, data, and third-party scripts meet | A ticket in the current sprint, plus QA re-test |
+---------------------------+----------------------------------------------------+----------------------------------------------------+----------------------------------------------------+
| Production | Scheduled crawls, regression diffs against a known-good baseline, alert routing | Everything you do not control: CMS edits, app installs, theme updates, marketing tags, expired content | Days to weeks, on a live site, with real users affected in the interim |
+---------------------------+----------------------------------------------------+----------------------------------------------------+----------------------------------------------------+
| (Not caught) | Nothing | Found by a plaintiff's tester, a screen reader user, or opposing counsel | An audit finding, a demand letter, or a settlement |
+---------------------------+----------------------------------------------------+----------------------------------------------------+----------------------------------------------------+The first two rows are where the cost curve is flattest and where most teams have nothing installed. The last row is not a stage β it is what happens when the other five are missing.
Why does catching a violation earlier cost less?
The same missing form label costs a developer perhaps thirty seconds in the editor, an hour once it is in CI, a sprint ticket after release, and β in the worst case β legal fees. The defect does not change; the number of people and processes touching it does.
Cost escalates for three reasons. Context evaporates: the developer who wrote the component knows why the label is missing; the one handed a remediation ticket four months later does not. Blast radius grows: one unlabeled input in a shared component becomes 400 failures once it renders across a catalog. And the fix mechanism changes β an editor fix is a keystroke, a production fix is a code change, a review, a QA pass, and a deploy. Once a violation surfaces in an external audit, published agency remediation rates run roughly $100β$250 per primary page (agency rate cards as of August 2026); once it surfaces in a demand letter, the number is set by counsel. One public example from TestParty's customer base: Dorai Home received a $74,999 demand and, with documented source-code remediation, settled for $2,000.
What tooling runs at each stage?
The tooling is largely free and open source, and one rules engine β axe-core β powers most of it, so findings stay consistent as code moves from editor to production. What changes stage to stage is what the engine can see.
A linter sees source text: it catches static errors before the code runs but cannot evaluate contrast or computed accessible names. A test-suite check sees a rendered component in isolation; a CI gate sees a fully rendered page; a production crawler sees the page a customer actually gets, third-party scripts included. A minimal CI check using Playwright:
// Fails the build on serious/critical WCAG 2.2 A/AA violations only.
// Covers machine-checkable criteria: 1.1.1 alt text, 1.4.3 contrast,
// 3.3.2 form labels, 4.1.2 name/role/value.
import { test, expect } from '@playwright/test';
import AxeBuilder from '@axe-core/playwright';
test('product page has no serious accessibility violations', async ({ page }) => {
await page.goto('/products/example');
const results = await new AxeBuilder({ page })
.withTags(['wcag2a', 'wcag2aa', 'wcag21aa', 'wcag22aa'])
.analyze();
const blocking = results.violations.filter(v =>
['serious', 'critical'].includes(v.impact)
);
expect(blocking, JSON.stringify(blocking, null, 2)).toHaveLength(0);
});Rule consistency across stages matters more than tool choice. The W3C's ACT Rules Format 1.1, a Recommendation as of February 5, 2026, standardizes how accessibility test rules are written and compared β which is what makes "the same rule, everywhere in the pipeline" a realistic goal rather than a slogan. For the wiring details on GitHub Actions, GitLab, and CircleCI, see our walkthrough on adding accessibility checks to a build pipeline.
How do you gate builds without blocking shipping?
Gate on severity and on new violations only. A gate that fails every build on day one because the existing codebase has 4,000 legacy findings will be disabled within two weeks β and a disabled gate catches nothing.
Three rules make gating survivable:
- Severity thresholds. Fail on `critical` and `serious`; report `moderate` and `minor` as warnings. This maps roughly to what actually blocks a user from completing a task versus what degrades the experience.
- New-violations-only baselines. Snapshot current findings as a baseline file committed to the repo, then fail only on violations absent from that baseline. Legacy debt gets its own remediation track instead of holding releases hostage β the approach we outline in our WCAG remediation workflow.
- Expiring allowlists. Any suppressed rule carries an owner and a date. Without expiry, the allowlist becomes the codebase.
The failure mode to design against is alert fatigue. In our assessment, the most common reason continuous accessibility testing programs die is not false negatives but volume: a pipeline that emits hundreds of undifferentiated findings per build trains engineers to ignore it, which is strictly worse than having no gate at all because it also produces a paper record of ignored warnings.
What does continuous automated testing still miss?
Roughly 30% of accessibility issue volume β and a disproportionate share of the issues that actually stop a purchase. No amount of pipeline coverage changes the ceiling of what a rules engine can evaluate.
Automation cannot tell you whether alt text is meaningful, whether focus order follows the visual reading order, whether an error message explains how to fix the problem, or whether a screen reader user can complete checkout end to end. Those are judgment calls mapping to criteria no engine fully evaluates β see what automated scanners actually detect for the criteria-level breakdown, and where manual testing beats automation for the failure classes only humans find. The working model: continuous automation for regression prevention at every commit, plus expert manual audits on a recurring cadence and before major releases. Continuous testing does not replace the audit; it shrinks what the audit has to find.
Which regressions does continuous testing actually prevent?
The ones a remediation project cannot: theme updates that overwrite fixed templates, app installs that inject unreviewed markup sitewide, and content that ships without alt text. These are why sites regress after being fixed.
The market-level evidence is the WebAIM Million reversal: average detected errors per home page rose from 51 in 2025 to 56.1 in 2026, and the same six error types have topped the list for seven consecutive years. Sites are not failing at new things; they are re-failing at old things, because the fixes were point-in-time and the codebases were not. On ecommerce specifically, Shopify's theme requirements cover only 16β22% of WCAG success criteria, and Shopify does not review third-party apps for accessibility before App Store approval (TestParty analysis; Shopify developer documentation). In TestParty's daily scans of customer stores, app-injected code and new product content are the two most common sources of fresh violations β neither of which passes through a pipeline you control, which is why the production stage exists alongside the CI stage. Our guide to ongoing accessibility monitoring covers that stage in depth.
Where does TestParty's stack fit?
Disclosure: TestParty builds tooling in this category, so treat this as vendor description rather than recommendation. Our products cover three of the five stages β PreGame for editor and pre-commit linting, Bouncer for CI/CD gates, Spotlight for scheduled production scanning with regression diffs β paired with monthly expert manual audits for the ~30% automation cannot reach. Fixes arrive as source-code pull requests rather than dashboard tickets; customers typically spend 15β30 minutes a month reviewing them. The pipeline model in this article is not proprietary, though: every stage above can be assembled from open-source axe-core packages, and teams with engineering capacity routinely do exactly that.
How do you start? The minimal viable pipeline
Start with one CI gate on your five highest-traffic routes, configured to fail only on new critical and serious violations. That single change stops the bleeding; everything else is refinement.
+------------------+---------------------------------------------------+----------------------------------------------------+
| Stage | Small team (1β5 developers) | Enterprise / multi-team |
+------------------+---------------------------------------------------+----------------------------------------------------+
| Editor | `eslint-plugin-jsx-a11y` in the shared config | Linter rules enforced in the shared design-system repo |
+------------------+---------------------------------------------------+----------------------------------------------------+
| Pull request | Skip initially | Diff-scoped scan posting results as a PR comment |
+------------------+---------------------------------------------------+----------------------------------------------------+
| CI/CD | axe-core on 5 key routes; new-violations-only | Full route matrix, per-team baselines, dashboards by owner |
+------------------+---------------------------------------------------+----------------------------------------------------+
| Staging | Manual scan before major releases | Automated crawl with scripted interaction states |
+------------------+---------------------------------------------------+----------------------------------------------------+
| Production | Weekly scheduled scan | Daily scans, severity-routed alerts, deploy-linked diffs |
+------------------+---------------------------------------------------+----------------------------------------------------+
| Human review | Expert audit before major launches | Recurring expert audits plus assistive-technology user testing |
+------------------+---------------------------------------------------+----------------------------------------------------+Sequence matters more than completeness. Teams that install all five stages in one sprint generate an unmanageable finding count and abandon the program; teams that add one stage per quarter, starting at CI, tend to keep it. For the full evaluation methodology behind those expert reviews, see our accessibility audit process guide.
Frequently Asked Questions
What is the difference between continuous accessibility testing and accessibility monitoring? Monitoring is one stage of continuous testing. Monitoring watches the live production site for drift caused by content edits, app installs, and theme updates. Continuous testing covers the whole pipeline β editor, pull request, CI, staging, and production β so most violations never reach production to be monitored in the first place. Programs need both; the pipeline stages prevent what you control, monitoring catches what you do not.
Should an accessibility check fail the build or just warn? Fail on new critical and serious violations; warn on everything else. A gate that blocks releases over minor findings gets disabled, and a gate that never blocks anything gets ignored. The baseline approach resolves the tension: existing violations are recorded and tracked separately, while any newly introduced serious violation blocks the merge that introduced it.
How much does continuous accessibility testing slow down CI? An axe-core scan of a rendered page typically completes in a few seconds, so a five-route check adds well under a minute to a build. Full-site crawls are the expensive operation and belong on a schedule or a staging trigger, not on every commit. If accessibility checks are meaningfully slowing your pipeline, the scope is wrong, not the tooling.
Does continuous automated testing mean we can skip manual audits? No. Automated testing across every pipeline stage still detects roughly 60β70% of issue volume (TestParty remediation data, as of August 2026), and the remainder includes most of what blocks a real purchase β focus order, meaningful alt text, screen reader flow through checkout. Continuous testing reduces how much an audit has to find; it does not eliminate the need for one.
Can continuous testing catch issues introduced by third-party apps and scripts? Only at the production stage. Third-party scripts and platform apps never pass through your build, so CI cannot see them. This is the structural argument for keeping scheduled production scans even after your pipeline is well instrumented β on ecommerce platforms in particular, unreviewed app code is one of the most common sources of new violations.
This article was produced using TestParty's cyborg approach β AI-assisted research and drafting, validated and refined by our accessibility team. The analysis above represents TestParty's editorial opinions based on publicly available data. As a competitor in the accessibility market, we have a point of view β but we've cited our sources so you can verify every claim independently.
Stay informed
Accessibility insights delivered
straight to your inbox.


Automate the software work for accessibility compliance, end-to-end.
Empowering businesses with seamless digital accessibility solutionsβsimple, inclusive, effective.
Book a Demo