Back to BlogQuality & Testing

Regression Testing on Every Deploy Without Slowing Releases

Rupak Amin

Founder & Lead Engineer, RAITHub

10 min read

To regression test every deploy without slowing releases, split the suite into tiers. Run selected fast tests on each pull request and the full suite before deploy. Run smoke tests against production right after each deploy. Where traffic allows, send a small share of users to the new version first, as a canary. Give each tier a time budget, and let any failure stop the rollout.

This is the process for a team that already has tests and wants them on every deploy without waiting an hour for each one. If you have almost no tests yet, build the first suite and gates with how to set up QA for an early-stage SaaS, then come back here.

What does "regression testing after every deploy" actually mean?

Two different jobs that are often mixed up. Regression testing checks that behaviour which used to work still works after a change. Post-deploy verification checks that the deployed system, with its real configuration, database and third-party services, is healthy. The first mostly happens before the deploy, in CI. The second can only happen after it, against the environment itself.

A team that only does the first ships bugs caused by environment differences: a missing variable, a migration that ran differently, a third-party key for the wrong mode. A team that only does the second finds regressions after customers do. You need both, sized so neither slows the release.

How should you split a regression suite into tiers?

By when each tier runs and how long it may take. The budget is what keeps releases fast. When a tier goes over it, you fix the tier. You don't wait longer.

TierWhen it runsWhat runsTime budget (guide)On failure
1. Pull requestEvery push to a pull requestType-check, lint, unit tests, tests selected by changed files, every test tagged criticalUnder 10 minutesMerge blocked
2. Main branchAfter merge, before deployThe full unit, integration and end-to-end suite, in parallelUnder 20 minutesDeploy blocked
3. Post-deployRight after each deploySmoke tests against the live environmentUnder 5 minutesRoll back or stop the rollout
4. CanaryDuring a staged rolloutError rate and latency of the new version compared with the oldMinutes to an hour, by trafficAutomatic rollback
5. ScheduledNightlySlow suites: more browsers, load checks, long data scenariosNo hard limitTicket for the next day

The time budgets are RAITHub's working guide for small and mid-sized products, not an industry benchmark. The principle holds at any size: the earlier the tier, the faster it must be, because it runs most often and someone is waiting on it.

How do you choose which tests run on a pull request?

Run what the change could affect, plus a fixed set that always runs. Both main test runners can select by changed files. Vitest's --changed option runs tests affected by changed files and accepts a commit or branch to compare against (Vitest CLI). Playwright's --only-changed runs only test files that changed between HEAD and a given ref, and works only with Git (Playwright command line).

# Tier 1: selected tests on a pull request
npx vitest run --changed origin/main
npx playwright test --only-changed=origin/main
npx playwright test --grep "@critical"

Selection has a blind spot. It follows imports, so it misses changes that reach tests some other way: configuration, environment variables, database migrations, feature flags, shared fixtures. Three rules cover the gap:

  • Always run the critical set. Tag the tests for sign-in, payments, permissions and the core action, and run them on every pull request whatever changed.
  • Run everything when shared things change. If a pull request touches configuration, migrations, lockfiles or test setup, skip selection and run the full suite.
  • Run the full suite before deploy. Tier 2 exists so that anything selection missed is caught before production, not after.

When the full suite gets slow, split it across machines rather than trimming it. Both tools support sharding in the form --shard=1/4 (same two sources). Four shards of five minutes beat one twenty-minute job.

What should post-deploy smoke tests check?

That the live system works end to end, using its real configuration. Keep the set short (five to ten tests) and aim it at what CI cannot see:

  1. The app responds on its main pages and API health endpoint, with the expected status codes.
  2. A test account can sign in. This proves session, auth provider and database settings in one step.
  3. The core action completes for a test account, with data written and read back.
  4. Payments are in the right mode. Production keys in production, and the webhook endpoint reachable. Never run a real card in an automated test.
  5. Background jobs are running, checked with a heartbeat or last-run timestamp the job writes.

Trigger the run from the deploy itself. With a host that reports deployments to GitHub, the deployment_status event starts a workflow when a deployment status comes in. It does not fire for the inactive state (GitHub: events that trigger workflows).

# .github/workflows/post-deploy-smoke.yml
name: post-deploy-smoke
on: deployment_status
jobs:
  smoke:
    if: github.event.deployment_status.state == 'success'
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 22
          cache: npm
      - run: npm ci
      - run: npx playwright install --with-deps chromium
      - run: npx playwright test --grep "@smoke"
        env:
          BASE_URL: ${{ github.event.deployment_status.environment_url }}
          SMOKE_USER_PASSWORD: ${{ secrets.SMOKE_USER_PASSWORD }}

Point use.baseURL in playwright.config.ts at process.env.BASE_URL, and give the smoke account its own data so it never touches a real customer's records. Decide ahead of time what a failure means. For most teams, the rule is roll back first, investigate second.

When is a canary release worth it?

When you have enough traffic that a small slice gives a real signal, and a deploy platform that can split traffic. Google's SRE workbook defines canarying as "a partial and time-limited deployment of a change in a service and its evaluation" (Google SRE workbook: canarying releases). The new version serves a small share of users while the old version serves the rest, and you compare the two.

The workbook's advice on what to compare is practical. Rank metrics by how well they show problems users would notice, and use only the top few (perhaps no more than a dozen). HTTP status codes and response latency come before infrastructure numbers like CPU. It also strongly advises running only one canary at a time.

SignalCompareTypical action
Error rateShare of 5xx responses, canary vs the current versionRoll back if clearly higher
LatencyP95 response time on key routesPause if clearly slower
Business eventsCheckouts, sign-ups or core actions completed per requestRoll back if they drop
Client errorsFront-end exceptions reported per sessionPause and inspect

This section is engineering guidance. RAITHub has no published canary rollout to point to. For most small and mid-sized products, post-deploy smoke tests and fast rollbacks give much of the signal a canary would, at lower cost. Low-traffic products get little from a canary, because a one percent slice may see only a handful of requests before it has to be judged. For them, a feature flag that turns a risky change on for internal users first does a similar job.

How do you keep regression testing fast as the suite grows?

Treat pipeline time as a product metric, with an owner. Suites slow down one test at a time, and nobody notices until people start skipping them.

  • Push tests down the pyramid. An end-to-end test that really checks a pricing rule belongs in a unit test that runs in milliseconds.
  • Shard and parallelise before you delete anything.
  • Quarantine flaky tests with a deadline, never with open-ended retries. A test that fails at random teaches the team to ignore red builds. The flaky end-to-end tests guide covers the method.
  • Add a test with every bug fix, at the lowest level that reproduces it, so the suite grows where the product actually breaks.
  • Review the slowest ten tests each month. Most suites have a few tests that take a large share of the time.

For choosing the browser tool behind tiers 1 to 3, see Playwright vs Cypress in 2026.

Which numbers show the process is working?

Two DORA measures, plus your own pipeline time. DORA defines change failure rate as the ratio of deployments that require immediate intervention following a deployment, and failed deployment recovery time as how long it takes to recover from one (DORA metrics guide). A working regression process pushes the first down. Post-deploy smoke tests and a rollback rule push the second down. Track tier 1 and tier 2 durations next to them. If deploys get safer but slower every month, the process will be bypassed eventually.

For scale, PropDesk, the property management platform RAITHub built, has 1,024 automated tests run in CI: 932 unit and 92 end to end. TheSkinProof, the founder's own marketplace venture rather than a client project, has 750+ tests across 217 API endpoints. The PropDesk mix is the point: roughly ten unit tests for every end-to-end test keeps a large suite fast. The full method is in how RAITHub tests software.

Why RAITHub for this

RAITHub ships every product with a gated pipeline, and a failing tier blocks the release. The QA and test automation service sets up this tiered process on an existing product. It adds tags and selection for pull requests, shards the full suite, writes the post-deploy smoke set and the rollback rule, and hands it all over in your repository. The quote is fixed and written after a free 15-minute technical audit, and the IP is yours.

When you don't need us

  • Your full suite already runs in under ten minutes. Run all of it on every pull request, add post-deploy smoke tests, and skip test selection entirely.
  • Your platform already does staged rollouts with automatic rollback on error rates. Configure it; don't rebuild it.
  • You need manual regression testers for each release. RAITHub builds automated suites and does not place testers in your team.
  • You need a certified vendor. RAITHub is not SOC 2 or ISO 27001 certified.

To get regression testing on every deploy without a slower pipeline, book the free 15-minute technical audit.

Documentation checked on 29 September 2026.

Frequently asked questions

Should I run the full regression suite on every deploy?

Run the full suite before every production deploy, usually on the main branch after merge. On pull requests, a selected subset plus the critical tests is enough, as long as the full run still gates the deploy.

What is the difference between regression tests and smoke tests?

Regression tests check that existing behaviour still works after a change, mostly in CI. Smoke tests are a short set run against the deployed environment to confirm it is healthy with its real configuration.

How do I run only the tests affected by a change?

Vitest's --changed option and Playwright's --only-changed option select tests by changed files against a Git ref. Always add a tagged critical set, and run everything when configuration or migrations change.

What is a canary release?

A partial, time-limited rollout where a small share of traffic gets the new version, and its error rate and latency are compared with the old version before the rollout continues or rolls back.

Do small products need canary releases?

Usually not. With low traffic, a small slice gives too little signal. Post-deploy smoke tests, a clear rollback rule and feature flags for risky changes give most of the benefit.

How long should a CI pipeline take?

As a working guide, under 10 minutes for pull request checks and under 20 for the full pre-deploy suite. Shard and move tests down the pyramid before you let it grow further.

regression testingpost-deploy smoke testscanary releasetest selectioncontinuous deploymentCI pipeline

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.