Back to BlogQuality & Testing

Continuous Testing in CI/CD: What to Automate and Gate

Rupak Amin

Founder & Lead Engineer, RAITHub

9 min read

Continuous testing means running automated tests at every stage of your pipeline, so a change is checked the moment it is proposed and again before users see it. In practice: fast checks on each pull request, the full suite before deploy, and smoke tests against production after. Each stage is a required gate, so a failing check blocks the merge or release automatically.

If you would rather have this built and run for you, see how RAITHub would test this near the end. The rest is the method, so your own team can set it up.

What is continuous testing, and how is it different from "running tests in CI"?

Running tests in CI means your suite executes on a server. Continuous testing means tests run at the right stage, as gates, and the result decides whether the change moves forward. The difference is whether a red result can be ignored. Google's SRE book makes the same point about releases: the continuous build test targets should be the same targets that gate the release, and a release is cut from the last build that passed all of them (Google SRE Book, Release Engineering).

The other half is that the pipeline is reproducible. The SRE book calls for hermetic builds: "If two people attempt to build the same product at the same revision number in the source code repository on different machines, we expect identical results." A test that passes on one machine and fails on another is not a gate anyone will trust.

What should run at each stage of the pipeline?

Match the stage to how often it runs and how long it may take. The earlier the stage, the faster it must be, because someone is waiting on it.

StageWhen it runsWhat runsTime budget (guide)On failure
Pull requestEvery push to a PRType-check, lint, unit tests, tests for changed files, the critical setUnder 10 minutesMerge blocked
Main branchAfter merge, before deployFull unit, API and end-to-end suite, in parallelUnder 20 minutesDeploy blocked
Post-deployRight after each deploySmoke tests against the live environmentUnder 5 minutesRoll back or stop the rollout
ScheduledNightlySlow suites: more browsers, load tests, long scenariosNo hard limitTicket for the next day

The budgets are RAITHub's working guide for small and mid-sized products, not an industry standard. The principle holds at any size. The tiered approach, including test selection and canary rollouts, is covered in full in regression testing on every deploy.

What should be a required gate, and what should only warn?

Gate the checks that mean the change is wrong, and only warn on checks that are advisory or noisy. A gate that fails for reasons unrelated to the change teaches the team to click "merge anyway", which quietly turns every gate off.

  • Gate: type checks, lint, unit tests, API tests (tenant isolation, roles, contracts), the critical end-to-end journeys, and a performance threshold on key endpoints.
  • Warn, don't gate: full-matrix cross-browser runs, visual-diff checks that need a human to confirm, and coverage percentage. Coverage is a diagnostic, not a target: Martin Fowler notes it "helps you find which bits of your code aren't being tested" but is "of little value to management" as a number (Martin Fowler, Test Coverage).

Mind one GitHub Actions trap: a skipped job reports success and "will not prevent a pull request from merging, even if it is a required check" (GitHub Actions: using conditions to control job execution). If a conditional job is a gate, make sure it runs on the paths that matter, or it waves changes through.

What does a continuous-testing pipeline look like in GitHub Actions?

Run the cheap layers first and stop early. A failing type check should not wait for a browser to boot. Playwright's own CI guidance generates a workflow that runs "on each push and pull request into the main/master branch" and installs browsers with npx playwright install --with-deps (Playwright: Continuous Integration).

# .github/workflows/test.yml
name: test
on: [pull_request]
jobs:
  static-and-unit:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with: { node-version: 22, cache: npm }
      - run: npm ci
      - run: npx tsc --noEmit
      - run: npx eslint .
      - run: npx vitest run
  api:
    needs: static-and-unit
    runs-on: ubuntu-latest
    services:
      postgres:
        image: postgres:17
        env: { POSTGRES_PASSWORD: test }
        ports: ['5432:5432']
    env:
      DATABASE_URL: postgresql://postgres:test@localhost:5432/postgres
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with: { node-version: 22, cache: npm }
      - run: npm ci
      - run: npx prisma migrate deploy
      - run: npx playwright test tests/api
  e2e:
    needs: api
    runs-on: ubuntu-latest
    strategy:
      matrix: { shard: [1, 2] }
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with: { node-version: 22, cache: npm }
      - run: npm ci
      - run: npx playwright install --with-deps chromium
      - run: npx playwright test tests/e2e --shard=${{ matrix.shard }}/2

Each job's needs means a failing earlier job stops the later ones, so the pipeline fails fast and cheaply. Make the jobs required checks on the branch so a red result actually blocks the merge. Replace the Prisma line with your own migration command.

Should performance be a continuous-testing gate?

Yes, for the endpoints behind your money-making journeys. A load test with codified thresholds turns performance into a pass/fail check like any other. k6's thresholds "are the pass/fail criteria that you define for your test metrics", and if one fails the test exits with a non-zero code, which fails the CI step (k6 thresholds).

// load/checkout.js
import http from 'k6/http'

export const options = {
  thresholds: {
    http_req_duration: ['p(95)<200'],
    http_req_failed: ['rate<0.01'],
  },
}

export default function () {
  http.get(`${__ENV.BASE_URL}/api/health`)
}

Run this against a staging environment, not production, and keep the load realistic. For page-level performance, set budgets on Core Web Vitals too: the "good" thresholds on web.dev are an LCP of 2.5 seconds or less, an INP of 200 milliseconds or less, and a CLS of 0.1 or less at the 75th percentile.

How do you keep continuous testing fast as the suite grows?

Treat pipeline time as a metric with an owner. Suites slow one test at a time until people start skipping them.

  • Select tests on a pull request. Vitest's --changed and Playwright's --only-changed run tests affected by changed files; always add a tagged critical set, and run everything when config, migrations or lockfiles change.
  • Shard before you delete. Both runners support --shard=1/4. Four five-minute shards beat one twenty-minute job.
  • Push tests down the pyramid. An end-to-end test that really checks a pricing rule belongs in a unit test that runs in milliseconds. See the test pyramid for a SaaS.
  • Quarantine flaky tests with a deadline, never open-ended retries. A random failure teaches the team to ignore red builds; the flaky end-to-end tests guide covers the method.

Buy, build or hire?

OptionWhat you getChoose this when
A CI or testing platformHosted runners, test dashboards, some low-code testsYou have engineers to write and own the gates; the platform runs them
Freelancers or crowdtestingManual passes before a releaseYou need extra eyes once; it does not build a pipeline
An in-house QA or platform hireSomeone who owns the pipeline and the suiteYou ship often and can keep one person busy full time
A managed QA team (QAaaS)A team that builds the gates, writes the tests and reports weeklyYou want continuous testing running without hiring or managing testers

How RAITHub would test this

RAITHub sets up continuous testing on an existing product: it adds tags and test selection for pull requests, shards the full suite, writes the post-deploy smoke set and a rollback rule, adds a k6 performance threshold on the key endpoints, and makes the right jobs required checks.

  • Scope: audit how you ship today; gate the critical paths first; widen coverage and add performance thresholds; hand over the pipeline in your repository.
  • Ways to buy it: a monthly QA plan, a fixed-price one-off audit, or a dedicated QA team that RAITHub manages and bills monthly. Not staff augmentation; testers stay managed by RAITHub.
  • Timeline: fixed in the written quote after a free audit, based on your suite size and how often you deploy.
  • What you receive: tests, CI gates and playbooks in your repository, a weekly reliability report, IP assigned to you and an NDA as standard.

Proof is the measured test counts on platforms RAITHub built: 1,024 tests on PropDesk, 750+ on TheSkinProof (the founder's own venture, not a client), 530+ on Sundor Skin (CI replays all 76 migrations), and 400+ on this site. See the QA and test automation service and the QA as a Service hub, then book the free 15-minute technical audit.

Documentation checked on 9 October 2026.

Frequently asked questions

What is continuous testing in CI/CD?

Running automated tests at every stage of the pipeline as required gates, so a failing change is blocked before it merges or deploys. Fast checks run on each pull request, the full suite before deploy, and smoke tests against production after.

Which tests should be a required gate?

Type checks, lint, unit tests, API tests for tenant isolation and roles, the critical end-to-end journeys, and a performance threshold on key endpoints. Keep noisy or advisory checks, such as full cross-browser matrices and coverage percentage, as warnings.

How long should a CI pipeline take?

As a working guide, under 10 minutes for pull-request checks and under 20 for the full pre-deploy suite. Select tests on pull requests, shard the full suite, and push tests down the pyramid before you let it grow further.

Should performance tests gate a release?

Yes, for the endpoints behind revenue-earning journeys. Codify thresholds, for example 95% of requests under 200 milliseconds and under 1% failing, and let the test exit non-zero to fail the pipeline step.

Why do my required checks pass when a job was skipped?

In GitHub Actions a skipped job reports success and will not block a merge, even as a required check. If a conditional job is a gate, make sure it runs on the paths that matter, or changes slip through.

Is coverage percentage a good gate?

No. Coverage shows which code is untested but is a weak target; teams optimise for the number instead of good tests. Gate full coverage of billing, permissions and tenant-scoped endpoints, and treat the overall percentage as a diagnostic.

continuous testingci cdgithub actionsrelease gatestest automationpipeline

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.