Back to BlogQuality & Testing

Every Release Breaks Something: Add Regression Gates in a Week

Rupak Amin

Founder & Lead Engineer, RAITHub

11 min read

If every release breaks something, you can have working regression gates in five working days. List the last ten things that broke and write a failing test for each. Add smoke tests for the three to five journeys that make money. Then make those checks required before any merge. Freeze new features for that week. It will not catch everything, but it stops the same bugs coming back.

This is the emergency plan, for a team that is shipping and hurting now. It is not a full QA strategy. For building testing habits from nothing over a month, read how to set up QA for an early-stage SaaS. This post is the week you spend stopping the bleeding first.

Why does every release break something?

Usually because nothing checks the parts that break, or the checks exist but nothing forces anyone to run them. Five causes account for most cases, and each one has a gate that stops it.

What you noticeLikely causeGate that stops it
The same bug returns a few releases laterFixes ship without a test that fails on the old codeOne regression test per fixed bug, required in CI
A change in one area breaks an unrelated screenShared code with no tests around its callersSmoke tests on the main journeys, run on every pull request
Tests pass locally, production breaksDifferent install, build or environment settingsCI builds with the same commands as the deploy target
Tests exist but bugs still shipTests are optional, slow or flaky, so people merge anywayRequired status checks that block the merge
Each release is huge and hard to reviewWeeks of changes shipped at onceSmaller releases, each one gated

None of this needs a QA department. It needs a short list of the right tests and a rule that the pipeline, not a person's memory, decides whether code can merge.

What does a one-week regression plan look like?

One deliverable a day, each one usable on its own. If the week is cut short, what you have already built still protects you.

DayDeliverableDone when
1A breakage list: the last 10 to 20 things that broke, grouped by areaEveryone agrees on the top five areas
2Smoke tests for the three to five journeys that make moneyThey run in CI in under five minutes
3A regression test for each recent bug on the listEach test fails on the code before its fix
4Required checks: nothing merges while a test failsA deliberately broken pull request cannot be merged
5A release routine: smoke tests after deploy and a rollback ruleThe team has done one gated release end to end

Freeze new features for the week. The developers who would build them write the tests instead, because they know where the code is fragile. A week without features is cheaper than another month of hotfixes.

Day 1: how do you find what actually keeps breaking?

From evidence, not memory. Pull the last two or three months of bug tickets, support complaints, hotfix commits and incident notes into one list. For each item, write the area, what the user saw, and whether it has happened before. Then sort by how often an area appears and how much it costs when it breaks. Payments, sign-in and anything that writes customer data usually sit at the top.

If something is broken right now and nobody knows which change caused it, git bisect finds the commit for you. It runs a binary search between a known-good and a known-bad commit. With git bisect run, a script decides each step: exit code 0 means good, and 1 to 127 (except 125) means bad (git-bisect documentation).

git bisect start
git bisect bad HEAD
git bisect good v1.42.0
git bisect run npx vitest run tests/checkout-total.test.ts
git bisect reset

That test file becomes the first regression test on your list. Finding the commit also tells you which kind of change keeps causing trouble, which is often more useful than the fix itself.

Day 2: which smoke tests should you write first?

A smoke test checks that a journey still completes. It does not check every detail. Pick the three to five journeys that would cost you money or trust within an hour of breaking: sign up, log in, the core action customers pay for, checkout or billing, and one admin action if your staff depend on it.

Tag them so they can run on their own. In Playwright, a test gets a tag through a details object or an @ token in its title, and --grep @smoke runs only the tagged tests (Playwright: annotations and tags).

// tests/smoke/checkout.spec.ts
import { test, expect } from '@playwright/test'

test('customer can buy a product', { tag: '@smoke' }, async ({ page }) => {
  await page.goto('/products/sample-item')
  await page.getByRole('button', { name: 'Add to cart' }).click()
  await page.getByRole('link', { name: 'Checkout' }).click()
  await page.getByLabel('Email').fill('smoke-test@example.com')
  await page.getByRole('button', { name: 'Place order' }).click()
  await expect(page.getByRole('heading', { name: 'Order confirmed' })).toBeVisible()
})

Keep smoke tests few and stable. Five reliable tests that everyone trusts are worth more than fifty that fail at random, because a flaky gate teaches people to ignore red builds. If tests start failing without a code change, the guide to flaky end-to-end tests shows how to fix them rather than re-run them. If you have not chosen a browser test tool yet, see Playwright vs Cypress in 2026.

Day 3: how do you write regression tests for bugs you already fixed?

Write each one at the lowest level that reproduces the bug, and prove it by running it against the code from before the fix. A regression test that passes on the broken code protects nothing.

  1. Reproduce the original bug from the ticket: the input, the account type, the data.
  2. Choose the level. A wrong total is a unit test on the pricing function. A permission leak is an integration test against the API. A broken form is an end-to-end test.
  3. Check out the commit before the fix and confirm the test fails. Then confirm it passes on the current code.
  4. Name it after the behaviour, not the ticket number, so a failure explains itself: "a partial update changes only the fields sent".

That last example is real. On this website in September 2026, zod 4's .partial() re-applied default values to partial updates, which would have wiped fields on a simple save. The fix added a regression test for every update schema, 19 tests that took the suite from 381 to 400 (the full write-up). One bug, fixed once, guarded for good.

Day 4: how do you make the tests actually block a bad release?

Make them required. In GitHub, a branch protection rule with required status checks means a pull request cannot be merged until those checks succeed. By default, though, branch protection rules don't apply to people with admin permissions. Turn on "Do not allow bypassing the above settings" if the founder or lead is the one who tends to merge in a hurry (GitHub: about protected branches).

A minimal workflow that runs the unit, regression and smoke tests on every pull request:

# .github/workflows/regression.yml
name: regression
on: [pull_request]
jobs:
  gate:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 22
          cache: npm
      - run: npm ci
      - run: npm run build
      - run: npx vitest run
      - run: npx playwright install --with-deps chromium
      - run: npx playwright test --grep @smoke

Test the gate itself. Open a pull request that breaks one of the smoke tests on purpose and confirm the merge button is blocked. A gate that has never stopped anything might not be connected.

Day 5: what should happen after each deploy?

Run the same smoke tests against production, and agree in advance what triggers a rollback. Point the Playwright baseURL at the live site, run --grep @smoke with a dedicated test account, and alert the team if it fails. Write the rollback rule down: if a smoke test fails after deploy, roll back first and investigate second. Deciding that before an incident saves the argument during one.

Then do one full gated release as a team: pull request, green checks, deploy, post-deploy smoke run. The first one will be slower than usual. That is the cost of finding out now, rather than from a customer.

How do you know the week worked?

Measure breakage per release, not test count. DORA defines change failure rate as "the ratio of deployments that require immediate intervention following a deployment" (DORA metrics guide). Count it by hand if you have to: releases in a week, and how many needed a hotfix or a rollback. Track two more numbers alongside it:

  • Repeat bugs: how many new bug reports match something already on your Day 1 list. This should drop to near zero first.
  • Escaped bugs by area: where the remaining breakages happen. Each one tells you where the next test goes.

Expect the first gains within two to three releases, mostly from repeats disappearing. Bugs in new code will still get through, until testing new code becomes part of every change.

What should you do in week two?

Turn the emergency into a habit. Every bug fix from now on ships with its regression test in the same pull request. Every new feature ships with tests for its rules. Grow the gates in the order laid out in the early-stage QA setup guide: integration tests on a real database, a clean migration replay, then more journeys. For how a full pipeline looks once it matures, see how RAITHub tests software. PropDesk, the property management platform RAITHub built, has 1,024 automated tests, 932 unit and 92 end to end, all run in CI. That number is where steady habits lead, not a week-one target.

Why RAITHub for this

Because regression gates are how RAITHub builds by default, not an add-on. Every engagement includes a test suite with CI gates, and a failing layer blocks the release. For a team in the situation above, the QA and test automation service runs this plan with you. It ranks your breakage list, writes the first smoke and regression tests, wires up required checks, and hands it over so your developers keep it going. Scope and price are fixed and written after a free 15-minute technical audit, and the tests belong to you.

When you don't need us

  • Your team can take the week. This plan is written so a small team can run it without outside help.
  • You want manual testers clicking through each release. RAITHub builds automated suites and does not place testers in your team.
  • The product is being replaced soon. Gate only the journeys that must survive until the switch, and spend the rest on the new system.
  • You need a certified vendor. RAITHub is not SOC 2 or ISO 27001 certified.

To put regression gates on a product that keeps breaking, book the free 15-minute technical audit.

Documentation checked on 29 September 2026.

Frequently asked questions

Why does every release break something?

Usually because fixes ship without regression tests, the main journeys have no smoke tests, or tests exist but are not required to pass before a merge. Large, infrequent releases make it worse.

Can you really add regression testing in one week?

You can add the gates that stop repeats: smoke tests for three to five key journeys, a regression test for each recent bug, and required CI checks. Full coverage takes longer and grows with each change after that.

What is the difference between a smoke test and a regression test?

A smoke test checks that a key journey still completes, such as checkout. A regression test checks that one specific bug that was fixed has not come back. You need both.

Should we stop shipping features while we add tests?

For one week, yes. The developers who know the fragile code are the right people to write the first tests, and a short freeze costs less than continued hotfixes.

How do we stop people merging when tests fail?

Use branch protection with required status checks, and turn on the setting that stops admins bypassing it. Then open a deliberately failing pull request to prove the gate blocks it.

How do we measure whether releases are getting better?

Track change failure rate: the share of deployments that need a hotfix or rollback. Also count repeat bugs, which should fall first, and note where the remaining breakages happen.

every release breaks somethingregression testingregression gatessmoke testsCI gatesrelease quality

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.