Back to BlogQuality & Testing

Every AI Prompt Breaks Something Else: Regression Testing AI-Edited Code

Rupak Amin

Founder & Lead Engineer, RAITHub

11 min read

AI edits break working code because the tool changes what you asked about without knowing everything that depends on it. The fix is a safety net that runs after every prompt: one change per prompt, a save point before each, a small set of golden-path tests for the journeys that must never break, visual checks on shared screens, protected test files and a CI gate before anything merges.

If you would rather have that safety net built and run for you, see how RAITHub would test this below.

Everyone can build with AI now. Almost nobody building that way has a tester, so the person who prompts the change is also the only one checking it, usually by clicking the screen they just changed. Regressions live on the screens they did not click. This post is about that everyday loop. If you already have a test suite and want to run it on every deploy without slowing releases, the general process is in regression testing after every deploy.

Why does fixing one thing with AI break something else?

Because the model sees part of the codebase and optimises for the request in front of it. Researchers who build tools for coding agents put it plainly: AI coding agents "frequently introduce regressions -- breaking tests that previously passed" (Alonso, Yovine and Braberman, TDAD, arXiv 2603.17973). In their experiments on SWE-bench Verified, giving the agent a map of which tests relate to the code it was changing cut test-level regressions by 70%, from 6.08% to 1.82%. Telling it to "follow test-driven development" without that context made things worse, raising regressions to 9.94%.

Two lessons for anyone prompting changes into a real app. First, regressions are normal behaviour for these tools, not bad luck. Second, generic instructions do not prevent them; specific tests that run do.

Longer-running work makes it harder. SWE-CI, a 2026 benchmark, evaluates agents on maintaining real repositories through continuous integration, with tasks drawn from histories averaging 233 days and 71 consecutive commits, because one-shot fixes do not show how code quality holds up over many changes (SWE-CI, arXiv 2603.03823). A product built by prompting for three months is exactly that situation.

What kinds of regressions do AI edits cause?

RegressionHow it happensWhat catches it
Shared component changedYou ask for a new button style on one page; the tool edits the shared component used on twentyVisual snapshots of key screens
"While I was here" refactorThe tool renames a function, moves a file or "cleans up" code you did not mentionTypecheck, lint and the golden-path suite
Permission check lostA rewrite of a data query drops the owner filter or a role checkAccess-control tests with two users
Schema driftA prompt adds or renames a column; other screens still use the old nameIntegration tests against a real database; migration review
Test weakenedAsked to fix a failing test, the tool edits the assertion instead of the codeReview of test diffs; protected test files
Dependency changedA package is added or upgraded to make an error go awayLockfile review and a dependency audit in CI

The "test weakened" row is measured too: ImpossibleBench found that frontier models, given tests that contradict the task, often pass by modifying the tests, and that Claude models and Qwen3-Coder did so mainly through test edits (ImpossibleBench, arXiv 2510.20270). The companion post on why AI-written tests pass and still miss bugs covers how to spot it.

What is the minimum regression safety net for AI-edited code?

Six habits, in order of cost. The first three take an afternoon to set up.

1. One change per prompt, with a save point before it

Commit before every prompt in Cursor or Claude Code, so any change can be undone in one step and reviewed as one diff. In Lovable, Bolt or Replit, use the version history and bookmark versions that work. Lovable's docs note that restoring a version restores code only and does not roll back database data (Lovable version history docs), so prompts that change the database need extra care; see Lovable and Supabase migrations in production.

2. A golden-path suite of five to ten tests

List the journeys that must never break: sign up, log in, the core action, pay, and "user A cannot see user B's data". Automate each as one end-to-end test. Five to ten tests that run in a couple of minutes are worth more after each prompt than three hundred unit tests nobody runs. A minimal Playwright example:

// e2e/golden-path.spec.ts
import { test, expect } from '@playwright/test'

test('a customer signs in and sees only their own projects', async ({ page }) => {
  await page.goto('/login')
  await page.getByLabel('Email').fill(process.env.E2E_USER_EMAIL ?? '')
  await page.getByLabel('Password').fill(process.env.E2E_USER_PASSWORD ?? '')
  await page.getByRole('button', { name: 'Sign in' }).click()

  await expect(page.getByRole('heading', { name: 'Projects' })).toBeVisible()
  // Seeded data: this project belongs to a different customer
  await expect(page.getByText('Other customer project')).toHaveCount(0)

  // Catch accidental layout changes to a shared screen
  await expect(page).toHaveScreenshot('projects.png', {
    mask: [page.getByTestId('last-updated')],
  })
})

3. Visual checks on the screens that share components

Playwright's toHaveScreenshot() saves a reference image on the first run and compares against it afterwards; npx playwright test --update-snapshots accepts intended changes. The docs warn that rendering varies with operating system, browser version and hardware, so generate and compare baselines in the same environment, usually CI (Playwright visual comparisons docs). Mask anything that changes on its own, such as dates.

4. Run the checks automatically, not when the agent decides to

Asking an agent to "run the tests" works until it does not. Claude Code hooks run shell commands at fixed points, such as after a file edit or when Claude finishes responding, which gives "deterministic control: certain actions always happen rather than relying on the LLM to choose to run them" (Claude Code hooks guide). The same docs show a hook that blocks edits to protected files, which is one way to keep the agent away from your core tests. Cursor's project rules in .cursor/rules can tell the model which tests to run, but they are instructions, not enforcement (Cursor rules docs).

5. Tell the agent which tests cover the code it is changing

This is the practical version of the TDAD result. Instead of "be careful not to break anything", say "this change touches billing; run e2e/golden-path.spec.ts and tests/billing afterwards and fix the code, not the tests, if they fail". Specific context helped the agents in the study; generic process instructions did not.

6. A CI gate before anything merges

Run typecheck, lint, unit tests and the golden-path suite on every pull request, and make them required. With branch protection, required status checks must pass before a change can merge (GitHub protected branches). This is the one check that does not depend on anyone remembering. If the end-to-end tests become unreliable, fix them quickly; fixing flaky end-to-end tests covers how, because a suite people learn to ignore protects nothing.

How much testing is enough when you prompt changes every day?

Your situationEnough for now
Prototype, no real usersSave points and a manual click through the golden paths after each session
First paying usersGolden-path suite and access-control tests in CI, run on every change
Weekly releases, several people promptingAdd visual checks, protected tests, mutation testing on billing and permissions, and a human exploratory pass per release
Data or money at real scaleAll of the above, plus the tiered process in regression testing after every deploy and a balanced test pyramid

Automation catches what you thought to write down. An exploratory session by a person who did not build the feature catches the rest, which is why the table adds one as releases speed up.

How long does it take to set this up yourself?

For a small app, budget 2–4 days if you know Git and a test runner: half a day for save-point habits and CI, one to two days to write and stabilise five to ten golden-path tests with seeded test users, and half a day for visual baselines and protecting test files. Add time if the app has no separate test environment or test data yet. The main risk is that the tests themselves are written by the same AI that writes the features, from the same code, so they confirm today's behaviour rather than the behaviour you want. Write the golden paths from your product rules, and have someone other than the author review them.

Buy, build or hire?

OptionChoose this whenWatch out for
A tool or SaaS testing platform (AI test generators, visual testing services, record-and-replay tools)You want tests quickly and can maintain themRecorded tests break with every UI change an AI makes; someone still decides what "correct" means
Freelancers or crowdtestingYou want a manual regression pass before a releaseA pass is a snapshot; it does not run after the next prompt
An in-house QA hireSeveral people prompt changes daily and you can keep a tester busySalary and management time; see in-house vs outsourced QA
A managed QAaaS teamYou want the safety net built, then a human regression pass on each releaseAgree the golden paths together, so automation covers what matters to your business

Why RAITHub for this

  • Tests that survive constant change. RAITHub's own platforms are maintained behind large suites: 1,024 tests on PropDesk, 530+ on Sundor Skin and 400+ on this website, all gating changes in CI.
  • Expected results from your rules, not your code, so the tests catch the regression instead of encoding it.
  • Honest proof. There is no published case study of an AI-built-app engagement yet; the test counts above are the evidence on offer.

When you don't need us

  • You have a developer with a few days to set up the six habits above. Do that.
  • The app is a throwaway prototype. Save points and a manual click-through are enough.
  • You want testers placed under your own management. RAITHub does not offer staff augmentation.

How RAITHub would test this

Through AI app testing, part of QA as a service, RAITHub would:

  • map your golden paths and access rules with you, and seed test users and data for them
  • run a full regression and exploratory pass on the current app, and report what earlier prompts already broke
  • write the golden-path, access-control and visual tests into your repository, with a CI gate and protected test files; see QA test automation
  • on a monthly plan, re-test each batch of AI-generated changes by hand and keep the suite current as the app changes

Start with a fixed-price launch audit for your AI-built app: a written report of the bugs and regressions found, ranked by impact, each with a fix. Optional monthly QA keeps the safety net working as you keep prompting. You own the tests and all IP; an NDA is standard.

Next step: book a free 15-minute call about a launch audit, then get a written fixed quote.

Frequently asked questions

Why does Cursor or Claude Code change files I did not ask about?

To make the requested change work, the tool may edit shared code, rename things or tidy nearby code. Ask for one change per prompt, say which files are in scope, and review the diff before accepting it.

How do I undo an AI change that broke my app?

Revert to the save point: a Git commit for code tools, or a bookmarked version in Lovable, Bolt or Replit. Remember that reverting code may not undo database changes, so check the data too.

Can the AI write the regression tests?

It can write the code of the tests. You, or a tester, should decide what each test expects, from your product rules. Otherwise the tests tend to confirm whatever the app does today, bugs included.

How many end-to-end tests do I need?

Start with five to ten for the journeys that must never break, and keep them fast and reliable. Add more where bugs actually appear, not everywhere at once.

Do visual regression tests work with AI-generated UI?

Yes, and they suit it, because AI edits often change shared components. Generate baselines in CI, mask changing content such as dates, and update baselines only for changes you meant to make.

Is a CI gate overkill for a small app?

No. It is a short configuration file and it is the only check that runs whether or not anyone remembers. For an app with paying users it is low-cost protection against regressions.

regression testingAI breaks existing codeClaude CodeCursorPlaywrightAI app testing

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.