Back to BlogQuality & Testing

Flaky End-to-End Tests: How to Find, Fix and Quarantine Them

Rupak Amin

Founder & Lead Engineer, RAITHub

11 min read

A flaky test passes and fails on the same code. Prove it with repeated runs, then fix the cause, which is usually a fixed sleep instead of a real wait, state shared between tests, or data that depends on order or the clock. Quarantine anything you cannot fix today behind a tag, so it stops blocking merges but keeps running. Retries should report flakiness, never hide it.

The examples use Playwright, and every Playwright behaviour described links to its documentation, checked on 29 September 2026. The causes and fixes apply to any end-to-end (E2E) tool. If you are still choosing one, see Playwright vs Cypress in 2026. This is engineering guidance from how RAITHub runs its own suites, not a report of a particular client's flaky suite.

Why do flaky tests matter so much?

Because a suite people do not trust stops protecting anything. Once a red build is usually "just the flaky one", the team re-runs until it goes green and merges, and a real failure gets waved through the same way. A test that fails every time is useful; a test that fails sometimes teaches people to ignore failures.

Playwright's own reporting makes the distinction precise: a test is flaky when it failed on the first run but passed when retried (Playwright: retries). That label is the starting point for everything below.

What makes an end-to-end test flaky?

Almost always one of six causes. Each has a recognisable symptom, which tells you where to look first.

CauseTypical symptomFix
Fixed sleeps or manual checks instead of real waitsFails more on slow CI machines than on laptopsWeb-first assertions and waiting for the specific response
State shared between testsPasses alone, fails in the full suite or in parallelA fresh browser context per test and per-test data
Order-dependent or leftover dataFails on the second run, or when test order changesSeed unique data per test through an API; clean up after
Time and datesFails near midnight, month end or on a different time zoneControl the clock in the test; set the time zone in config
Third-party servicesFails when a payment sandbox or email provider is slowMock the external call in E2E; test the integration separately
A real race condition in the appFails rarely, and users occasionally report the same thingFix the application, not the test

The last row is the one to take seriously. Some flaky tests are accurate: the app really does double-submit a form on a fast double click, or render before its data arrives. Before you change a test, ask whether a user could hit the same failure.

How do you find which tests are flaky?

Run the suspect test many times in parallel, and make CI fail when anything is flagged as flaky. Both are built into the Playwright CLI (Playwright command line):

# Run one file 30 times across 4 workers; any failure is a signal
npx playwright test tests/checkout.spec.ts --repeat-each=30 --workers=4

# In CI: retry failures, but fail the run if anything needed a retry
npx playwright test --retries=2 --fail-on-flaky-tests

# Re-run only what failed last time, while you investigate
npx playwright test --last-failed

--repeat-each runs each test N times, and --fail-on-flaky-tests fails the run if any test is flagged as flaky. Running with several workers matters: tests that share state often pass in a single worker and only fail when they run side by side.

Keep a record. The HTML report marks flaky tests on every run; a JSON reporter lets you count, per test, how often it needed a retry over the last few weeks. The tests at the top of that list are where to spend time first.

How do you fix waiting and timing problems?

Replace every fixed sleep and every one-off check with a web-first assertion, and wait for the specific network response when a step depends on it. Playwright auto-waits before actions: it checks that an element is visible, stable, able to receive events and enabled before it clicks (Playwright: auto-waiting). Its web-first assertions retry until the condition is met or the timeout expires. A fixed sleep does neither.

// Flaky: a guess at how long saving takes, then a one-off check
await page.getByRole('button', { name: 'Save' }).click()
await page.waitForTimeout(2000)
expect(await page.getByText('Saved').isVisible()).toBe(true)

// Stable: the assertion retries until the text appears or it times out
await page.getByRole('button', { name: 'Save' }).click()
await expect(page.getByText('Saved')).toBeVisible()

// Stable, when the next step depends on the server having finished
const saved = page.waitForResponse(
  (res) => res.url().endsWith('/api/projects') && res.request().method() === 'POST',
)
await page.getByRole('button', { name: 'Save' }).click()
expect((await saved).ok()).toBe(true)

Playwright's best practices warn against the first pattern directly: a manual assertion such as expect(await locator.isVisible()) does not wait, so the test checks once and moves on (Playwright best practices). Search your suite for waitForTimeout and isVisible() inside expect; each hit is a candidate.

For dates, control the clock rather than hoping the test does not run at midnight. Playwright can fix or advance time inside the page (Playwright: clock), and setting timezoneId in the config makes CI and laptops agree on what "today" means.

How do you stop tests interfering with each other?

Give each test its own browser state and its own data. Playwright's guidance is that each test should run independently, with its own local storage, session storage, data and cookies (Playwright best practices). The browser part is free: Playwright creates a fresh browser context for every test. The data part is your job.

Login is the usual shared state. Sign in once in a setup project and reuse the saved storageState, rather than logging in through the UI in every test. When tests change server-side state for the signed-in user, such as settings or a cart, use one account per parallel worker, keyed on testInfo.parallelIndex, so two workers never edit the same account (Playwright: authentication). Keep the saved state files out of Git; they contain session cookies.

How should you seed test data?

Create what each test needs through an API, with unique names, in a fixture that also cleans up. Seeding through the UI is slow and adds its own flakiness; relying on rows left by an earlier test is the order dependency from the table above.

// tests/fixtures.ts
import { test as base, expect } from '@playwright/test'
import { randomUUID } from 'node:crypto'

type Project = { id: string; name: string }

export const test = base.extend<{ project: Project }>({
  project: async ({ request }, use) => {
    // A test-only endpoint, disabled in production.
    const res = await request.post('/api/test/projects', {
      data: { name: 'e2e-' + randomUUID() },
    })
    expect(res.ok()).toBe(true)
    const project: Project = await res.json()

    await use(project)

    await request.delete('/api/test/projects/' + project.id)
  },
})

export { expect }

// tests/rename-project.spec.ts
import { test, expect } from './fixtures'

test('renames a project', async ({ page, project }) => {
  await page.goto('/projects/' + project.id)
  await page.getByLabel('Project name').fill('Renamed')
  await page.getByRole('button', { name: 'Save' }).click()
  await expect(page.getByRole('heading', { name: 'Renamed' })).toBeVisible()
})

A random suffix means parallel workers and repeated runs never collide, and the fixture's teardown keeps the database from filling with leftovers. If a test-only endpoint is not an option, seed directly into a dedicated test database before the run, and never point E2E tests at production data.

For third-party services, route the call in E2E and return a fixed response with page.route() and route.fulfill() (Playwright: mock APIs). Playwright's advice is to test only what you control. Prove the real integration in a separate, smaller set of tests against the provider's sandbox.

What retry policy should you use?

Retry in CI, never locally, and treat every retry as a defect to fix. Retries exist to keep one bad run from blocking the team, not to make a flaky test look healthy.

// playwright.config.ts
import { defineConfig } from '@playwright/test'

export default defineConfig({
  retries: process.env.CI ? 2 : 0,
  use: {
    trace: 'on-first-retry',   // record a trace only when a retry happens
    timezoneId: 'UTC',
  },
  reporter: [['html'], ['json', { outputFile: 'test-results/results.json' }]],
})
  • Zero retries locally, so developers see flakiness while they are writing the test.
  • Two retries in CI at most. More than that hides real problems and makes slow runs slower.
  • A trace on the first retry, so every flaky run leaves the evidence needed to fix it.
  • A weekly look at the flaky list, with the top offender assigned to someone. Add --fail-on-flaky-tests once the list is short enough that failing the build is fair.

How do you quarantine a flaky test without losing it?

Tag it, link it to an issue, exclude the tag from the blocking CI job, and run it in a separate job that reports but does not block. Playwright supports tags and annotations on any test, and --grep and --grep-invert select by tag (Playwright: annotations and tags).

test('exports orders as CSV', {
  tag: '@quarantine',
  annotation: { type: 'issue', description: 'https://github.com/your-org/your-app/issues/123' },
}, async ({ page }) => {
  // ...
})

# Blocking job: everything except quarantined tests
npx playwright test --grep-invert @quarantine

# Non-blocking job: quarantined tests only, so you can see when they stabilise
npx playwright test --grep @quarantine

Set two rules and keep them. A quarantined test has an owner and an issue. And quarantine has a time limit, for example two weeks, after which the test is fixed, rewritten at a lower level, or deleted on purpose. Quarantine without a limit is just a slower way of deleting tests.

Some flaky E2E tests are testing the wrong thing at the wrong level. If a test is really checking a calculation or a validation rule, move it to a unit or integration test, where it is faster and cannot be flaky for browser reasons. That is the test pyramid doing its job.

Why RAITHub for this, and when to use a tool instead

RAITHub's QA-first method puts E2E tests at the top of a gated pyramid: type-checks, unit, integration and contract tests run first, and a Playwright journey only proves what the lower layers cannot. Every layer blocks the release, which only works if the tests are trusted, so flaky tests are treated as defects. The method, and the test counts behind the platforms RAITHub built, including 400+ tests on this website, are in how RAITHub tests software. The QA and test automation service applies the same approach to an existing product: stabilise the suite, move tests to the right level and gate the merge.

When a tool or your own team is enough:

  • A handful of flaky tests. The repeat-run command and the fixes above are a day's work for a developer on your team.
  • You need visibility, not fixes. A hosted test dashboard that tracks flaky rates over time may be all you need.
  • You want manual exploratory testing. RAITHub builds automated suites gated in CI; a manual-testing service is a better fit for that.

If your team has stopped trusting its E2E suite, book the free 15-minute technical audit and bring your flaky list.

Playwright documentation checked on 29 September 2026.

Frequently asked questions

What is a flaky test?

A test that passes and fails on the same code without any change. Playwright labels a test flaky when it failed on the first attempt and passed on a retry.

How do I find flaky tests in Playwright?

Run the suspect test many times with --repeat-each and several workers, and use --fail-on-flaky-tests in CI so any test that needed a retry fails the run. The HTML report marks flaky tests on each run.

Should I use retries for flaky tests?

In CI, up to two retries, with a trace recorded on the first retry. Locally, none. Treat every retried test as a defect to fix rather than a pass.

Why do my tests pass locally but fail in CI?

Usually because CI machines are slower and run tests in parallel. Fixed sleeps become too short, and tests that share accounts or data start to collide. Replace sleeps with web-first assertions and give each test its own data.

Is waitForTimeout bad in Playwright?

For making tests wait, yes. It guesses a duration instead of waiting for a condition. Use web-first assertions or wait for the specific network response instead.

How long should a test stay in quarantine?

Set a limit, such as two weeks, and give each quarantined test an owner and an issue. After the limit, fix it, rewrite it at a lower level, or delete it deliberately.

flaky testsflaky e2e testsPlaywrighttest isolationtest data seedingtest retriestest quarantine

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.