Back to BlogQuality & Testing

QA Metrics That Actually Matter (and Vanity Ones to Drop)

Rupak Amin

Founder & Lead Engineer, RAITHub

9 min read

The QA metrics that matter measure what reaches users and how fast you recover, not how much testing you did. Track escaped defects (bugs found in production), change failure rate, failed-deployment recovery time, and pipeline time. Drop raw test count, overall coverage percentage and pass rate as goals: they rise without quality improving, so the team optimises the number instead of the product.

If you would rather have QA run and reported for you, see how RAITHub would test this near the end. The rest is how to pick and read the metrics yourself.

Why do most QA dashboards measure the wrong thing?

Because the easy numbers to collect are the least useful. Test count, coverage percentage and pass rate are all outputs of the testing, not outcomes for users. They go up when you add weak tests, and they stay green while a real bug ships. The useful question is never "how many tests do we have"; it is "what reaches users, and how fast do we recover when something does". Once a number becomes a target, people optimise the number, so choose targets that only improve when the product improves.

Which QA metrics actually matter?

Four, plus one for teams that sell on latency. Each measures an outcome, not an activity.

MetricWhat it measuresWhy it matters
Escaped defectsBugs found in production that testing should have caught before releaseThe direct measure of whether QA is working. Trending up means coverage is missing where the product actually breaks
Change failure rateThe share of deployments that need an immediate fix or rollbackTells you how often shipping goes wrong, regardless of test count
Failed-deployment recovery timeHow long it takes to recover from a bad deployYou will ship some bad changes; this is how much they hurt
Pipeline timeHow long CI takes on a pull request and before deployA slow pipeline gets bypassed, which quietly turns the gates off
P95 latency on key routesThe 95th-percentile response time of revenue-earning endpointsA performance regression is a defect users feel; make it a tracked number

Three of these come straight from DORA's research. DORA defines change failure rate as "the ratio of deployments that require immediate intervention following a deployment", and failed deployment recovery time as "the time it takes to recover from a deployment that fails and requires immediate intervention" (DORA's four keys). DORA has since added a fifth key, deployment rework rate: "the ratio of deployments that are unplanned but happen as a result of an incident in production" — another honest way to see escaped defects.

What is an escaped defect, and how do you measure it?

An escaped defect is a bug that reached production when a test or review should have caught it first. Count them per release, and tag each one with the layer that should have caught it: a unit test, an API test, an end-to-end test, or manual exploratory testing. The tag is the useful part, because it tells you where to add coverage.

  • Count, don't just list. Escaped defects per release, trended over time, is the headline.
  • Classify by severity. One data-loss bug matters more than ten cosmetic ones; weight the count or report the two separately.
  • Write a test for every escape, at the lowest level that reproduces it. The suite then grows exactly where the product breaks. This is the habit behind the measured counts on platforms RAITHub built, and the method is in regression testing on every deploy.

Why are test count and coverage misleading?

Because both rise without quality rising. A thousand tests that never touch tenant isolation or billing look impressive and protect nothing. Coverage is worse as a target: it is easy to hit a high number with tests that execute code without checking its behaviour.

Martin Fowler is blunt about coverage as a management metric. It "helps you find which bits of your code aren't being tested", but is "of little value to management since you need a technical background to understand whether the tests are good or whether the uncovered code is a problem" (Martin Fowler, Test Coverage). Use coverage to find untested code, never as a KPI. The real measure of a suite, he argues, is rare bugs escaping to production and the confidence to change code without fear — which is exactly the escaped-defect metric above.

Vanity metricWhy it misleadsTrack instead
Total test countGoes up with weak tests; says nothing about what is coveredEscaped defects by layer
Overall coverage %Hit with tests that run code without asserting behaviourCoverage of billing, permissions and tenant-scoped endpoints specifically
Pass rate99% green is normal; the 1% and the flaky reruns hide the riskFlaky-test count and quarantine age
Bugs found by QARewards finding many small bugs late over preventing themEscaped defects and change failure rate

How should you read these metrics together?

No single number tells the story; the pairs do. Watch how two move together before you act.

  • Escaped defects down, pipeline time up. You are buying quality with speed. Fine for a while, but a pipeline that grows every month gets bypassed, so shard and trim before it crosses your budget.
  • Change failure rate down, recovery time up. You ship safely but recover slowly. Invest in post-deploy smoke tests and a rehearsed rollback, covered in regression testing on every deploy.
  • Pass rate high, escaped defects high. Your tests check the wrong things. Add tests at the layer where defects escape, usually the API layer in a SaaS; see the test pyramid for a SaaS.
  • Flaky-test count rising. Trust in the suite is draining. Quarantine with a deadline, not open-ended retries; the flaky end-to-end tests guide covers it.

How do you set targets without gaming them?

Set a direction, not a magic number. "Escaped defects trending down quarter on quarter" is a better target than "95% coverage", because it cannot be hit by writing weak tests. Where you must have a threshold, put it on the specific thing that matters: full coverage of billing, permissions and tenant-scoped endpoints, and a P95 latency ceiling on the checkout path. Review the metrics in a short monthly session and change one thing in response; a dashboard nobody acts on is its own vanity metric.

Buy, build or hire?

OptionWhat you getChoose this when
A reporting tool or platformDashboards for test runs, coverage and DORA metricsYou have someone to read and act on them; the tool only draws the charts
Freelancers or crowdtestingA burst of manual testing; defect counts for one engagementYou need extra eyes once, not an ongoing measure
An in-house QA hireSomeone who owns the metrics and the suiteYou ship often and can keep one person busy full time
A managed QA team (QAaaS)A team that runs the testing and reports escaped defects and release health weeklyYou want the measures acted on, not just displayed

How RAITHub would test this

RAITHub runs QA as a managed service and reports the outcome metrics, not the vanity ones: escaped defects by layer, change failure rate, recovery time, pipeline time and P95 latency on your key routes, in a weekly report.

  • Scope: baseline what is tested now and what escapes to production; add tests at the layers where defects actually get through; wire DORA-style metrics into your pipeline; review them monthly and act on one thing.
  • Ways to buy it: a monthly QA plan, a fixed-price one-off audit, or a dedicated QA team that RAITHub manages and bills monthly. Not staff augmentation.
  • Timeline: fixed in the written quote after a free audit.
  • What you receive: tests and CI in your repository, a weekly QA report, IP assigned to you and an NDA as standard.

Proof is the measured test counts on platforms RAITHub built: 1,024 tests on PropDesk, 750+ on TheSkinProof (the founder's own venture, not a client), 530+ on Sundor Skin, and 400+ on this site. These are counts, shown for discipline; there is no QA case study yet, and none is implied. See the QA as a Service hub and the manual and exploratory testing service, then book the free 15-minute technical audit.

Documentation checked on 9 October 2026.

Frequently asked questions

What are the most important QA metrics?

Escaped defects (bugs found in production), change failure rate, failed-deployment recovery time and pipeline time. For products that sell on speed, add the P95 latency of key routes. These measure outcomes for users, not how much testing was done.

Why is test coverage a bad metric to target?

Coverage rises with tests that execute code without checking its behaviour, so a high percentage can hide weak testing. Use coverage to find untested code, and target full coverage of billing, permissions and tenant-scoped endpoints specifically, not an overall number.

What is an escaped defect?

A bug that reached production when a test or review should have caught it first. Count them per release and tag each with the layer that should have caught it, so you know where to add tests.

What is a good change failure rate?

Lower is better, but the useful signal is the trend for your own team, not a single benchmark. DORA defines it as the share of deployments that need immediate intervention, such as a rollback or hotfix. Watch it fall as your gates improve.

Is test count a useful QA metric?

Not as a goal. A large suite that never tests tenant isolation or billing protects nothing. Count escaped defects by layer instead, and let the test count grow as a by-product of writing a test for every bug.

How often should we review QA metrics?

Monthly is enough for most teams. Keep the session short, read the metrics in pairs, and change one thing in response. A dashboard nobody acts on is itself a vanity metric.

qa metricsdora metricsescaped defectschange failure ratetest coveragequality measurement

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.