Founder & Lead Engineer, RAITHub
The QA metrics that matter measure what reaches users and how fast you recover, not how much testing you did. Track escaped defects (bugs found in production), change failure rate, failed-deployment recovery time, and pipeline time. Drop raw test count, overall coverage percentage and pass rate as goals: they rise without quality improving, so the team optimises the number instead of the product.
If you would rather have QA run and reported for you, see how RAITHub would test this near the end. The rest is how to pick and read the metrics yourself.
Why do most QA dashboards measure the wrong thing?
Because the easy numbers to collect are the least useful. Test count, coverage percentage and pass rate are all outputs of the testing, not outcomes for users. They go up when you add weak tests, and they stay green while a real bug ships. The useful question is never "how many tests do we have"; it is "what reaches users, and how fast do we recover when something does". Once a number becomes a target, people optimise the number, so choose targets that only improve when the product improves.
Which QA metrics actually matter?
Four, plus one for teams that sell on latency. Each measures an outcome, not an activity.
| Metric | What it measures | Why it matters |
|---|---|---|
| Escaped defects | Bugs found in production that testing should have caught before release | The direct measure of whether QA is working. Trending up means coverage is missing where the product actually breaks |
| Change failure rate | The share of deployments that need an immediate fix or rollback | Tells you how often shipping goes wrong, regardless of test count |
| Failed-deployment recovery time | How long it takes to recover from a bad deploy | You will ship some bad changes; this is how much they hurt |
| Pipeline time | How long CI takes on a pull request and before deploy | A slow pipeline gets bypassed, which quietly turns the gates off |
| P95 latency on key routes | The 95th-percentile response time of revenue-earning endpoints | A performance regression is a defect users feel; make it a tracked number |
Three of these come straight from DORA's research. DORA defines change failure rate as "the ratio of deployments that require immediate intervention following a deployment", and failed deployment recovery time as "the time it takes to recover from a deployment that fails and requires immediate intervention" (DORA's four keys). DORA has since added a fifth key, deployment rework rate: "the ratio of deployments that are unplanned but happen as a result of an incident in production" — another honest way to see escaped defects.
What is an escaped defect, and how do you measure it?
An escaped defect is a bug that reached production when a test or review should have caught it first. Count them per release, and tag each one with the layer that should have caught it: a unit test, an API test, an end-to-end test, or manual exploratory testing. The tag is the useful part, because it tells you where to add coverage.
- Count, don't just list. Escaped defects per release, trended over time, is the headline.
- Classify by severity. One data-loss bug matters more than ten cosmetic ones; weight the count or report the two separately.
- Write a test for every escape, at the lowest level that reproduces it. The suite then grows exactly where the product breaks. This is the habit behind the measured counts on platforms RAITHub built, and the method is in regression testing on every deploy.
Why are test count and coverage misleading?
Because both rise without quality rising. A thousand tests that never touch tenant isolation or billing look impressive and protect nothing. Coverage is worse as a target: it is easy to hit a high number with tests that execute code without checking its behaviour.
Martin Fowler is blunt about coverage as a management metric. It "helps you find which bits of your code aren't being tested", but is "of little value to management since you need a technical background to understand whether the tests are good or whether the uncovered code is a problem" (Martin Fowler, Test Coverage). Use coverage to find untested code, never as a KPI. The real measure of a suite, he argues, is rare bugs escaping to production and the confidence to change code without fear — which is exactly the escaped-defect metric above.
| Vanity metric | Why it misleads | Track instead |
|---|---|---|
| Total test count | Goes up with weak tests; says nothing about what is covered | Escaped defects by layer |
| Overall coverage % | Hit with tests that run code without asserting behaviour | Coverage of billing, permissions and tenant-scoped endpoints specifically |
| Pass rate | 99% green is normal; the 1% and the flaky reruns hide the risk | Flaky-test count and quarantine age |
| Bugs found by QA | Rewards finding many small bugs late over preventing them | Escaped defects and change failure rate |
How should you read these metrics together?
No single number tells the story; the pairs do. Watch how two move together before you act.
- Escaped defects down, pipeline time up. You are buying quality with speed. Fine for a while, but a pipeline that grows every month gets bypassed, so shard and trim before it crosses your budget.
- Change failure rate down, recovery time up. You ship safely but recover slowly. Invest in post-deploy smoke tests and a rehearsed rollback, covered in regression testing on every deploy.
- Pass rate high, escaped defects high. Your tests check the wrong things. Add tests at the layer where defects escape, usually the API layer in a SaaS; see the test pyramid for a SaaS.
- Flaky-test count rising. Trust in the suite is draining. Quarantine with a deadline, not open-ended retries; the flaky end-to-end tests guide covers it.
How do you set targets without gaming them?
Set a direction, not a magic number. "Escaped defects trending down quarter on quarter" is a better target than "95% coverage", because it cannot be hit by writing weak tests. Where you must have a threshold, put it on the specific thing that matters: full coverage of billing, permissions and tenant-scoped endpoints, and a P95 latency ceiling on the checkout path. Review the metrics in a short monthly session and change one thing in response; a dashboard nobody acts on is its own vanity metric.
Buy, build or hire?
| Option | What you get | Choose this when |
|---|---|---|
| A reporting tool or platform | Dashboards for test runs, coverage and DORA metrics | You have someone to read and act on them; the tool only draws the charts |
| Freelancers or crowdtesting | A burst of manual testing; defect counts for one engagement | You need extra eyes once, not an ongoing measure |
| An in-house QA hire | Someone who owns the metrics and the suite | You ship often and can keep one person busy full time |
| A managed QA team (QAaaS) | A team that runs the testing and reports escaped defects and release health weekly | You want the measures acted on, not just displayed |
How RAITHub would test this
RAITHub runs QA as a managed service and reports the outcome metrics, not the vanity ones: escaped defects by layer, change failure rate, recovery time, pipeline time and P95 latency on your key routes, in a weekly report.
- Scope: baseline what is tested now and what escapes to production; add tests at the layers where defects actually get through; wire DORA-style metrics into your pipeline; review them monthly and act on one thing.
- Ways to buy it: a monthly QA plan, a fixed-price one-off audit, or a dedicated QA team that RAITHub manages and bills monthly. Not staff augmentation.
- Timeline: fixed in the written quote after a free audit.
- What you receive: tests and CI in your repository, a weekly QA report, IP assigned to you and an NDA as standard.
Proof is the measured test counts on platforms RAITHub built: 1,024 tests on PropDesk, 750+ on TheSkinProof (the founder's own venture, not a client), 530+ on Sundor Skin, and 400+ on this site. These are counts, shown for discipline; there is no QA case study yet, and none is implied. See the QA as a Service hub and the manual and exploratory testing service, then book the free 15-minute technical audit.
Documentation checked on 9 October 2026.
Frequently asked questions
What are the most important QA metrics?
Escaped defects (bugs found in production), change failure rate, failed-deployment recovery time and pipeline time. For products that sell on speed, add the P95 latency of key routes. These measure outcomes for users, not how much testing was done.
Why is test coverage a bad metric to target?
Coverage rises with tests that execute code without checking its behaviour, so a high percentage can hide weak testing. Use coverage to find untested code, and target full coverage of billing, permissions and tenant-scoped endpoints specifically, not an overall number.
What is an escaped defect?
A bug that reached production when a test or review should have caught it first. Count them per release and tag each with the layer that should have caught it, so you know where to add tests.
What is a good change failure rate?
Lower is better, but the useful signal is the trend for your own team, not a single benchmark. DORA defines it as the share of deployments that need immediate intervention, such as a rollback or hotfix. Watch it fall as your gates improve.
Is test count a useful QA metric?
Not as a goal. A large suite that never tests tenant isolation or billing protects nothing. Count escaped defects by layer instead, and let the test count grow as a by-product of writing a test for every bug.
How often should we review QA metrics?
Monthly is enough for most teams. Keep the session short, read the metrics in pairs, and change one thing in response. A dashboard nobody acts on is itself a vanity metric.
Related posts
Ready to discuss your project?
Book a free 15-minute technical audit with our engineering team.