Founder & Lead Engineer, RAITHub
For most launches, enough performance testing means three things: a load test at two to three times your estimated peak traffic, held for 10 to 15 minutes, with p95 response times and error rates inside written budgets; a short spike test; and Core Web Vitals in Google's "good" range on key pages. Find and fix the first bottleneck, then rerun.
If you would rather have the load test designed and run for you, see how RAITHub would test this near the end.
Why test performance before launch at all?
Because launch-day traffic is the one load you cannot rehearse in production. Bugs that only appear under concurrency (exhausted database connections, slow queries without an index, a third-party API rate limit, a missing cache) do not show up when one developer clicks through staging. A load test lets you meet them on your schedule instead of your customers'.
It does not need to be a big project. A focused test on the five or six endpoints that matter, run against a production-like environment, catches most of what will hurt.
How much load should you test for?
Estimate your peak, then test above it. A worked example, using illustrative numbers:
- Expected visitors in the busiest hour: say 2,000, from the launch email list, ads and press you have planned.
- Requests per visit to your backend: say 15 API calls across sign-up, browsing and one purchase. Read it from your browser's network tab on a real journey.
- Average rate: 2,000 × 15 ÷ 3,600 seconds is about 8 requests per second.
- Peak within the hour: traffic arrives in bursts, after an email or a post. Assume the busiest minute is three to five times the hourly average, so 25 to 40 requests per second.
- Test target: hold the top of that range, then push beyond it to find where it breaks.
The exact numbers matter less than writing them down. A test without a target load proves only that the system handled whatever you happened to send.
Which types of performance test do you need before launch?
Grafana's k6 documentation names six test types (k6: load test types). Not all are needed for a first launch.
| Test type | What it answers | Before launch? |
|---|---|---|
| Smoke | Does the script work, and does the system cope with minimal load? | Yes, always first |
| Average-load | Does it behave under normal expected traffic? | Yes |
| Stress | What happens above expected load, for example at 2 to 3 times peak? | Yes, this is the main one |
| Spike | Does it survive a sudden burst, like a launch email going out? | Yes, if you will announce the launch |
| Soak | Does it degrade over hours: memory leaks, filling disks, growing queues? | Useful; often run after launch |
| Breakpoint | Where exactly is the capacity limit? | Optional; good for capacity planning |
What should the pass or fail budgets be?
Write them before the test, so the result is a decision rather than an argument. A sensible starting set for a web app's API:
- Error rate below 1% of requests at target load.
- p95 response time (95% of requests faster than this) under 500 ms for reads, and under 800 ms to 1 second for writes such as checkout. Use percentiles, not averages; an average hides the slow tail that users feel.
- p99 under about 1.5 seconds, so the worst experiences are bounded too.
- No resource ceiling hit: database CPU and connections, memory, and queue depth stay below roughly 70 to 80% at target load, leaving headroom.
These are engineering defaults, not a standard; tighten them for search or interactive features. For the pages themselves, Google's Core Web Vitals give published targets: Largest Contentful Paint within 2.5 seconds, Interaction to Next Paint of 200 ms or less, and Cumulative Layout Shift of 0.1 or less, measured at the 75th percentile of page loads (web.dev: Web Vitals). Server load tests do not measure those; check them with Lighthouse and real-user monitoring.
What does a minimal k6 load test look like?
k6 scripts are JavaScript. Thresholds turn your budgets into a pass or fail result, so a failing run exits with a non-zero code and can fail a CI job (k6 thresholds docs). This script uses the ramping-arrival-rate executor, which starts requests at a set rate regardless of how fast the system responds, the way real users arrive (k6 ramping-arrival-rate docs):
// load/launch.js: run with k6 run -e BASE_URL=https://staging.example.com load/launch.js
import http from 'k6/http'
import { check } from 'k6'
export const options = {
scenarios: {
launch_peak: {
executor: 'ramping-arrival-rate',
startRate: 5,
timeUnit: '1s',
preAllocatedVUs: 50,
maxVUs: 300,
stages: [
{ target: 15, duration: '3m' }, // normal traffic
{ target: 40, duration: '2m' }, // ramp to estimated peak
{ target: 100, duration: '10m' }, // hold at about 2.5x peak
{ target: 0, duration: '1m' },
],
},
},
thresholds: {
http_req_failed: ['rate<0.01'],
http_req_duration: ['p(95)<500', 'p(99)<1500'],
'http_req_duration{name:checkout}': ['p(95)<800'],
},
}
const BASE = __ENV.BASE_URL
const JSON_HEADERS = { 'Content-Type': 'application/json' }
export default function () {
const list = http.get(BASE + '/api/products', { tags: { name: 'products' } })
check(list, { 'products 200': (r) => r.status === 200 })
const order = http.post(
BASE + '/api/checkout',
JSON.stringify({ sku: 'LOADTEST-1', qty: 1 }),
{ headers: JSON_HEADERS, tags: { name: 'checkout' } },
)
check(order, { 'checkout 201': (r) => r.status === 201 })
}
Point the checkout at the payment provider's sandbox or a stub, never at live payments. Tag each request so the results break down by endpoint, and keep the script in your repository next to your other tests.
What should you watch while the load test runs?
The k6 summary tells you that something got slow. The server side tells you why. Watch these together:
- Database: active connections against the pool limit, slow-query log, CPU, and lock waits. The first bottleneck is usually here.
- Application: CPU, memory, event-loop lag for Node, and error logs.
- Serverless and hosting limits: concurrency caps, cold starts, function time-outs and plan quotas.
- Third parties: rate limits on email, search, AI or payment APIs, which your own code may hit long before your servers struggle. API rate limiting explained covers both sides.
- Caching: hit rates on the CDN and application cache.
Fix the first bottleneck you find, then rerun. Each fix usually exposes the next one. Two or three rounds is normal.
What usually breaks first, and what are the common mistakes?
| Symptom under load | Common cause | Typical fix |
|---|---|---|
| Errors jump suddenly at a certain rate | Database connection pool exhausted | A pooler, a right-sized pool, fewer connections per function |
| p95 climbs steadily with load | A slow query or missing index on a hot path | Add the index, cut N+1 queries, cache the result |
| 429 errors from a provider | A third-party rate limit | Queue and batch the calls; ask for a higher limit |
| Slow first requests after quiet periods | Serverless cold starts | Keep critical functions warm, or move them to a long-running service |
| Memory grows over a long run | A leak or an unbounded cache | Profile, bound the cache, find the leak in a soak test |
The mistakes that make results meaningless: testing a staging environment much smaller than production; testing from the same machine as the server; testing only the home page while the expensive work sits behind login; reusing one test user so everything is cached; and not seeding realistic data volumes, so every query is fast against an empty table.
Doing it yourself takes about 2 to 4 days for someone who knows the stack: a day to write and validate the script, a day or two to run, read and fix, and a rerun. The main risk is a misleading pass from an environment or data set that does not resemble production.
Buy, build or hire?
| Option | Choose this when | Watch out for |
|---|---|---|
| A tool or SaaS testing platform (hosted load generators) | You can write scripts and need large or distributed load | The tool generates traffic; it does not tell you why the system slowed down |
| Freelancers or crowdtesting | You need a one-off test script and report | Little help diagnosing the bottleneck in your code |
| An in-house QA hire | Performance is a permanent concern and you release often | Performance skills combine testing, profiling and infrastructure; rare in one hire |
| A managed QAaaS team | You want the test designed, run, diagnosed, and turned into a CI gate | Agree on a production-like environment before the test, or the result will not mean much |
Why RAITHub for performance testing?
Because the people who run the test also read the code and the database. RAITHub uses k6 for load and latency testing, and the performance work on this site, including cutting a database hosting bill and adding rate limiting without Redis, came from tracing real bottlenecks rather than adding servers. Load testing is one line of the wider pre-launch QA checklist, and the endpoints you load-test should already have functional tests; see the API testing guide. If the test shows the architecture needs real changes, code rescue covers that.
When you don't need RAITHub: if your launch is a waitlist page on a static host, or your expected traffic is a few hundred visitors a day on a managed platform, a smoke test and a Lighthouse run are enough.
How RAITHub would test this
Through QA and test automation, part of QA as a service, RAITHub would:
- agree the target load and written budgets with you, from your launch plan and analytics
- write k6 scripts for the key journeys, with realistic test users and data, and stub live payments
- run smoke, average-load, stress and spike tests against a production-like environment while watching the database, application and providers
- report each bottleneck with its cause and a specific fix, then rerun after fixes
- leave the scripts in your repository and a smaller version in CI as a performance gate
Buy it as a fixed-price pre-launch audit, inside a monthly QA plan, or with a dedicated QA team that RAITHub manages and bills monthly. Testers are never placed under your management. You own the scripts and the IP; an NDA is signed first. Start with a free 15-minute call, then a written fixed quote. Ask for a pre-launch performance test.
Frequently asked questions
How much load testing is enough before launch?
For most products: a smoke test, then a stress test at two to three times your estimated peak held for 10 to 15 minutes, plus a short spike test, all passing written p95 and error-rate budgets.
What is a good p95 response time for a web app?
Under 500 ms for typical API reads and under about a second for heavier writes is a common starting budget. It is an engineering default, not a standard; set budgets that fit your product.
What is the difference between load testing and performance testing?
Performance testing is the umbrella: how fast and stable the system is. Load testing is one kind of it, measuring behaviour at a given number of users or requests. Stress, spike and soak tests are others.
Can I load test in production?
Before launch, test a production-like staging environment instead. After launch, small controlled tests in production are possible, but never against live payments or real customer data.
Is k6 free?
k6 itself is open source and runs locally or in CI (k6 docs). Grafana also sells a hosted cloud service for larger or distributed load, which is optional.
Does a load test check Core Web Vitals?
No. Load tests measure your servers. Core Web Vitals measure the page in the browser; check them with Lighthouse in CI and with real-user monitoring after launch.
Related posts
Ready to discuss your project?
Book a free 15-minute technical audit with our engineering team.