Founder & Lead Engineer, RAITHub
RAITHub ships and tests production software. See QA as a Service or talk to us.
A SaaS that was fast at launch and slows as customers grow is almost always hitting the database, not running out of servers. Find which requests are slow and why, read the slow-query log, fix missing indexes and N+1 query patterns, then cache what repeats. Add hardware last: a bigger server only hides the real problem for one more growth stage.
If you would rather have the diagnosis and the fix done for you, see how RAITHub would fix this below. This is a prioritised order for founders and small teams whose product works but is getting slower as data and users pile up. It is the SaaS-scale companion to the general why a Next.js app is slow in production guide, focused on what changes when a multi-tenant database fills up.
Why does a SaaS slow down only after it grows?
Because the slow parts were invisible when the tables were small. A query that scans a whole table is instant on 500 rows and painful on 5 million. Nothing in the code changed; the data under it did.
At launch, almost any query returns fast because the database holds everything in memory and a full scan of a tiny table costs nothing. As each customer adds rows, three things happen at once: scans take longer, the working set stops fitting in memory, and small per-request inefficiencies multiply by every customer hitting them. The symptom is a product that feels fine on your test account and slow on your biggest customer's.
So the first instinct, buy a bigger server, treats the symptom. It works for a while because more memory and CPU paper over an unindexed query, which is exactly why the real cause stays hidden until the next growth stage brings it back, now more expensive to fix.
In what order should you diagnose a slow SaaS?
From the quickest and most revealing to the most expensive. Measure before you change anything, so you fix the request that is actually slow, not the one you assume is.
| Step | What to do | Why it comes here | How to verify |
|---|---|---|---|
| 1 | Measure which endpoints are slow, and at what percentile | You cannot fix what you have not measured; averages hide the slow tail | A dashboard of P95 latency per route; pick the worst three |
| 2 | Read the database slow-query log for those routes | Most SaaS slowness is one or two queries, not the whole app | EXPLAIN ANALYZE on the top query shows a sequential scan or a huge row count |
| 3 | Fix N+1 queries and missing indexes | The lowest-effort, largest wins; often milliseconds instead of seconds | Query count per request drops; EXPLAIN shows an index scan |
| 4 | Cache what is read often and changes rarely | Removes repeated work the database should not keep doing | Database stays idle between writes; cache hit rate is high |
| 5 | Add connection pooling and tune the database | Serverless opens many short connections; Postgres has a limit | No "too many connections" errors; stable connection count |
| 6 | Scale hardware or add read replicas | Only once the queries are efficient and still saturate one machine | Before and after numbers for the saturated resource |
Nine times out of ten the fix is in steps 2 to 4. Hardware in step 6 is real, but it is the last lever, not the first.
How do you find the one query that is slow?
Turn on slow-query logging, reproduce the slow page, and read the log. PostgreSQL will record any statement over a threshold you set, which turns "the app feels slow" into "this exact query took 4.2 seconds".
-- Log every statement slower than 500 ms (session or postgresql.conf)
SET log_min_duration_statement = 500;
-- Then run the slow request and inspect the worst query with:
EXPLAIN (ANALYZE, BUFFERS)
SELECT * FROM invoices
WHERE tenant_id = $1
ORDER BY created_at DESC
LIMIT 50;
In the plan, two words tell you most of the story. Seq Scan on a large table means no usable index. A row count in the thousands feeding into a small LIMIT means the database read far more than it returned. The PostgreSQL EXPLAIN documentation describes how to read the rest. On a multi-tenant SaaS, always include the tenant filter in the query you test, because your own small test tenant can look fast while a large one is slow.
What usually causes the slowdown: indexes and N+1 queries
Two patterns cause most SaaS slowness. A missing index turns a lookup into a full-table scan. An N+1 query runs one query to fetch a list, then one more query per item in the list, so a page showing 100 rows quietly runs 101 queries.
The index fix is usually a single statement, built without locking the table by using CONCURRENTLY:
-- A tenant's invoices, newest first, is a common SaaS access pattern.
-- This composite index serves the WHERE and the ORDER BY together.
CREATE INDEX CONCURRENTLY IF NOT EXISTS idx_invoices_tenant_created
ON invoices (tenant_id, created_at DESC);
The order of columns matters: filter columns first, then the column you sort by. Why indexes work and when they do not is covered in the PostgreSQL indexing guide. For the N+1 pattern, fetch related rows in one query (a join or a single WHERE id IN (...)) instead of looping. Most ORMs have an "include" or "eager load" option that does this; the fix is to turn it on for the list endpoints, not to loop in application code.
Indexes are not free: every index slows writes a little and uses storage. Add the ones your slow queries ask for, not one per column. On Sundor Skin, a B2B wholesale platform RAITHub built with 146 PostgreSQL tables and 530+ tests, tenant-scoped indexes and row-level security are part of the schema from the first migration rather than retrofitted under load.
When should you cache, and what should you cache?
Cache data that is read far more often than it changes, and read on pages your customers hit constantly: dashboards, lists, public pages. The goal is to stop the database doing the same work on every request.
This also protects your bill, not just your latency. On a serverless database that sleeps when idle, anything touching it on every request keeps it awake and metering. This site cut its own database cost by caching the pages crawlers hit so ordinary traffic no longer woke the database, as described in what a SaaS costs to run per month. Cache invalidation is the hard part: cache by a key that changes when the underlying data changes, and keep the cache window short enough that stale data is a minor annoyance, not a correctness bug. Never cache one tenant's data under a key another tenant can read.
What about connections and hardware?
Serverless functions each open their own database connection, and Postgres caps how many it allows. Past a few hundred customers you can exhaust the limit before you exhaust the CPU, which looks like slowness or outright errors under load.
A connection pooler (such as the one Neon and most managed Postgres providers offer) sits between your app and the database and reuses a small set of connections. Turn it on before you reach for a bigger machine. Only after queries are indexed, hot reads are cached and pooling is in place does hardware become the answer: a larger instance, or a read replica to take reporting and dashboard traffic off the primary. Scaling a well-tuned system is covered in scaling an MVP to production, where performance work is deliberately the last priority.
How long does fixing a slow SaaS take yourself?
If you can read a query plan, the first big win is often a day: measure, find the worst query, add an index, confirm the plan changed. A full pass across the worst endpoints is usually one to two weeks for one engineer. The main risk of doing it yourself is adding an index that locks a large table during business hours (use CONCURRENTLY), or caching tenant data under a shared key and leaking it between customers.
Buy, build or hire?
| Option | What you get | Choose this when |
|---|---|---|
| Managed database with built-in insights | A pooler and a slow-query dashboard from your database provider | You want the measurement tools without setting them up, and your team can act on them |
| An APM or monitoring tool | Per-route latency, traces and database timing in one place | You need to find the slow routes first and have nobody watching latency today |
| Fix it in-house | Your own engineer indexes, de-N+1s and caches the worst paths | Someone on the team can read a query plan and has the time |
| Hire a fixed-scope performance pass | A measured before/after, the fixes in your repo with tests | The slowness is costing you customers and nobody in-house has the time or the Postgres depth |
How RAITHub would fix this
Because RAITHub runs this stack in production and has fixed its own cost and performance leaks, not just read about them.
- Scope: measure P95 latency per route; read the slow-query log and run EXPLAIN on the worst queries; add the indexes they need, concurrently; remove N+1 patterns on the list endpoints; cache hot, slow-changing reads with tenant-safe keys; add pooling; recommend hardware or a read replica only if the tuned system still saturates.
- Timeline: a performance pass fits a code rescue engagement (2–4 weeks), or a fixed-scope hardening job inside a backend engagement (6–12 weeks for larger work), confirmed in the written quote.
- You receive: before-and-after numbers for each endpoint changed, the fixes and new indexes in your repository, regression and performance tests gated in CI, handover docs, and full IP under NDA.
- Proof: this site runs on Next.js and Neon with 400+ tests and a documented database-cost fix; Sundor Skin has tenant-scoped indexes and row-level security across 146 tables with 530+ tests; PropDesk has 1,024 tests. See the SaaS development service.
When you don't need us
- One obvious query is slow and you can read its plan. Add the index and confirm, no outside help needed.
- Your data is still small. If your largest table is in the thousands of rows, the slowness is elsewhere; profile before you optimise the database.
- You want engineers placed in your team. RAITHub offers fixed-scope work and dedicated teams, not staff augmentation.
If your product is getting slower as it grows and you do not know which query to blame, send RAITHub your slowest page and your stack, and book the free 15-minute technical audit.
Documentation checked on 11 October 2026.
Frequently asked questions
Why is my SaaS slow only for big customers?
Because the slow queries scan tables that grow with each customer's data. Your small test account hits few rows and feels fast; a large tenant hits millions and triggers full-table scans. Test every slow query with a large tenant's ID, not your own account.
Should I add a bigger server to fix a slow SaaS?
Only after the queries are efficient. A bigger machine hides an unindexed query for one more growth stage, then the latency and the bill both return. Measure, index and cache first; scale hardware last, with before-and-after numbers.
What is an N+1 query and how do I know I have one?
It is running one query for a list, then one more per item, so a 100-row page runs 101 queries. You spot it by counting queries per request in your logs or APM. Fix it by fetching related rows in one query with a join or an IN clause.
How do I find the slowest query in PostgreSQL?
Set log_min_duration_statement to log statements over a threshold such as 500 ms, reproduce the slow page, then run EXPLAIN (ANALYZE, BUFFERS) on the logged query. A sequential scan on a large table or a huge row count feeding a small LIMIT points at the fix.
Will adding indexes slow down my writes?
A little, and each index uses storage, so add only the indexes your slow queries need, not one per column. Build them with CREATE INDEX CONCURRENTLY so the build does not lock the table during business hours.
Can caching leak one tenant's data to another?
Yes, if you cache under a shared key. Always include the tenant ID in the cache key, keep windows short, and never cache a response that mixes tenants. On a multi-tenant SaaS, a caching bug can be a data-exposure bug.
Related posts
Ready to discuss your project?
Book a free 15-minute technical audit with our engineering team.