SaaS Backup and Disaster Recovery: RPO, RTO and Restore Drills
Founder & Lead Engineer, RAITHub
A SaaS backup and disaster recovery plan starts with two numbers: RPO, how much recent data you can afford to lose, and RTO, how long you can be down. Turn on point-in-time recovery at your database provider to meet the RPO. Keep a second, independent copy in another account. Then prove the RTO with a timed restore drill every month, because an untested backup is only a hope.
If you would rather have the plan and the drills built for you, see how RAITHub would build this near the end. The rest is the method, written for a small team running PostgreSQL on a managed provider.
What do RPO and RTO mean for a SaaS?
AWS's disaster recovery guidance defines the recovery time objective (RTO) as the maximum acceptable delay between the interruption of service and restoration of service, and the recovery point objective (RPO) as the maximum acceptable amount of time since the last data recovery point (AWS: Disaster Recovery of Workloads on AWS). Both are set by the business, not by the engineers.
In SaaS terms: RPO is "if we restore, how many minutes of customer work disappear?" RTO is "how long are customers locked out while we restore?" The same AWS paper adds a useful brake: if a recovery strategy costs more than the failure it protects against, it should not be put in place unless something else, such as regulation, requires it.
Set the numbers per data store, not per company. Invoices and customer records might need an RPO of minutes. Analytics events you can rebuild from logs might tolerate a day.
What can go wrong, and which protection covers it?
Different disasters need different protections. Replication and high availability help when hardware fails, but they copy mistakes too: a bad DELETE reaches the replica in seconds. AWS makes the same point: continuous replication may not protect against data corruption or malicious deletion as well as point-in-time backups do (AWS: disaster recovery options in the cloud).
| Event | Example | What recovers it |
|---|---|---|
| Hardware or zone failure | The database host or a data centre goes down | Provider high availability (standby or multi-zone) |
| Human or code error | A migration or script deletes or overwrites rows | Point-in-time recovery to just before the error |
| One tenant's data damaged | A customer's admin bulk-deletes records and wants them back | Restore to a separate database, copy that tenant's rows back |
| Account compromise or provider loss | Stolen credentials, a deleted project, a billing lockout | An independent copy in a different account or provider |
| Region-wide outage | A whole cloud region is unavailable | Cross-region copies and infrastructure as code to rebuild |
What do managed Postgres providers give you by default?
Less than many teams assume, and the defaults are worth checking today.
- Amazon RDS. Automated backups can be retained for 0 to 35 days; 0 disables them. The default is seven days when you create the instance in the console, but one day if you create it through the API or CLI without setting it (Amazon RDS backup retention period). RDS uploads transaction logs to S3 every five minutes, and a point-in-time restore creates a new DB instance rather than overwriting the source (Amazon RDS point-in-time restore). Automated backups are deleted with the instance unless you choose to retain them (Amazon RDS backups).
- Neon. The history window that powers instant restore is up to 6 hours (capped at 1 GB of history) on the Free plan, up to 7 days on Launch and up to 30 days on Scale, with a default of 1 day on the paid plans (Neon history window). Shortening it lowers storage cost but limits how far back you can restore.
Two lessons follow. First, a default may be shorter than you think, so set retention on purpose. Second, provider backups live in the same account as production. Anyone who can delete the database may be able to delete its backups too.
Do you still need your own backups if the provider has point-in-time recovery?
Yes, one independent copy. Point-in-time recovery is your main tool for the common case, a bad write you notice within hours. A separate logical backup covers the rare but fatal cases: a compromised or closed account, a deleted project, or a mistake older than your retention window.
For PostgreSQL, pg_dump makes consistent exports even while the database is in use, and does not block readers or writers (PostgreSQL pg_dump). Its custom format (-Fc) is compressed and lets pg_restore restore selectively. Note its limit: PostgreSQL's documentation says logical dumps cannot be used for continuous archiving or WAL replay (PostgreSQL continuous archiving and PITR). A nightly dump therefore gives you an RPO of up to a day; it is the safety net, not the main restore path.
Store the dump in a different cloud account, with write-only credentials for the job that uploads it and object versioning or a retention lock so the job cannot delete old copies. AWS notes that cross-account backup copies help protect against insider threats or account compromise (same AWS paper as above).
What else has to be backed up besides the database?
- Uploaded files in object storage: turn on versioning, and replicate or copy to a second location.
- Secrets and configuration: environment variables, API keys, DNS records and OAuth app settings, kept in a password manager or secrets store you can reach if the main cloud account is gone.
- Infrastructure as code, so you can rebuild servers, queues and networks rather than click them together from memory during an outage.
- Source code and CI configuration, mirrored outside your main Git host if losing it would stop you shipping.
- Third-party data you depend on. If Stripe is the source of truth for subscriptions, your restore plan must reconcile your database with it after a restore. See testing payments and webhooks for replaying events.
Which disaster recovery strategy fits a SaaS of your size?
AWS describes four broad strategies, in rising cost and complexity: backup and restore, pilot light, warm standby, and multi-site active/active (AWS: disaster recovery options in the cloud).
| Strategy | What runs in the recovery location | Fits |
|---|---|---|
| Backup and restore | Nothing until needed; data copies plus infrastructure as code | Most early and mid-stage SaaS products |
| Pilot light | Replicated data; app servers defined but switched off | Products with enterprise contracts that name an RTO in hours |
| Warm standby | A scaled-down, working copy that can take traffic | Products where hours of downtime lose major customers |
| Multi-site active/active | Full production in more than one region | Large platforms with near-zero downtime commitments |
The same paper notes that a multi-zone setup within one region may already cover much of the risk, and that even active/active setups still need point-in-time backups, because corrupted data replicates everywhere. For most SaaS teams, backup and restore done well, with tested drills, is the right first step.
How do you restore a single tenant without rolling back everyone?
Restore to a new database, never over production, then copy back only that tenant's rows. On a shared schema with a tenant_id column, that means exporting the tenant's rows from the restored copy and upserting them into production, in foreign-key order. Rows the customer created after the restore point stay untouched. Design for this early: every tenant-scoped table should carry tenant_id directly, which also helps isolation (see the row-level security guide and database per tenant vs shared schema).
What does a restore drill look like?
Restore the latest backup into a scratch database, run checks that prove the data is real and recent, time it, and record the result. Run it monthly, ideally from CI so it cannot be forgotten. AWS's guidance is blunt: a backup strategy must include testing your backups.
#!/usr/bin/env bash
# scripts/restore-drill.sh: restore the latest dump into a scratch DB and check it.
# Needs: BACKUP_FILE (a pg_dump -Fc file) and SCRATCH_URL (an empty database).
set -euo pipefail
start=$(date +%s)
pg_restore --no-owner --no-acl --exit-on-error \
--dbname="$SCRATCH_URL" "$BACKUP_FILE"
restore_seconds=$(( $(date +%s) - start ))
# 1. Core tables are not empty.
accounts=$(psql "$SCRATCH_URL" -tAc "SELECT count(*) FROM accounts")
invoices=$(psql "$SCRATCH_URL" -tAc "SELECT count(*) FROM invoices")
# 2. Data is recent: newest invoice within the expected RPO (26 hours here).
age_hours=$(psql "$SCRATCH_URL" -tAc \
"SELECT floor(extract(epoch FROM now() - max(created_at)) / 3600) FROM invoices")
echo "restore_seconds=$restore_seconds accounts=$accounts invoices=$invoices age_hours=$age_hours"
if [ "$accounts" -eq 0 ] || [ "$invoices" -eq 0 ]; then
echo "FAIL: empty core table" >&2; exit 1
fi
if [ "$age_hours" -gt 26 ]; then
echo "FAIL: newest data is $age_hours hours old" >&2; exit 1
fi
echo "PASS"
Adjust the table names and the age limit to your schema and dump schedule. For point-in-time recovery, run a second drill through your provider's restore feature into a new instance or branch, with the same checks. Record the time each drill takes: that number, not a guess, is your real RTO. If the drill fails, treat it like a production incident.
How long does it take to set this up yourself?
About two to four days for an engineer who knows your cloud account: set retention, schedule the dump to a second account, write the drill script, run it once by hand, then schedule it. The main risk of doing it yourself is stopping after the first successful drill. Schemas change, credentials expire and storage fills up, and backups that worked in March fail quietly by June.
Buy, build or hire?
| Option | What you get | Choose this when |
|---|---|---|
| Your provider's built-in backups | Automated backups and point-in-time recovery with a few settings | Always, as the base layer; check retention and who can delete it |
| A backup tool or service | Scheduled copies to another location, sometimes with restore testing | You want an independent copy without writing scripts, and the tool supports your database |
| Custom build (in-house or hired) | Cross-account copies, tenant-level restore, scheduled drills in CI and a written runbook | Customers ask about RPO and RTO, or a lost tenant would cost you the account |
Why RAITHub for this
RAITHub treats restores as something to test, like any other code path. Its products are built on PostgreSQL with tenant-scoped data, such as Sundor Skin's 146 tables with row-level security, which is the design that makes single-tenant restores practical. The SaaS development service includes a backup plan and restore drill at launch. RAITHub has no published disaster recovery incident to point to, so this guide is engineering practice, not a case study.
When you don't need us
- You are pre-launch with test data. Turn on your provider's backups, set retention, and come back to this before the first paying customer.
- Your provider offers restore testing and you have time to set it up and watch it. Use it.
- You need a certified business continuity programme for an audit. RAITHub is not SOC 2 or ISO 27001 certified; it can build the technical controls, but the certification needs an auditor.
How RAITHub would build this
- Scope: agree RPO and RTO per data store; set provider retention and point-in-time recovery; add a nightly logical dump to a separate, locked-down account; build a single-tenant restore script; schedule a restore drill in CI with alerts on failure.
- Timeline: included in a new SaaS build (4–6 weeks for a fixed-scope MVP), or as a fixed-scope job on an existing product, confirmed in the written quote.
- What you receive: scripts and CI jobs in your repository, a disaster recovery runbook with the drill results, tests, IP assigned to you and an NDA as standard.
- Next step: a free 15-minute technical audit, then a fixed written quote.
To find out how long a restore would really take you, book the free 15-minute technical audit.
Documentation checked on 7 October 2026.
Frequently asked questions
What is the difference between RPO and RTO?
RPO is how much recent data you can afford to lose, measured as time since the last recovery point. RTO is how long the service can be down before it is restored. Both are business decisions.
How often should a SaaS test its backups?
At least monthly, and after any major schema or infrastructure change. Automate the drill in CI so it runs without anyone remembering, and alert when it fails.
Is database replication a backup?
No. Replication copies every change, including accidental deletes and corrupted data, within seconds. It protects against hardware failure. You still need point-in-time backups to go back to before a mistake.
How long should backups be kept?
Long enough to cover the time it might take to notice a problem, plus any contract or legal requirement. Many teams keep point-in-time recovery for one to four weeks and longer-term logical copies for months.
Can I restore one customer's data without affecting others?
Yes, if every tenant-scoped table carries a tenant ID. Restore the backup into a separate database, then copy that tenant's rows back into production. Never restore over the live database for a single tenant.
What should a SaaS tell customers about backups?
The RPO and RTO you can actually show from drills, where backups are stored, and how a restore request works. Promise only what your drill results support, and put the numbers in your SLA or security documentation.
Related posts
Technical SEO Checklist for 2026: The Foundation That Lets You Rank
7 min readLocal and Geo SEO for Service Businesses: Rank Where Your Customers Are
7 min readSaaS Entitlements: Enforcing Plans, Limits and Add-ons in Code
14 min readReady to discuss your project?
Book a free 15-minute technical audit with our engineering team.