Back to BlogArchitecture & Engineering

SaaS Backup and Disaster Recovery: RPO, RTO and Restore Drills

Rupak Amin

Founder & Lead Engineer, RAITHub

12 min read

A SaaS backup and disaster recovery plan starts with two numbers: RPO, how much recent data you can afford to lose, and RTO, how long you can be down. Turn on point-in-time recovery at your database provider to meet the RPO. Keep a second, independent copy in another account. Then prove the RTO with a timed restore drill every month, because an untested backup is only a hope.

If you would rather have the plan and the drills built for you, see how RAITHub would build this near the end. The rest is the method, written for a small team running PostgreSQL on a managed provider.

What do RPO and RTO mean for a SaaS?

AWS's disaster recovery guidance defines the recovery time objective (RTO) as the maximum acceptable delay between the interruption of service and restoration of service, and the recovery point objective (RPO) as the maximum acceptable amount of time since the last data recovery point (AWS: Disaster Recovery of Workloads on AWS). Both are set by the business, not by the engineers.

In SaaS terms: RPO is "if we restore, how many minutes of customer work disappear?" RTO is "how long are customers locked out while we restore?" The same AWS paper adds a useful brake: if a recovery strategy costs more than the failure it protects against, it should not be put in place unless something else, such as regulation, requires it.

Set the numbers per data store, not per company. Invoices and customer records might need an RPO of minutes. Analytics events you can rebuild from logs might tolerate a day.

What can go wrong, and which protection covers it?

Different disasters need different protections. Replication and high availability help when hardware fails, but they copy mistakes too: a bad DELETE reaches the replica in seconds. AWS makes the same point: continuous replication may not protect against data corruption or malicious deletion as well as point-in-time backups do (AWS: disaster recovery options in the cloud).

EventExampleWhat recovers it
Hardware or zone failureThe database host or a data centre goes downProvider high availability (standby or multi-zone)
Human or code errorA migration or script deletes or overwrites rowsPoint-in-time recovery to just before the error
One tenant's data damagedA customer's admin bulk-deletes records and wants them backRestore to a separate database, copy that tenant's rows back
Account compromise or provider lossStolen credentials, a deleted project, a billing lockoutAn independent copy in a different account or provider
Region-wide outageA whole cloud region is unavailableCross-region copies and infrastructure as code to rebuild

What do managed Postgres providers give you by default?

Less than many teams assume, and the defaults are worth checking today.

  • Amazon RDS. Automated backups can be retained for 0 to 35 days; 0 disables them. The default is seven days when you create the instance in the console, but one day if you create it through the API or CLI without setting it (Amazon RDS backup retention period). RDS uploads transaction logs to S3 every five minutes, and a point-in-time restore creates a new DB instance rather than overwriting the source (Amazon RDS point-in-time restore). Automated backups are deleted with the instance unless you choose to retain them (Amazon RDS backups).
  • Neon. The history window that powers instant restore is up to 6 hours (capped at 1 GB of history) on the Free plan, up to 7 days on Launch and up to 30 days on Scale, with a default of 1 day on the paid plans (Neon history window). Shortening it lowers storage cost but limits how far back you can restore.

Two lessons follow. First, a default may be shorter than you think, so set retention on purpose. Second, provider backups live in the same account as production. Anyone who can delete the database may be able to delete its backups too.

Do you still need your own backups if the provider has point-in-time recovery?

Yes, one independent copy. Point-in-time recovery is your main tool for the common case, a bad write you notice within hours. A separate logical backup covers the rare but fatal cases: a compromised or closed account, a deleted project, or a mistake older than your retention window.

For PostgreSQL, pg_dump makes consistent exports even while the database is in use, and does not block readers or writers (PostgreSQL pg_dump). Its custom format (-Fc) is compressed and lets pg_restore restore selectively. Note its limit: PostgreSQL's documentation says logical dumps cannot be used for continuous archiving or WAL replay (PostgreSQL continuous archiving and PITR). A nightly dump therefore gives you an RPO of up to a day; it is the safety net, not the main restore path.

Store the dump in a different cloud account, with write-only credentials for the job that uploads it and object versioning or a retention lock so the job cannot delete old copies. AWS notes that cross-account backup copies help protect against insider threats or account compromise (same AWS paper as above).

What else has to be backed up besides the database?

  • Uploaded files in object storage: turn on versioning, and replicate or copy to a second location.
  • Secrets and configuration: environment variables, API keys, DNS records and OAuth app settings, kept in a password manager or secrets store you can reach if the main cloud account is gone.
  • Infrastructure as code, so you can rebuild servers, queues and networks rather than click them together from memory during an outage.
  • Source code and CI configuration, mirrored outside your main Git host if losing it would stop you shipping.
  • Third-party data you depend on. If Stripe is the source of truth for subscriptions, your restore plan must reconcile your database with it after a restore. See testing payments and webhooks for replaying events.

Which disaster recovery strategy fits a SaaS of your size?

AWS describes four broad strategies, in rising cost and complexity: backup and restore, pilot light, warm standby, and multi-site active/active (AWS: disaster recovery options in the cloud).

StrategyWhat runs in the recovery locationFits
Backup and restoreNothing until needed; data copies plus infrastructure as codeMost early and mid-stage SaaS products
Pilot lightReplicated data; app servers defined but switched offProducts with enterprise contracts that name an RTO in hours
Warm standbyA scaled-down, working copy that can take trafficProducts where hours of downtime lose major customers
Multi-site active/activeFull production in more than one regionLarge platforms with near-zero downtime commitments

The same paper notes that a multi-zone setup within one region may already cover much of the risk, and that even active/active setups still need point-in-time backups, because corrupted data replicates everywhere. For most SaaS teams, backup and restore done well, with tested drills, is the right first step.

How do you restore a single tenant without rolling back everyone?

Restore to a new database, never over production, then copy back only that tenant's rows. On a shared schema with a tenant_id column, that means exporting the tenant's rows from the restored copy and upserting them into production, in foreign-key order. Rows the customer created after the restore point stay untouched. Design for this early: every tenant-scoped table should carry tenant_id directly, which also helps isolation (see the row-level security guide and database per tenant vs shared schema).

What does a restore drill look like?

Restore the latest backup into a scratch database, run checks that prove the data is real and recent, time it, and record the result. Run it monthly, ideally from CI so it cannot be forgotten. AWS's guidance is blunt: a backup strategy must include testing your backups.

#!/usr/bin/env bash
# scripts/restore-drill.sh: restore the latest dump into a scratch DB and check it.
# Needs: BACKUP_FILE (a pg_dump -Fc file) and SCRATCH_URL (an empty database).
set -euo pipefail

start=$(date +%s)

pg_restore --no-owner --no-acl --exit-on-error \
  --dbname="$SCRATCH_URL" "$BACKUP_FILE"

restore_seconds=$(( $(date +%s) - start ))

# 1. Core tables are not empty.
accounts=$(psql "$SCRATCH_URL" -tAc "SELECT count(*) FROM accounts")
invoices=$(psql "$SCRATCH_URL" -tAc "SELECT count(*) FROM invoices")

# 2. Data is recent: newest invoice within the expected RPO (26 hours here).
age_hours=$(psql "$SCRATCH_URL" -tAc \
  "SELECT floor(extract(epoch FROM now() - max(created_at)) / 3600) FROM invoices")

echo "restore_seconds=$restore_seconds accounts=$accounts invoices=$invoices age_hours=$age_hours"

if [ "$accounts" -eq 0 ] || [ "$invoices" -eq 0 ]; then
  echo "FAIL: empty core table" >&2; exit 1
fi
if [ "$age_hours" -gt 26 ]; then
  echo "FAIL: newest data is $age_hours hours old" >&2; exit 1
fi
echo "PASS"

Adjust the table names and the age limit to your schema and dump schedule. For point-in-time recovery, run a second drill through your provider's restore feature into a new instance or branch, with the same checks. Record the time each drill takes: that number, not a guess, is your real RTO. If the drill fails, treat it like a production incident.

How long does it take to set this up yourself?

About two to four days for an engineer who knows your cloud account: set retention, schedule the dump to a second account, write the drill script, run it once by hand, then schedule it. The main risk of doing it yourself is stopping after the first successful drill. Schemas change, credentials expire and storage fills up, and backups that worked in March fail quietly by June.

Buy, build or hire?

OptionWhat you getChoose this when
Your provider's built-in backupsAutomated backups and point-in-time recovery with a few settingsAlways, as the base layer; check retention and who can delete it
A backup tool or serviceScheduled copies to another location, sometimes with restore testingYou want an independent copy without writing scripts, and the tool supports your database
Custom build (in-house or hired)Cross-account copies, tenant-level restore, scheduled drills in CI and a written runbookCustomers ask about RPO and RTO, or a lost tenant would cost you the account

Why RAITHub for this

RAITHub treats restores as something to test, like any other code path. Its products are built on PostgreSQL with tenant-scoped data, such as Sundor Skin's 146 tables with row-level security, which is the design that makes single-tenant restores practical. The SaaS development service includes a backup plan and restore drill at launch. RAITHub has no published disaster recovery incident to point to, so this guide is engineering practice, not a case study.

When you don't need us

  • You are pre-launch with test data. Turn on your provider's backups, set retention, and come back to this before the first paying customer.
  • Your provider offers restore testing and you have time to set it up and watch it. Use it.
  • You need a certified business continuity programme for an audit. RAITHub is not SOC 2 or ISO 27001 certified; it can build the technical controls, but the certification needs an auditor.

How RAITHub would build this

  • Scope: agree RPO and RTO per data store; set provider retention and point-in-time recovery; add a nightly logical dump to a separate, locked-down account; build a single-tenant restore script; schedule a restore drill in CI with alerts on failure.
  • Timeline: included in a new SaaS build (4–6 weeks for a fixed-scope MVP), or as a fixed-scope job on an existing product, confirmed in the written quote.
  • What you receive: scripts and CI jobs in your repository, a disaster recovery runbook with the drill results, tests, IP assigned to you and an NDA as standard.
  • Next step: a free 15-minute technical audit, then a fixed written quote.

To find out how long a restore would really take you, book the free 15-minute technical audit.

Documentation checked on 7 October 2026.

Frequently asked questions

What is the difference between RPO and RTO?

RPO is how much recent data you can afford to lose, measured as time since the last recovery point. RTO is how long the service can be down before it is restored. Both are business decisions.

How often should a SaaS test its backups?

At least monthly, and after any major schema or infrastructure change. Automate the drill in CI so it runs without anyone remembering, and alert when it fails.

Is database replication a backup?

No. Replication copies every change, including accidental deletes and corrupted data, within seconds. It protects against hardware failure. You still need point-in-time backups to go back to before a mistake.

How long should backups be kept?

Long enough to cover the time it might take to notice a problem, plus any contract or legal requirement. Many teams keep point-in-time recovery for one to four weeks and longer-term logical copies for months.

Can I restore one customer's data without affecting others?

Yes, if every tenant-scoped table carries a tenant ID. Restore the backup into a separate database, then copy that tenant's rows back into production. Never restore over the live database for a single tenant.

What should a SaaS tell customers about backups?

The RPO and RTO you can actually show from drills, where backups are stored, and how a restore request works. Promise only what your drill results support, and put the numbers in your SLA or security documentation.

saas backupdisaster recoveryRPO and RTOpoint-in-time recoveryrestore drillPostgreSQL backup

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.