Back to BlogArchitecture & Engineering

SaaS SLAs, Uptime and Status Pages: What to Promise Customers

Rupak Amin

Founder & Lead Engineer, RAITHub

11 min read

Most early SaaS products should promise 99.5% to 99.9% monthly uptime, and only after measuring that they already beat it. Set a stricter internal target than the one in the contract, define exactly what counts as downtime, exclude announced maintenance, and offer service credits rather than open-ended liability. Publish a status page so customers hear about incidents from you first.

If you would rather have the monitoring, health checks and status page built for you, see how RAITHub would build this near the end. Contract wording is general information here; confirm the legal terms with your adviser.

What is the difference between an SLI, an SLO and an SLA?

Google's SRE book gives the standard definitions (Google SRE book, Service Level Objectives):

  • SLI (service level indicator): a carefully defined quantitative measure of some aspect of the service, such as the share of requests that succeed.
  • SLO (service level objective): a target value or range for an SLI, such as 99.9% of requests succeeding over 30 days.
  • SLA (service level agreement): an explicit or implicit contract with users that includes consequences of meeting or missing the SLOs it contains.

Put simply: the SLI is what you measure, the SLO is what you aim for, and the SLA is what you pay for missing. The same chapter recommends keeping a safety margin by using a tighter internal SLO than the one you advertise, so you can react to problems before customers see a breach.

How much downtime does each uptime percentage allow?

Less than most people expect at the top end. Figures below assume a 30-day month and a 365-day year.

UptimeDowntime per 30-day monthDowntime per yearRealistic for
99%7 h 12 min87.6 hInternal tools, early betas
99.5%3 h 36 min43.8 hAn early SaaS on a single database instance
99.9%43.2 min8.76 hA mature SaaS with high-availability database and tested deploys
99.95%21.6 min4.38 hTeams with on-call cover and redundancy across zones
99.99%4.32 min52.56 minLarge platforms with dedicated reliability engineering

The 99.99% yearly figure matches Google's own example: a system with a 99.99% target can be down for up to 52.56 minutes in a year (Google SRE book, Embracing Risk). One slow deploy or one bad migration can use a whole month's allowance at 99.95% and above.

Why can't your SLA be higher than your dependencies?

Because your service is down when something it needs is down. If your app depends on one database, one host and one auth provider, each failing independently, your best-case availability is roughly the product of theirs, and lower than any one of them.

Check the published SLAs of what you run on. Amazon RDS, for example, offers service credits when a Multi-AZ instance falls below 99.95% monthly uptime, but a Single-AZ instance only below 99.5% (Amazon RDS service level agreement). A SaaS running on a single database instance has no business promising 99.95%. Note too that a provider SLA is a credit on their bill, not compensation for your customers' losses.

Should an early SaaS offer an SLA at all?

Only when a customer asks for one, and then a modest one. Google's SRE book argues that 100% is probably never the right reliability target: it is impossible, and typically more reliability than users want or notice (Embracing Risk). The same book warns against overachieving: if real performance is far better than the stated objective, users come to rely on the real performance (Service Level Objectives).

A practical path:

  1. Measure first. Run uptime monitoring and an SLI for at least two or three months before promising anything.
  2. Set an internal SLO slightly above what you will put in contracts, such as 99.9% internally for a 99.5% SLA.
  3. Offer the SLA on paid tiers that need it, usually business or enterprise plans, not the self-serve tier.
  4. Use the gap as an error budget. The difference between target and actual is the budget of unreliability left for the period (Embracing Risk). When it is nearly spent, slow releases and fix reliability; when there is plenty left, ship.

What should a SaaS SLA document define?

Most SLA disputes come from vague definitions, not from outages. A short SLA should pin down:

TermWhat to defineCommon choice for an early SaaS
Covered serviceWhich parts count: web app, API, background jobsWeb app and public API; not beta features
DowntimeWhat counts as "down", measured how and from whereHealth check failing from multiple regions for consecutive minutes, or error rate above a threshold
Measurement periodMonthly or quarterlyCalendar month
ExclusionsAnnounced maintenance, customer-caused issues, third-party outages beyond your control, force majeureMaintenance announced in advance and capped per month
RemedyService credits as a share of the monthly fee, by tierCredits only, capped, with credits as the sole remedy
Claim processHow and by when a customer claims creditsWritten request within a set number of days
Support responseTime to first response by severity (not time to fix)Separate from uptime, by plan

This is general information, not legal advice; confirm SLA wording, liability caps and the interaction with your terms of service with your adviser.

How do you measure uptime honestly?

Two signals, used together. External checks call a health endpoint every minute from several regions; they catch DNS, certificate and hosting failures. Request-based SLIs count the share of real requests that succeed (non-5xx, under a latency threshold) from your logs or metrics; they catch partial outages that a single health check misses.

A good health endpoint checks what users need, quickly, without leaking detail:

// app/api/health/route.ts (Next.js App Router)
import { NextResponse } from 'next/server'
import { db } from '@/lib/db'

export const dynamic = 'force-dynamic'

export async function GET() {
  const started = Date.now()
  try {
    // A trivial query proves the app can reach the database.
    await Promise.race([
      db.$queryRaw`SELECT 1`,
      new Promise((_, reject) => setTimeout(() => reject(new Error('timeout')), 2000)),
    ])
    return NextResponse.json({ status: 'ok', ms: Date.now() - started })
  } catch {
    return NextResponse.json({ status: 'degraded' }, { status: 503 })
  }
}

Keep it cheap, since it runs every minute from every checker. Do not call slow third-party APIs from it; monitor those separately so a payment provider's hiccup does not mark your whole app as down. Count a minute as down only when checks from more than one region fail, to filter out network blips between one checker and your host.

What should a SaaS status page show?

Components that customers recognise, their current state, active incidents with timed updates, scheduled maintenance and recent history. Atlassian Statuspage's component states are a common vocabulary: degraded performance means working but slow or impacted in a minor way; partial outage means completely broken for a subset of customers; major outage means completely unavailable; under maintenance means being worked on (Atlassian Statuspage: what is a component).

  • Name components by what customers do: "Sign-in", "Dashboard", "API", "Email notifications", "Payments", not internal service names.
  • Host it separately from your app, on a different provider or a hosted status service, so it stays up when you are down.
  • Post the first update within minutes, even if it only says you are investigating. Then update on a fixed rhythm until resolved.
  • Let customers subscribe by email or webhook, so enterprise customers' own teams hear directly.
  • Do not automate the headline from a single check. Automated alerts should page your team; a person should set the public status.

Pair the status page with your release process: announce maintenance there, and link it from the release checklist in a SaaS release process for weekly releases.

Which engineering work actually raises uptime?

Your cloud provider already covers its own layer with its SLA. The part of uptime you control most directly is your own release path, and that is where the least costly gains usually are:

How long does it take to set this up yourself?

About two to five days for one engineer: a health endpoint, external uptime checks from several regions, a request-based SLI dashboard, alerting to whoever is on call, a hosted status page with components, and a one-page SLA draft for your adviser to review. The main risk of doing it yourself is signing an SLA before you have measured anything, then discovering the number you promised is one your stack cannot reach.

Buy, build or hire?

OptionWhat you getChoose this when
Off-the-shelf monitoring and status page serviceHosted multi-region checks, alerting and a status page in an afternoonAlmost always, for the checks and the page itself
Template or open-source status pageA self-hosted page you controlYou want no extra subscription and can host it away from your main app
Custom build (in-house or hired)Request-based SLIs, error budget dashboards, health checks per component and release gates tied to the budgetYou are signing SLAs with enterprise customers and need numbers you can defend

Why RAITHub for this

RAITHub builds SaaS products with the release safety that uptime depends on: gated CI, migrations replayed on every change (all 76 on Sundor Skin) and large automated suites, such as 1,024 tests on PropDesk. The SaaS development service adds health checks, monitoring and a status page at launch. RAITHub does not draft legal terms, and its working hours are agreed per client rather than round the clock, so it says so up front.

When you don't need us

  • No customer has asked for an SLA yet. Set up a hosted uptime monitor and status page yourself, measure for a few months, and decide later.
  • You need round-the-clock incident response. That needs an on-call team in your customers' time zones or a managed operations provider.
  • You need the SLA contract written. That is a job for your lawyer; RAITHub can supply the measured numbers.

How RAITHub would build this

  • Scope: health endpoints per component; multi-region uptime checks and alerting; a request-based SLI and error budget dashboard; a hosted status page with customer-facing components; the measured numbers and a draft SLO for your adviser to turn into SLA terms.
  • Timeline: included in a new SaaS build (4–6 weeks for a fixed-scope MVP), or as a fixed-scope job on an existing product, confirmed in the written quote.
  • What you receive: code, dashboards and CI checks in your accounts and repository, an incident and status-page runbook, tests, IP assigned to you and an NDA as standard.
  • Next step: a free 15-minute technical audit, then a fixed written quote.

To find out what uptime you can safely promise, book the free 15-minute technical audit.

Documentation checked on 7 October 2026.

Frequently asked questions

What uptime should a SaaS promise?

For most early products, 99.5% to 99.9% monthly, and only after measuring that you already exceed it. Keep a stricter internal target so you have room to fix problems before a breach.

How much downtime is 99.9% uptime?

About 43 minutes in a 30-day month, or 8.76 hours in a year. At 99.99%, it is about 4.3 minutes a month and 52.56 minutes a year.

Does scheduled maintenance count as downtime?

Only if your SLA says so. Most SaaS SLAs exclude maintenance announced in advance, often with a monthly cap. Define the notice period and cap in the SLA itself.

What is an error budget?

The amount of unreliability your SLO allows in a period, such as 43 minutes a month at 99.9%. Teams use it to decide when to slow releases and focus on reliability, and when to ship freely.

Should the status page update automatically?

Alerts should be automatic, but the public status should be set by a person. A single failed check can be a false alarm, and an automated "major outage" banner can cause more support load than the incident itself.

What do service credits usually look like?

A percentage of the monthly fee, rising with how far uptime fell below the commitment, usually capped and claimed in writing within a set period. Confirm the terms and liability position with your adviser.

saas slauptimestatus pageSLOerror budgetsaas reliability

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.