Founder & Lead Engineer, RAITHub
Feature flags in a SaaS do four jobs: release, experiment, plan entitlement and kill switch. Target them by tenant first, so a whole customer sees the same product, and roll out by percentage with a stable hash so the same 10% stays in as you widen it. A small in-house table covers most early products; buy a hosted tool once you need experiments and audit trails.
If you would rather have flags built into your product properly, see how RAITHub would build this below.
What is a feature flag, and what are the four types a SaaS needs?
A feature flag is a runtime switch that decides whether a piece of code runs for a given request, without a new deploy. In a SaaS the same mechanism ends up doing four very different jobs, and mixing them up is where most flag mess comes from.
Pete Hodgson's long-standing feature toggles article on martinfowler.com sorts toggles by how long they live and how often they change: release toggles live for weeks, experiment toggles for hours or weeks, ops toggles from short to long, and permissioning toggles for years. Mapped onto a B2B SaaS:
| Flag type | Job | Typical lifetime | Who flips it | Default when the flag service is down |
|---|---|---|---|---|
| Release | Ship unfinished or risky code dark, then open it gradually | Days to weeks; delete after full rollout | Engineering | Off (old behaviour) |
| Experiment | Split users between variants and measure a metric | One test cycle, usually weeks | Product or growth | Control variant |
| Plan entitlement (permission) | Turn features on per plan, add-on or tenant contract | Years; it is part of the product model | Billing state, not a person | Whatever the stored entitlement says |
| Kill switch (ops) | Turn off a costly or failing feature in seconds | Permanent for risky paths | On-call engineer | On (feature keeps working) |
The defaults column matters more than it looks. A release flag should fail closed, so an outage of your flag store never exposes half-built work. A kill switch should fail open, so the same outage does not switch off a working feature for every customer.
Should plan features be feature flags or entitlements?
Plan features should be entitlements: rows in your own database, written by your billing webhooks, read by the same evaluator as your flags. They should not be flags someone toggles by hand in a dashboard.
The difference is the source of truth. A release flag is a decision an engineer makes. An entitlement is a fact about what a customer has paid for, and it must change when Stripe says the subscription changed, not when someone remembers. Our guide to building multi-tenant B2B SaaS explains why entitlements live in your database and are updated from idempotent webhooks; how plans are priced is in SaaS billing models explained.
Entitlements are also not the same as roles. "This tenant is on the Pro plan" and "this user is an admin inside that tenant" are separate checks, and a feature usually needs both. Role design is covered in SaaS authorization and RBAC design.
How do you target a feature flag by tenant, user and plan?
Evaluate in a fixed order: a global switch, explicit tenant denies, explicit tenant allows, plan rules, then the percentage rollout. The first rule that decides wins.
- Tenant first. In B2B, if half the users at one customer see a new screen and half do not, support tickets follow. Roll out by tenant ID by default, and by user ID only for consumer-style features or experiments.
- Allow lists for design partners. Named tenants get the feature early, whatever the percentage.
- Deny lists for risk. A large customer in a change freeze can be held back explicitly.
- Plan rules from entitlements. The evaluator reads the tenant's plan from your database, never from a value the browser sends.
- Evaluate on the server. A flag hidden in the browser is a suggestion, not a control. If a feature is gated, the API must check the same flag.
How do percentage rollouts stay stable for each user?
Hash the flag key together with the tenant or user ID into a bucket from 0 to 99.99, and show the feature when the bucket is below the rollout percentage. The same input always gives the same bucket, so nobody flickers between versions on refresh.
This is how the open-source tools do it. The Unleash stickiness documentation describes hashing a context field such as userId together with the strategy's groupId, which defaults to the flag name, using MurmurHash to get a number between 0 and 100; anyone at or below the rollout percentage sees the feature.
Two properties fall out of that design, and both are worth keeping in your own code:
- Widening is monotonic. Moving from 10% to 25% keeps the original 10% in, because their buckets have not changed. Never use random numbers per request.
- Salting by flag key. Because the flag key is part of the hash, the same unlucky tenants are not first in line for every risky release.
What does a minimal TypeScript feature flag evaluator look like?
About 50 lines. This one uses the FNV-1a 32-bit hash because it is short, deterministic and needs no dependency; it is not a cryptographic hash and does not need to be.
type Plan = 'free' | 'pro' | 'enterprise'
type FlagKind = 'release' | 'experiment' | 'entitlement' | 'kill_switch'
export interface FlagContext {
tenantId: string
userId: string
plan: Plan // read from your entitlements table, never from the client
}
export interface FlagRule {
key: string
kind: FlagKind
enabled: boolean // master switch: false means off for everyone
denyTenants?: string[]
allowTenants?: string[]
plans?: Plan[]
rolloutPercent?: number // 0 to 100; undefined means 100
rolloutBy?: 'tenant' | 'user'
}
// FNV-1a, 32-bit: same input, same output, on every server and runtime.
function fnv1a(input: string): number {
let hash = 0x811c9dc5
for (let i = 0; i < input.length; i++) {
hash ^= input.charCodeAt(i)
hash = Math.imul(hash, 0x01000193)
}
return hash >>> 0
}
// Bucket 0..9999, so rollouts can go in steps of 0.01%.
export function bucket(flagKey: string, unitId: string): number {
return fnv1a(flagKey + ':' + unitId) % 10000
}
const SAFE_DEFAULT: Record<FlagKind, boolean> = {
release: false,
experiment: false,
entitlement: false,
kill_switch: true, // a missing kill switch must not switch the feature off
}
export function isEnabled(
flag: FlagRule | undefined,
ctx: FlagContext,
kind: FlagKind,
): boolean {
if (!flag) return SAFE_DEFAULT[kind]
if (!flag.enabled) return false
if (flag.denyTenants?.includes(ctx.tenantId)) return false
if (flag.allowTenants?.includes(ctx.tenantId)) return true
if (flag.plans && !flag.plans.includes(ctx.plan)) return false
if (flag.rolloutPercent === undefined) return true
const unit = flag.rolloutBy === 'user' ? ctx.userId : ctx.tenantId
return bucket(flag.key, unit) < flag.rolloutPercent * 100
}
Call it as isEnabled(flags.get('new-invoice-editor'), ctx, 'release'). Passing the kind at the call site means a typo in a flag key gets the safe default for that kind rather than an accidental "on".
Store the rules in one table, cache them in memory for a few seconds per server instance, and write every change to an audit log with who changed what. A minimal Postgres table:
CREATE TABLE feature_flags (
key text PRIMARY KEY,
kind text NOT NULL
CHECK (kind IN ('release', 'experiment', 'entitlement', 'kill_switch')),
enabled boolean NOT NULL DEFAULT false,
rules jsonb NOT NULL DEFAULT '{}',
owner text NOT NULL,
expires_at date, -- required for release and experiment flags
updated_at timestamptz NOT NULL DEFAULT now()
);
The owner and expires_at columns are the lowest-cost flag-debt control you will ever add. If you already keep an audit log, flag changes belong in it; the pattern is in SaaS audit log design.
Do-it-yourself estimate: 1–3 days for the table, evaluator, cache, admin screen and tests, if you already know your stack. The main risk is not the code; it is the flags nobody deletes, and a cache that serves a stale kill switch for minutes during an incident.
Should you buy LaunchDarkly, Unleash, Flagsmith or PostHog, or build your own?
Buy when you need experiments with statistics, approval workflows, or many teams flipping flags. Build a small table when you have a handful of release flags and plan entitlements that already live in your database. Prices below are from each vendor's pricing page on 2 October 2026; check them again before you decide.
| Tool | Free tier | Paid entry point | Self-host option |
|---|---|---|---|
| LaunchDarkly | Developer plan: unlimited seats, 5 service connections, 1,000 client-side MAU per month | Foundation: $10 per service connection per month and $8.33 per 1,000 client-side MAU (billed yearly) | No; Enterprise is custom-priced (LaunchDarkly pricing) |
| Unleash | Open Source edition, self-hosted, no seat limit | Pay-as-you-go cloud at $75 per seat per month | Yes, open source (Unleash pricing) |
| Flagsmith | Up to 50,000 requests per month, 1 team member | Start-Up: $40 per month billed annually ($45 monthly), 1,000,000 requests, 3 members | Yes, on Enterprise, on-prem or in your own cloud (Flagsmith pricing) |
| PostHog | 1 million feature flag requests per month | Usage-based beyond the free allowance; see the page for current rates | Hosted product (PostHog pricing) |
Note the units differ: seats, service connections, monthly active users and requests. Model your own numbers before comparing headline prices; a server-side B2B product with few seats can land very differently on each.
Buy, build or hire?
| Option | Examples | Choose this when | Watch out for |
|---|---|---|---|
| Off-the-shelf hosted service | LaunchDarkly, Flagsmith cloud, PostHog | Several teams ship daily, you run experiments, or you need approvals and audit history now | Pricing units that grow with traffic or seats; another vendor in your security questionnaire |
| Open-source, self-hosted (the template route) | Unleash Open Source, Flagsmith self-hosted | You want a ready admin UI and SDKs but your data must stay in your cloud | You now run, patch and back up another service |
| Small in-house table | The evaluator and table above | Fewer than a few dozen flags, entitlements already in your database, one team | No experiment statistics; you must build the audit trail and the cleanup habit yourself |
| Hire a team to build it in | RAITHub or another studio | Flags, entitlements and billing must agree, and nobody on your team owns that layer | Make sure the handover includes the tests and the cleanup rules, not only the code |
How do you stop feature flags turning into technical debt?
Give every release and experiment flag an owner and an expiry date when it is created, and make the build fail when a flag outlives its date. Hodgson's article calls toggles "inventory which comes with a carrying cost" and suggests expiry dates, removal tasks in the backlog and "time bomb" tests.
- Create the removal ticket with the flag. The pull request that adds a release flag links the ticket that deletes it.
- Fail CI on expired flags. A test reads the flag table or a flag registry and fails when a release flag passes its expires_at.
- Delete in two steps. First set the flag to 100% and leave it for a release; then remove the branch in code; then drop the row.
- Cap the count. If you have more live release flags than engineers, stop adding and start removing.
- Never reuse a key. Old SDK caches or stale rows can switch an old meaning back on.
How do you test code behind a feature flag?
Test the configuration you will run in production and the configuration with the flag off. You do not need every combination of every flag.
Hodgson recommends exactly that: test the expected production configuration and the fallback with toggles off, and optionally everything on, with the convention that "off" means the old behaviour. In practice:
- Unit tests for the evaluator: allow beats rollout, deny beats allow, plan rules hold, the same ID always gets the same bucket, and a 20% rollout lands near 20% over 10,000 sample IDs.
- Both paths of each live release flag in your integration or end-to-end suite, with the flag injected per test rather than read from a shared environment.
- A server-side check test: call the gated API as a tenant without the flag and expect a refusal, not just a hidden button.
- A kill-switch drill: flip it in staging and time how long until every instance obeys. That number is your cache TTL plus propagation.
Flags reduce the blast radius of a release; they do not replace regression tests. If releases already break things, start with why every release breaks something and how to set up QA in an early-stage SaaS.
Why RAITHub for this
Flags in a SaaS sit where tenancy, permissions and billing meet, and that is the layer RAITHub builds and tests on every project. RAITHub has not published a standalone feature-flag system; the evidence is in the surrounding layers.
- Permission models at scale. Sundor Skin has 12 staff roles built from 88 permission codes and 530+ automated tests; PropDesk serves 4 roles from one codebase with 1,024 tests.
- Tenant isolation. BlockEstate is a multi-tenant listing platform that reached MVP in 6 weeks, so tenant-scoped targeting is familiar ground.
- Billing tied to access. PropDesk collects rent through Stripe, and PadhAI runs 9 payment gateways, so entitlements that follow payment state are routine work.
- QA first. The evaluator, both flag paths and the server-side checks ship with tests and CI, not as a later task.
When you don't need us
- You have three flags and one developer. The code above, or a free tier from the table, is enough.
- You need experiment statistics now. A hosted tool with built-in analysis will get you there faster than a custom build.
- You want developers placed in your team. RAITHub does fixed-scope builds and dedicated teams, not staff augmentation.
How RAITHub would build this
As part of a SaaS build or as a focused addition to an existing product, scoped in writing first.
- Flag and entitlement model: one table for release, experiment and kill-switch flags, entitlements written from billing webhooks, and safe defaults per type.
- Server-side evaluator: tenant, user and plan targeting with deterministic hashing, a short in-memory cache and a measured kill-switch propagation time.
- Admin screen with audit: who changed which flag, for which tenants, and when.
- Debt controls: owners, expiry dates and a CI check that fails on expired flags.
- Or integration instead: if a hosted tool fits better, wiring LaunchDarkly, Unleash, Flagsmith or PostHog into the same server-side checks.
Timeline: inside a new product, this is part of the 4–6 week fixed-scope SaaS and MVP build; added to an existing backend, it fits within the 6–12 week backend and API range alongside other work, with the exact scope in the quote.
You receive: automated tests and CI for the evaluator and both paths of each flag, handover docs and runbooks (including the kill-switch drill), and full IP assigned to you under NDA.
Next step: a free 15-minute technical audit, then a written fixed quote. See the SaaS development service and the SaaS industry page, or book the audit.
Frequently asked questions
What are feature flags used for in SaaS?
Four jobs: releasing code gradually, running experiments, gating features by plan or contract, and switching off a failing or costly feature quickly. Each has a different lifetime, owner and safe default, so keep them labelled by type.
Should feature flags be evaluated on the client or the server?
On the server for anything that gates data or paid features. Client-side flags are fine for cosmetic changes, but a user can change what the browser sees, so the API must check the same flag.
How do percentage rollouts keep the same users in?
By hashing the flag key with a stable ID, such as the tenant or user ID, into a fixed bucket. A user below the rollout percentage stays in as the percentage grows, because the bucket never changes.
Is LaunchDarkly worth it for a small SaaS?
It has a free Developer plan, per its pricing page, and paid usage is billed by service connections and client-side monthly active users. Whether it is worth it depends on whether you need experiments, approvals and audit history now; a small table works for a handful of flags.
How many feature flags is too many?
There is no fixed number, but release flags without an owner and expiry date are too many at any count. Many teams cap live release flags and fail the build when one passes its expiry date.
Can feature flags replace plan entitlements?
They can share an evaluator, but entitlements should be written by your billing system, not toggled by hand. That way a downgrade or failed payment changes access automatically and consistently.
Related posts
Designing a Public API for Your SaaS: Keys, Versioning and Limits
12 min readSaaS Entitlements: Enforcing Plans, Limits and Add-ons in Code
14 min readSending Webhooks to Your SaaS Customers: Retries, Signing and Logs
12 min readReady to discuss your project?
Book a free 15-minute technical audit with our engineering team.