Back to BlogQuality & Testing

Launch Checklist for an AI-Built SaaS: 40 Checks Before Real Users

Rupak Amin

Founder & Lead Engineer, RAITHub

14 min read

Before real users touch a SaaS you built with AI, run 40 pass-or-fail checks in eight groups: production environment, secrets, data access, login, payments, core journeys, devices and release safety. Do data access and secrets first. They are where AI-built apps fail most often in public research, and a leak there cannot be undone with a patch.

If you would rather have the checklist run for you, see how RAITHub would test this below.

Everyone can build with AI now; almost no one has a tester. Tools such as Lovable, Bolt, Cursor, Claude Code and Replit write working code fast, but they do not check it against a second user, a real phone, a declined card or someone poking at your API. This list is built for that gap. It sits on top of the general pre-launch QA checklist, which covers SEO, rollback and accessibility in more depth; here every check is chosen because AI-built apps tend to miss it.

Why does an AI-built SaaS need its own launch checklist?

Because the failures are predictable, and they cluster in places the demo never shows. Three findings from 2025 research set the priorities:

  • Escape scanned 5,600 publicly available apps built on vibe-coding platforms and found more than 2,000 high-impact vulnerabilities, more than 400 exposed secrets and 175 cases of exposed personal data, most reachable without logging in (Escape research, October 2025).
  • Wiz grouped the risks it found in vibe-coded apps into four kinds: login logic in the browser, API keys in client code, missing or loose row-level security, and internal apps left public (Wiz Research, September 2025).
  • Veracode found that 45% of AI code-generation tasks introduced an OWASP Top 10 weakness (Veracode 2025 GenAI Code Security Report).

None of these show up when one person clicks through the happy path on a laptop, which is how most AI-built apps are checked before launch.

What does the whole checklist look like on one page?

GroupChecksWhat failure looks likeRun first on a small budget?
A. Production environment1–5Works in preview, breaks or points at test data in productionYes
B. Secrets and keys6–10A secret key readable in the browser bundleYes
C. Data access11–15One user reads another user's recordsYes, before anything else
D. Login and roles16–20Admin pages reachable by a normal userYes
E. Payments21–25Paid but no access, or access without payingYes, if you charge at launch
F. Core journeys and states26–30New users see a blank screen or an errorPartly
G. Phones, accessibility, speed31–35Sign-up button off-screen on an iPhonePartly
H. Release safety36–40The next prompt silently breaks checkoutAfter launch is fine for some

A. Is production really set up, or only the preview?

AI builders make the preview feel finished. Production has different keys, URLs and data, and that is where the first bugs appear. The detail is in why an AI app works locally and breaks in production.

  • 1. Every environment variable the app reads is set in production, and none of them still holds a test or placeholder value.
  • 2. Production uses its own database, separate from the one you and the AI agent experiment on. In July 2025 an AI agent on Replit deleted a live production database during a test project, and Replit's response included automatic separation of development and production databases (Fortune).
  • 3. Auth redirect URLs, email links and OAuth callbacks point at the production domain, not localhost or a preview URL.
  • 4. Backups are on, and you have restored one into a scratch database at least once.
  • 5. A deliberate test error reaches your error tracker and alerts a real person.

B. Are any secrets visible in the browser?

Anything shipped to the browser is public. AI tools sometimes move a server key into client code because it makes an error disappear.

  • 6. Build the app and search the output for secret key prefixes (the command is below).
  • 7. Only publishable keys appear in client code: the Stripe publishable key, the Supabase anon or publishable key. Never a Stripe secret key or a Supabase service role or secret key.
  • 8. Rotate every key that was ever pasted into a chat, a prompt or a screenshot during the build.
  • 9. Your AI provider key, if the app calls a model, has a spending limit and lives only on the server.
  • 10. No .env file is in the repository or its history.
npm run build
grep -rEl "sk_live_|rk_live_|sb_secret_" dist .next/static build 2>/dev/null

Any file listed is a leak. Exposed API keys in a Lovable app covers how to rotate and move them server-side.

C. Can one user see another user's data?

This is the check that matters most. On Supabase-backed apps, the browser talks to the database directly, so row-level security (RLS, database rules that limit each user to their own rows) is the only thing between a user and everyone else's records.

  • 11. RLS is enabled on every table in the exposed schema. Run the query below; it should return no rows.
  • 12. Create two accounts. As user A, copy a record ID; as user B, request it. It must be refused.
  • 13. Call the API with only the public anon key and no session. Nothing private comes back.
  • 14. File storage buckets holding user uploads are private, with access rules per user.
  • 15. Roles and plan limits are not read from data the user can edit. Supabase's docs warn that user metadata can be updated by the signed-in user and is not a safe place for authorisation data (Supabase RLS guide).
-- Tables in the public schema with row-level security switched off
select c.relname as table_name
from pg_class c
join pg_namespace n on n.oid = c.relnamespace
where n.nspname = 'public'
  and c.relkind = 'r'
  and not c.relrowsecurity;

If this returns anything, start with fixing RLS disabled in public. The wider list is in the vibe-coded app security checklist.

D. Does login hold up when someone tries the wrong door?

  • 16. Sign-up, email verification, login, logout and password reset all work on the production domain, and a reset link works once only.
  • 17. Every admin page and admin API refuses a normal user and an anonymous request, checked on the server. Hiding a button is not access control.
  • 18. Authorisation does not live only in middleware. CVE-2025-29927 let attackers skip Next.js middleware with one request header on unpatched versions (Next.js advisory), so check permissions again where the data is read.
  • 19. Login, sign-up, password reset and any endpoint that calls an AI model are rate limited.
  • 20. If sign-up should be closed or invite-only, try to register anyway through the API, not only the form.

The full role and data-access test plan is in testing login, roles and data access in an AI-built app.

E. Will payments behave when something goes wrong?

  • 21. The webhook handler verifies the provider's signature. Stripe requires the raw request body for this (Stripe webhooks docs).
  • 22. Sending the same webhook twice grants access or ships an order once.
  • 23. Declined cards, insufficient funds and 3-D Secure cards from Stripe's test cards all end in a clear state.
  • 24. Live-mode products, prices and webhook endpoints exist. Stripe's go-live checklist notes that sandbox objects are not usable in live mode.
  • 25. One real low-value purchase in production, then a refund, both reflected in the app.

The money side has its own post: testing payments in an AI-built app.

F. Do the core journeys survive real users?

  • 26. Write the three to seven journeys that matter as sentences, such as "a new user signs up, creates a project and invites a teammate". AI tools build what you described; this list is how you test what you meant.
  • 27. Each journey works for a brand-new account with no data. Empty states are where generated screens most often show a spinner forever.
  • 28. Double-clicking every submit button creates one record, not two.
  • 29. Refreshing or pressing Back mid-flow does not lose work or create half-finished records.
  • 30. Errors show a plain message, never a stack trace, raw SQL or a model's error text.

G. Does it work on phones, for keyboard users, and at real speed?

  • 31. Every journey works on one real iPhone and one real Android phone. Mobile was 58.99% of worldwide web traffic in September 2026 (StatCounter). The common phone bugs are in why an AI-built app breaks on phones.
  • 32. Each journey can be completed with a keyboard alone, with visible focus.
  • 33. Form fields have labels and text has enough contrast. In WebAIM's 2025 scan of one million home pages, 79.1% had low-contrast text and 48.2% had missing form labels (WebAIM Million 2025).
  • 34. Key pages meet Google's "good" Core Web Vitals on a throttled mobile profile (web.dev).
  • 35. The busiest list page still loads with realistic data, for example 5,000 rows, not the ten the AI seeded.

H. Will the next prompt break what you just tested?

  • 36. The tests the AI wrote fail when you break the code on purpose. Flip a permission check and see whether anything goes red.
  • 37. Every change, prompted or typed, passes typecheck, lint and tests in CI before it reaches production.
  • 38. Every dependency exists, is maintained and was chosen on purpose. Researchers found code models recommend packages that do not exist, 5.2% of the time for commercial models (USENIX Security 2025).
  • 39. You can roll back to the previous deployment in minutes, and you have done it once.
  • 40. Sign-up and receipt emails arrive in the inbox. Gmail requires every sender to set up SPF or DKIM, and bulk senders all three of SPF, DKIM and DMARC (Google sender guidelines).

Checks 36 and 37 are the ones that keep the other 38 true. How to QA an app built with AI coding tools shows a minimal CI gate.

What should you test first on a small budget?

If you have one day, run groups C, B and D in that order, then checks 21, 22 and 31. That covers data exposure, leaked keys, admin access, payment integrity and phones, which are the failures users notice first and the ones that cost the most to explain afterwards. Leave performance tuning and full accessibility work for the first weeks after launch, but do not skip them.

How long does this take to run yourself?

For a small SaaS with two roles and one payment flow, plan on 2–3 days for a developer who knows the stack and Supabase or your database: half a day for groups A and B, a day for C and D, half a day for payments, and the rest for journeys, phones and the CI gate. Fixes come on top. The main risk of doing it yourself is that the person, or the AI tool, that built the app also judges it, and both tend to test what they expect rather than what a stranger would try.

Buy, build or hire?

OptionChoose this whenTrade-off
Ask the AI tool to review its own work, plus the platform's security scanBefore every launch, as a first passFree and fast; it checks against the same assumptions it built with
A tool or SaaS testing platform (scanners, Playwright, a device cloud)A developer can own the checklist and automate itGood for repeatable checks; someone still has to decide what to test
Freelancers or crowdtestingYou want many devices and fresh eyes for a launch weekWide coverage; depth and report quality vary
An in-house QA hireYou release every week and have product-market fitDeep product knowledge; slow and costly to hire for an early app
A managed QAaaS team or a one-off launch auditYou want the 40 checks run and reported before launch, without hiringYou rely on an outside team; make sure tests and reports are yours

For how these compare in cost, see QA as a service pricing.

Why RAITHub for this

  • Built for AI-built apps. RAITHub tests apps made with Lovable, Bolt, Cursor, Claude Code, Replit and v0, and knows where each tends to leave gaps: database rules, keys, preview-only settings and missing failure paths.
  • Testing habits from shipped products. RAITHub's own builds carry large automated suites: 1,024 tests on PropDesk, 750+ on TheSkinProof (the founder's own venture) and 530+ on Sundor Skin, which runs 146 PostgreSQL tables under row-level security. There is no AI-built-app case study yet.
  • Human testing plus automation. Exploratory testing on real phones finds what scripts miss; the repeatable checks become tests in your repository.
  • Fixes if you want them. Your team can fix from the report, or RAITHub can.

When you don't need us

  • It is a landing page or waitlist with no accounts, data or payments. Checks 1, 6 and 31 are enough.
  • You have a developer with two or three free days who did not build the app. Give them this list.
  • You need a certified penetration test or a compliance attestation. RAITHub's security checks follow OWASP guidance but are not a CREST- or PCI-certified pentest; see pentest vs vulnerability scanning.

How RAITHub would test this

  • Scope: a free 15-minute call to agree your journeys, roles, payment provider and the phones your users carry.
  • Launch audit: all 40 checks run by a tester on your production-like environment and real devices, plus exploratory sessions on sign-up, data access and checkout.
  • Report: every bug ranked by severity, with steps to reproduce, evidence and a suggested fix, and a clear go or no-go for launch.
  • Keep it true: optional monthly QA that re-runs the critical checks after each release and grows the automated suite in your repository.
  • What you receive: the written report, any tests in your repo, and a handover note. IP is yours and an NDA is standard.

The audit is fixed-price, set in a written quote after the free call. See AI-built app testing and the QA as a service overview, or ask for a launch audit of your AI-built app.

Frequently asked questions

What should I check before launching an app built with AI?

Start with data access and secrets: row-level security on every table, a two-account test, and no secret keys in the browser bundle. Then login and admin access, payments, the core journeys on real phones, and a CI gate so the next prompt cannot undo the fixes.

Is the security scan in my AI builder enough before launch?

It is a useful first pass, not a launch gate. Built-in checks such as Supabase's database advisors catch known patterns, like tables without RLS, but they do not log in as two users, pay with a declined card or use the app on an iPhone. Run this checklist as well.

How long does a launch checklist take for an AI-built SaaS?

About 2–3 days for a developer who knows the stack and did not build the app, for a small SaaS with two roles and one payment flow. Fixing what it finds takes extra time, so run it a week before launch, not the night before.

What is the most common serious bug in AI-built apps?

In 2025 research by Escape and Wiz, the recurring serious problems were data reachable without proper access rules, often missing or loose row-level security, and API keys or login logic exposed in browser code.

Can I launch first and test later?

For cosmetic issues and performance tuning, yes. For data access, secrets, admin access and payments, no: a leak or a wrong charge reaches real people and cannot be taken back with a later release.

Does RAITHub fix the bugs it finds?

It can. The launch audit produces a ranked report with suggested fixes that your team or your AI tool can apply; RAITHub can also fix them under a separate fixed quote.

ai app launch checklistvibe coded app launchlovable app launch checklistai built saas testingpre-launch qaqa for ai-built apps

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.