Back to BlogCode Rescue & Fixes

From Vibe-Coded Prototype to Production: A 30-Day Hardening Plan

Rupak Amin

Founder & Lead Engineer, RAITHub

11 min read

To take a vibe-coded prototype to production, freeze and back it up first, then spend four weeks hardening it: security in week one, tests and CI on the money paths in week two, logging, backups and alerts in week three, and load, cost caps and a launch gate in week four. Most prototypes need this, not a rewrite.

This plan is for founders who built a working prototype with Lovable, Bolt, Cursor, Claude Code or a similar tool, have early users or a launch date, and need it to survive real customers. It assumes one experienced engineer, or a small team, working on an app of ordinary size: a few dozen screens, one database, one or two payment or email integrations. A larger app takes longer; the order stays the same.

If you only want to know which gaps your app has, run the 10-minute checks in how to make a vibe-coded app production-ready first. This post is the schedule for fixing them.

Why does a vibe-coded prototype need a hardening plan at all?

Because the tools optimise for code that runs, and production needs code that holds up when users are many, careless or hostile. Those are different tests, and nothing in a prototype runs the second one.

The evidence is consistent. Veracode's 2025 GenAI Code Security Report found that 45% of AI-generated code samples failed security tests and introduced OWASP Top 10 vulnerabilities, and that newer, larger models did not do better on security. In the 2025 Stack Overflow Developer Survey, the top frustration developers reported with AI tools, cited by 66%, was solutions that are almost right, but not quite; 45.2% said debugging AI-generated code takes more time than expected. "Almost right" is fine in a demo and expensive in production, which is why the plan below verifies every fix with a test.

What does the 30-day plan look like?

WhenFocusDone means
Days 1–2Freeze and ownYou own every account, a restore from backup has been tested, and a clean build works from the repository
Week 1 (days 3–9)SecurityNo secrets in client code, every table has row-level security or server-side ownership checks, every endpoint validates input
Week 2 (days 10–16)Tests and CIEnd-to-end tests on sign-up, login, the core workflow and payment, run on every pull request, blocking merge on failure
Week 3 (days 17–23)Operate itValidated environment config, schema as migrations, error tracking with alerts to a person, scheduled backups
Week 4 (days 24–30)Load, cost and launchRate limits on login and paid calls, spending caps, a load test at expected peak, and a signed-off launch gate

The order is deliberate: nothing is changed until it is backed up, nothing is "fixed" until a test proves it, and nothing launches until someone is alerted when it breaks.

What should you do in the first two days?

Stop changing things, take ownership, and prove you can rebuild the app from scratch. Everything later depends on this.

  • Freeze features. No new prompts to the AI tool that touch production for 30 days, except for fixes on this plan.
  • Own the accounts. Repository, hosting, database, domain, payments, email and AI provider, all in your company's name.
  • Back up and restore. Take a database backup and restore it somewhere else. A backup you have never restored is a hope, not a backup.
  • Rebuild from clean. Clone the repository to a fresh machine and run the exact install and build commands your host runs. On 27 September 2026 a security upgrade to the RAITHub website passed locally and failed on Vercel with ERESOLVE, a peer-dependency conflict; the write-up shows why the clean install is the one that counts.

This is the same first step RAITHub uses for any inherited codebase, set out in the Code Rescue Playbook.

Week 1: how do you close the security gaps?

Fix the three gaps that expose other people's data: secrets in client code, missing authorization, and missing input validation. Rotate every key that was ever exposed; deleting it from the code does not un-leak it.

For apps on Supabase, Supabase's row-level security guide is direct: enable RLS on every table in an exposed schema. Enabling it without policies blocks all access through the public API, so add a policy for each operation users need. A minimal owner-only pattern:

alter table public.projects enable row level security;

create policy "owners read their projects"
  on public.projects for select
  to authenticated
  using ((select auth.uid()) = owner_id);

create policy "owners update their projects"
  on public.projects for update
  to authenticated
  using ((select auth.uid()) = owner_id)
  with check ((select auth.uid()) = owner_id);

Then write the test that attacks it: sign in as user B and request user A's project by ID. It should fail before the policy exists and pass after. On Sundor Skin, a B2B wholesale platform RAITHub built, row-level security covers a 146-table PostgreSQL core, and an IDOR (insecure direct object reference) suite deliberately tries to read another buyer's data on every run.

For input validation, validate every request body on the server with a schema and accept only the fields the endpoint expects, so a request carrying "role": "admin" is rejected rather than saved. Then test what the schema actually outputs. RAITHub found that zod 4's .partial() still applies .default() values to omitted fields, which would have silently unpublished records on a partial update; the fix and regression tests are in zod 4 .partial() keeps default values.

Week 2: which tests should you write first?

End-to-end tests for the three to five journeys that make money or hold data, run in CI on every pull request, with a failed test blocking the merge. Unit tests for the rest can follow.

A minimal GitHub Actions workflow for a Node project with Playwright end-to-end tests:

name: ci
on:
  pull_request:
  push:
    branches: [main]

jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 22
          cache: npm
      - run: npm ci
      - run: npm run lint
      - run: npx tsc --noEmit
      - run: npm test
      - run: npx playwright install --with-deps chromium
      - run: npx playwright test

The Playwright config's webServer setting starts the app before the browser tests run. Point the tests at a disposable test database, never production. Then turn on branch protection so the test job must pass before anything merges. From this point on, every change the AI tool makes is checked by the same gate as a human's.

Which journeys to cover first:

  1. Sign-up, login and password reset.
  2. The core workflow a customer pays for, start to finish.
  3. Payment, including a declined card and a webhook arriving twice.
  4. A second user trying to read the first user's data.
  5. Any admin action that changes other users' records.

Week 3: how do you make the app operable?

Make the configuration impossible to get wrong, make the schema reproducible, and make sure a person hears about errors before customers do.

Validated config. Parse environment variables once at start-up and fail the build if a required one is missing. On the RAITHub website, the build fails when a required production variable is absent, so a half-configured deploy never goes live.

import { z } from 'zod'

const Env = z.object({
  DATABASE_URL: z.string().url(),
  STRIPE_SECRET_KEY: z.string().startsWith('sk_'),
  NEXT_PUBLIC_SUPABASE_URL: z.string().url(),
})

// Throws at start-up, with every missing or malformed variable listed.
export const env = Env.parse(process.env)

Schema as migrations. Prototypes built through a dashboard often have no migration files. Capture the current schema as a baseline migration, make every future change a migration, and replay them on an empty database in CI. If the replay fails, production and the repository have drifted.

Errors and alerts. Add error tracking on client and server, structured server logs, and one alert rule that pages a named person when the error rate jumps. Write a one-page runbook: how to deploy, roll back and restore.

Week 4: how do you prepare for launch traffic and bills?

Put a rate limit and a login in front of everything that costs money, set spending caps with every provider, load-test your expected peak, and run a written launch gate.

  • Rate limits on login, sign-up, password reset and every AI or paid API call. Mind the lock-out trap: when the RAITHub website limited login only per email address, anyone who knew the admin's email could keep the admin locked out. It now uses three buckets: email plus IP, per IP, and per email across IPs. The patterns are in the rate limiting explainer.
  • Spending caps and billing alerts on the AI provider, hosting, database and email.
  • A load test shaped like launch day: many sign-ups at once, the core workflow in parallel.
  • The launch gate. Work through the pre-launch QA checklist and sign it off in writing.

Is 30 days realistic for every vibe-coded app?

For a typical prototype with one experienced engineer, yes for the essentials above. Two situations take longer: a very large app, and a data model that contradicts the business.

What you findWhat it usually means
Security gaps, no tests, no migrationsNormal. The 30-day plan covers it
Hundreds of screens or several apps sharing one databaseSame plan, more weeks; fix in risk order, money paths first
The data model cannot represent how the business worksRebuild the core module; keep the rest. Rewrite vs refactor covers the decision
Nobody can say what the app should doStop hardening; write the spec first

Can you keep using AI coding tools after launch?

Yes, and you should, as long as CI checks their work. The difference after hardening is that a generated change has to pass the same tests, type-checks and migration replay as any other before it reaches customers.

Keep the prompts small and the changes reviewable, and add a test with every generated fix that touches money or data. RAITHub's own gates are described in how RAITHub tests software; this website alone runs 400+ tests in CI.

Why RAITHub for hardening a vibe-coded prototype

  • The fixes above are RAITHub's own incidents. The zod 4 .partial() bug, the npm ERESOLVE conflict and the rate-limit lock-out were found and fixed on this website, each with a regression test.
  • Isolation proven at scale. Sundor Skin's RLS and IDOR suite, 530+ tests in all, is the pattern week 1 installs.
  • Keep what works. Much of a generated app is fine. The first two days separate what to keep from what to harden.
  • Fixed scope. A free 15-minute technical audit, then a fixed written quote. You own the code and the IP, and an NDA is standard.

RAITHub has not published a case study of hardening an AI-generated app, so this post does not claim one; the evidence comes from platforms RAITHub built and from its own website.

When you don't need us

  • It is still a prototype with no real user data. Keep iterating; harden it when users arrive.
  • You have one specific problem, such as a failing deploy. That is a single fix: use the fix one issue page.
  • You have an engineer with time. The plan above is complete enough to follow in-house.
  • You need a certified vendor. RAITHub is not SOC 2 or ISO 27001 certified.

If several of the week-1 checks fail and launch is close, the Code Rescue service runs this plan for you. To start, book the free 15-minute technical audit and send the results of your checks.

Last reviewed: 29 September 2026. External sources checked on 29 September 2026.

Frequently asked questions

How long does it take to make a vibe-coded prototype production-ready?

For a typical prototype with one experienced engineer, about 30 days for the essentials: a two-day freeze, then a week each on security, tests and CI, operations, and load, cost and launch. Large apps or a broken data model take longer.

Do I need to rewrite my AI-generated app?

Usually not. Most gaps are fixed in place. A rewrite of one module makes sense when its data model cannot represent how the business works.

What should I fix first in a Lovable or Bolt app?

Secrets in client code, missing row-level security or ownership checks, and missing server-side input validation. These are the gaps that expose other people's data.

Which tests matter most for a vibe-coded app?

End-to-end tests for sign-up and login, the core paid workflow, payments, and a second user trying to read the first user's data, run in CI on every pull request.

Can I keep using Cursor or Claude Code after launch?

Yes. Once CI blocks merges on failed tests, type errors and migration drift, generated changes are checked by the same gate as human ones.

What does it cost to harden a prototype?

It depends on the app's size and how many gaps are open. RAITHub publishes no rates; it gives a fixed written quote after a free 15-minute technical audit.

vibe codingvibe-coded prototype to productionAI-generated codeproduction hardeningrow-level securityCIlaunch checklist

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.