Back to BlogAI Features & LLMs

How to QA an App Built with AI Coding Tools

Rupak Amin

Founder & Lead Engineer, RAITHub

10 min read

To QA an app built with AI coding tools, test behaviour rather than trusting the code: write down what each role should be able to do, test access control and money flows first, verify every dependency exists and is the one you meant, check that AI-written tests fail when the code is broken, and gate every change in CI. Treat the app as code from an unknown contributor.

If you would rather have an AI-built app tested for you, see how RAITHub would test this below.

Why does AI-generated code need different QA?

Because the tools optimise for code that runs and looks right, and the bugs that matter do not stop code running. Three numbers frame the problem:

  • Veracode tested code from more than 100 language models on 80 tasks and found that 45% of cases introduced an OWASP Top 10 vulnerability, with no improvement in security as models got larger (Veracode 2025 GenAI Code Security Report).
  • In the 2025 Stack Overflow survey, 66% of developers named "AI solutions that are almost right, but not quite" as a frustration, and 46% said they distrust the accuracy of AI tools (Stack Overflow Developer Survey 2025).
  • Researchers at USENIX Security 2025 found that code-generating models recommend packages that do not exist, at an average of 5.2% for commercial models and 21.7% for open-source ones (Spracklen et al.). An attacker can register those names.

"Almost right" is the QA problem in two words. The happy path works in the demo; the edge cases, the second user and the failure modes were never specified, so the model guessed. The public example is CVE-2025-48757, where generated apps without row-level security exposed data to anyone holding the public key.

A note on scope: this post is about testing code that AI wrote. Testing an AI feature inside your product, such as a chatbot, is a different job, covered in how to test LLM features.

What does AI-generated code usually get wrong?

FailureWhy AI tools produce itTest that catches it
Missing authorisation on the serverThe UI hides a button, so the generated code looks finishedCall the API directly as the wrong user or role
Data visible across users or tenantsQueries filter by ID but not by owner; row-level security offTwo-account IDOR tests on every data endpoint
Secrets in the browser bundleA server key used in client code because it made the error go awaySearch the build output for key prefixes
Validation only in the formClient-side checks generated with the form; the API trusts the bodySend invalid and extra fields straight to the API
Hallucinated or unexpected packagesThe model names a plausible package that does not exist, or an abandoned oneReview every new dependency; audit in CI
Error paths that fail openThe prompt described success; failures were never specifiedMock timeouts and errors from each external service
Tests that assert nothing usefulThe model wrote tests that pass against whatever the code doesMutation testing, or break the code by hand and see what fails
Inconsistent patterns across filesEach prompt solved its problem locallyLint, typecheck and a code review of the core flows

What is the QA plan, step by step?

1. Write down what the app should do

List the roles, and for each one what it can see, create, change and delete. Add the money and data flows: sign-up, payment, refund, export, delete account. This one page is the test oracle. Without it, you can only test whether the app does what the code does, which always passes.

2. Test access control first

Create two accounts per role. For every endpoint that returns or changes data, make the request as the wrong user. On Supabase-backed apps, also check that row-level security is on for every exposed table; the vibe-coded app security checklist has the queries. These are the bugs that expose other people's data, so they come before anything cosmetic.

3. Verify the dependencies

Open the package manifest and confirm that every package is one someone chose on purpose, exists on the registry with real history, and is maintained. Then run an audit in a clean clone. A model that hallucinates a package name will often hallucinate it again for the next user, which is what makes the attack practical.

4. Exercise the failure paths

Turn off the network to each external service in turn: payments, email, storage, the AI API. Send malformed requests. Double-submit forms. Check that nothing is marked paid, sent or deleted that should not be, and that errors show a plain message, not a stack trace.

5. Check the tests the AI wrote

A green test suite proves little if the tests would pass against broken code. Mutation testing makes this measurable: a tool such as Stryker introduces small changes to your code and reruns the tests; a change that no test catches marks a weak spot. If you cannot run mutation testing, do it by hand on the critical paths: flip a comparison in the permission check and see whether anything goes red.

6. Put a gate in front of every prompt

Every AI-generated change should pass the same CI checks as a human one, before it merges. A minimal GitHub Actions gate for a Node project:

name: ci
on: [pull_request]
jobs:
  check:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 22
      - run: npm ci
      - run: npm run typecheck
      - run: npm run lint
      - run: npm test
      - run: npm audit --omit=dev --audit-level=high

Add your access-control tests to npm test so a prompt that removes an ownership check fails the build. Regression testing after every deploy covers the end-to-end layer.

What does a minimal access-control test look like?

One test per data-holding endpoint, written once, run forever:

import { test, expect } from '@playwright/test'

test('a member cannot change another account plan', async ({ request }) => {
  const res = await request.patch('/api/accounts/' + process.env.OTHER_ACCOUNT_ID, {
    headers: { Authorization: 'Bearer ' + process.env.MEMBER_TOKEN },
    data: { plan: 'enterprise' },
  })
  expect([403, 404]).toContain(res.status())
})

RAITHub's own platforms use this pattern at scale. Sundor Skin, a B2B wholesale platform RAITHub built, runs a 21-case IDOR suite in CI that tries to read other buyers' data (case study).

How long does it take to QA an AI-built app yourself?

For a small app with two or three roles and a payment flow, budget 3–5 days for a developer who knows the stack: a day to write the role map, one to two days for access-control and failure-path tests, a day for dependencies and the CI gate, and time to fix what turns up. The main risk of doing it yourself is that the same person, or the same AI tool, that built the app also judges it. Asking the model that wrote the code whether the code is secure tends to get a confident yes. A second reviewer, human or company, finds what the first one assumed.

When is the code beyond testing, and needs a rescue?

When tests keep finding the same class of bug in new places, when no one can explain how a core flow works, or when the data model itself is wrong, for example tenant data with no owner column. At that point testing documents the problem rather than fixing it. The code rescue playbook covers how to stabilise an inherited codebase, and making a vibe-coded app production-ready covers the wider launch checklist.

Buy, build or hire?

RouteChoose this when
A tool or SaaS testing platform (scanners, AI code review bots, the builder's own security scan)You want a fast first pass on known patterns. These tools do not know your roles, so they miss most access-control bugs.
Freelancers or crowdtestingYou need exploratory testing of the UI across devices and can write the role map yourself.
An in-house QA hireYou ship AI-generated changes daily and have steady testing work for a full-time person.
A managed QAaaS teamYou want the plan above done, with the tests left in your repository and a CI gate in place, without hiring.

Why RAITHub for this

  • Test-heavy by default. RAITHub's builds ship with large suites: 530+ tests on Sundor Skin and 1,024 on PropDesk, and 400+ on this website.
  • Access control is the first thing tested, because that is where AI-built apps leak data.
  • Testing that can turn into fixing. If the code needs a rescue rather than a test pass, RAITHub will say so and can scope that work separately.

When you don't need RAITHub

  • The app is a prototype with test data and no real users yet. Use the checklist and a scanner.
  • You need a certified penetration test or a compliance attestation. RAITHub's security testing is application-level testing against OWASP guidance, not a CREST- or PCI-certified pentest.
  • You want testers placed under your management. RAITHub does not offer staff augmentation.

How RAITHub would test this

  • Role and flow map: written from the product, not the code, and agreed with you.
  • Security pass: access control, tenant isolation, secrets, validation and authentication against the OWASP Top 10:2025; see the OWASP testing checklist.
  • Exploratory and failure-path testing on the journeys that hold data or money.
  • Automation: access-control and regression tests in your CI, plus a check on the strength of the tests that already exist.

Buy it as a fixed-price pre-launch audit, a monthly QA plan for teams shipping AI-generated changes every week, or a dedicated QA team RAITHub manages and bills monthly. You receive the findings with fixes, the tests in your repository and full IP under NDA. See QA as a service, security testing and QA test automation. Next step: book the free 15-minute technical audit, then get a written fixed quote.

Last reviewed: 7 October 2026.

Frequently asked questions

Is AI-generated code less secure than human code?

It is not reliably more secure. Veracode's 2025 study found an OWASP Top 10 flaw in 45% of AI-generated test cases. The bigger issue is volume: AI tools produce a lot of code quickly, and review rarely keeps up unless tests do the checking.

Can I ask the AI to test its own code?

You can ask it to write tests, and they help, but check them. Models often write tests that confirm what the code already does. Break the code on purpose, or run mutation testing, to see whether the tests notice.

What should I test first in a vibe-coded app?

Access control: whether one user can read or change another user's data through the API. Then secrets in the front end, server-side validation and payment flows. Those are the failures that hurt real users.

Do an AI app builder's built-in security scans replace testing?

No. Scans such as Lovable's pre-publish check (Lovable security docs) catch common configuration issues and are worth running. They do not know your roles and business rules, so cross-user access and logic bugs still need tests written for your app.

How do I stop AI tools adding fake or risky packages?

Review every new dependency in pull requests, confirm it exists with real history on the registry, commit the lockfile, and run a dependency audit in CI. Treat a package you have never heard of as a question, not a fix.

How is testing AI-generated code different from testing an AI feature?

Testing AI-generated code checks ordinary software that an AI happened to write. Testing an AI feature checks a model's outputs inside your product, which needs evaluation sets and guardrails rather than only pass/fail tests.

testing AI-generated codeAI code QAvibe codingCursorLovablemutation testingQA as a service

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.