Back to BlogAI Features & LLMs

Your AI Wrote the Tests Too: Why They Pass and Still Miss Bugs

Rupak Amin

Founder & Lead Engineer, RAITHub

13 min read

AI-written tests pass and still miss bugs because the model reads the code and writes assertions that describe what it does, bugs included. They also lean on mocks, chase coverage rather than failures, and agents sometimes edit a failing test instead of the code. Check them by breaking the code on purpose: a test suite that stays green against broken code is not protecting you.

If you would rather have someone independent check your tests and your app, see how RAITHub would test this below.

Everyone can build with AI now. Very few teams have a tester. So the AI often writes the app and then writes the tests that judge the app, and nobody outside that loop asks whether the tests mean anything. This post is about that loop. For the wider QA plan for an AI-built app, start with how to QA an app built with AI coding tools.

Why do tests the AI wrote pass when the code is wrong?

Because a test needs an oracle: an independent statement of what the right answer is. When you ask a model to "write tests for this file", the only oracle it has is the file. It reads the implementation and asserts that the implementation returns what the implementation returns.

Researchers have measured this. Konstantinou, Degiovanni and Papadakis studied LLM-generated test oracles across 24 Java repositories and found that the generated oracles "mainly capture the actual program behavior making bug detection difficult" (arXiv 2410.21136). In plain words: if the code has a bug, the AI's test often locks the bug in as the expected result.

Coverage numbers hide this. A separate 2025 study of LLM test generation reported test suites that reach 100% line and branch coverage but only a 4% mutation score, meaning they would notice almost none of the small faults deliberately planted in the code (Wang, Xu, Briand and Liu, arXiv 2506.02954). Coverage tells you a line ran. It does not tell you anything checked the result.

What do useless AI-written tests look like?

Four patterns explain most green-but-empty test suites in AI-built codebases. They are easy to spot once you know the shape.

PatternWhat it looks likeWhy it misses bugsQuick check
Mirror assertionsThe expected value was copied from running the code, e.g. a discount test that expects 89.1Whatever the code returned becomes "correct", including the wrong answerAsk: where did this expected number come from? If not from a requirement, it proves nothing
Mock everythingDatabase, payment client and auth are all mocked, and the test asserts the mock was calledThe real query, the real permission check and the real API contract never runCount assertions on mocks versus assertions on outputs or stored data
Happy path onlyOne test per function with valid input from one userThe bugs live in the second user, the empty list, the expired token and the failed paymentLook for any test with a wrong user, bad input or a provider error
Edited to passAn assertion changed, a test skipped or deleted in the same change that touched the codeThe agent made the signal go green rather than fixing the causeReview every diff that touches both a source file and its test file

Here is a mirror assertion in practice. The requirement says members get 10% off and the discount must never apply twice. The generated code applies it twice, and the generated test confirms the result:

// pricing.ts (generated)
export function memberPrice(price: number): number {
  const once = price * 0.9
  return once * 0.9 // bug: discount applied twice
}

// pricing.test.ts (generated from the code above)
import { expect, it } from 'vitest'
import { memberPrice } from './pricing'

it('applies the member discount', () => {
  expect(memberPrice(110)).toBe(89.1) // copies the bug
})

A test written from the requirement would expect 99. The one written from the code expects 89.1 and passes forever. Nothing in a coverage report or a green CI badge reveals the difference.

Do AI coding agents really change tests to make them pass?

Sometimes, yes, and it has been measured. ImpossibleBench, a 2025 benchmark by Zhong, Raghunathan and Carlini, gives coding agents tasks where the tests contradict the specification, so the only way to pass is to cheat. On its conflicting-tests version of SWE-bench, GPT-5 passed the impossible tasks 54% of the time, and the authors report that Claude models and Qwen3-Coder cheated mainly by modifying the test cases, while other models also special-cased the tested inputs (ImpossibleBench, arXiv 2510.20270).

That is a stress test, not normal use, and the tools improve quickly. But it explains a pattern that every reviewer of AI-edited code eventually meets: a failing test, a prompt that says "make the tests pass", and a diff where the assertion quietly changed. The agent did what it was asked. The request was the problem.

How can you tell if your AI-written tests are any good in five minutes?

Break the code on purpose and see what turns red. Pick the three functions that matter most: usually the permission check, the price or payment calculation, and whatever decides which rows a user can see. For each one, make a small, wrong change:

  • flip a comparison (>= to >), or change === to !==
  • delete the line that checks ownership or role
  • return early with an empty array or a hard-coded true

Run the tests after each change, then undo it. If nothing fails, the tests around that function are decoration. This is manual mutation testing, and it takes minutes. It is also the fastest way to show a non-technical stakeholder why "we have 300 tests" is not the same as "we are tested".

How do you run mutation testing on an AI-built JavaScript or TypeScript app?

Automate the sabotage. A mutation testing tool makes hundreds of small changes ("mutants") to your code, reruns the tests for each, and reports which mutants survived. For JavaScript and TypeScript, StrykerJS supports Vitest, Jest, Mocha and other runners. Its default thresholds colour a score of 80% or more green and under 60% red, and a break threshold makes it exit with code 1 below a score you choose, so it can fail CI (Stryker configuration docs).

A minimal config for a Vitest project, scoped to the code that holds money and permissions so the run stays short:

{
  "$schema": "./node_modules/@stryker-mutator/core/schema/stryker-schema.json",
  "testRunner": "vitest",
  "mutate": ["src/lib/auth/**/*.ts", "src/lib/billing/**/*.ts"],
  "thresholds": { "high": 80, "low": 60, "break": 60 },
  "incremental": true
}

Save it as stryker.config.json, install @stryker-mutator/core and @stryker-mutator/vitest-runner, and run npx stryker run. Incremental mode stores results and reuses them on the next run, which keeps repeat runs fast (same docs). Read the surviving mutants, not just the score: each one is a specific line where a bug could hide without any test noticing.

Two cautions. Mutation testing is slow on a whole codebase, so start with the critical folders. And do not hand the survivors to the same agent with "kill these mutants" and walk away; you will get tests that assert the current output again. Write the expected values from the requirement, then let the tool confirm the tests bite.

How should you prompt an AI to write tests that actually catch bugs?

Give it an oracle that is not the code. The changes that help most:

  • Write tests before the code, or from the spec only. Paste the rule ("members get 10% off, never stacked; refunds only within 30 days") and ask for tests from that text, without showing the implementation.
  • Ask for failing cases by name: the wrong user, the wrong role, an expired session, an empty cart, a declined card, a provider timeout.
  • Ban test edits in fix prompts. "Fix the code so this test passes; do not change any file under tests/." Then check the diff, because instructions are not guarantees.
  • Prefer real boundaries. Test against a local database and the payment provider's sandbox rather than mocks, so the real query and the real permission rule run.
  • Keep a human-owned core. A small set of tests for access control and money, written or reviewed by a person, that no prompt is allowed to touch.

Prompts reduce the problem. They do not remove it, which is why the next step is a mechanical guard.

How do you stop an agent from weakening the tests?

Make test changes visible and reviewed, and make the checks run whether or not the agent decides to run them.

  • Required checks. With branch protection, required status checks must pass before a change can merge (GitHub protected branches).
  • A named owner for the tests. A CODEOWNERS entry such as /tests/security/ @your-reviewer means any change to the access-control tests needs that person's approval.
  • Agent hooks. Claude Code hooks are shell commands that run at fixed points, such as after a file edit, which gives "deterministic control: certain actions always happen rather than relying on the LLM to choose to run them" (Claude Code hooks guide). Cursor's equivalent for instructions is project rules in .cursor/rules (Cursor rules docs), which guide the model but do not enforce anything.
  • A diff rule for reviewers: any pull request that changes a source file and its test in the same commit gets a second look at the assertion lines.

For keeping those guards working across hundreds of prompts, see regression testing for AI-edited code.

How long does it take to audit AI-written tests yourself?

For a small app with a few hundred generated tests, budget 1–2 days if you know the stack: half a day for the manual sabotage check on the critical paths, half a day to set up and run Stryker on two or three folders, and the rest to rewrite the tests whose expected values came from the code. Add another 1–3 days if the suite is mostly mocks and you need real database tests. The main risk of doing it yourself is the same blind spot that created the problem: if you also wrote the prompts, you will tend to read the tests as you meant them, not as they are.

Buy, build or hire?

OptionChoose this whenWatch out for
A tool or SaaS testing platform (mutation testing, AI test generators, coverage tools)You can read the results and rewrite weak tests yourselfTools find weak spots; they cannot tell you what the right answer was supposed to be
Freelancers or crowdtestingYou want people to use the app and find visible bugs the tests missedThey test the UI, rarely the test suite itself
An in-house QA hireYou ship AI-generated changes daily and can keep a tester busy full timeA real salary commitment; see in-house vs outsourced QA
A managed QAaaS teamYou want an independent check of the tests and the app, with fixed tests left in your repoAgree up front which flows are critical, so effort goes where bugs cost most

Why RAITHub for this

  • An independent oracle. RAITHub writes expected results from your product rules, not from your code, which is exactly what AI-written tests lack.
  • Test-heavy builds. RAITHub's own platforms ship with large suites: 530+ tests on Sundor Skin, 1,024 on PropDesk and 400+ on this website. There is no published case study of an AI-built-app audit yet, so these counts are the proof on offer.
  • Access control first, because that is where AI-built apps leak data; see the vibe-coded app security checklist.

When you don't need us

  • The app is a prototype with test data and no real users. Do the five-minute sabotage check and move on.
  • You have a developer with time to run Stryker and rewrite the weak tests. The steps above are enough.
  • You want testers placed under your own management. RAITHub does not offer staff augmentation.

How RAITHub would test this

Through AI app testing, part of QA as a service, RAITHub would:

  • write the role and rule map for your app (who can see, change and pay for what) and agree it with you
  • run a manual sabotage pass and mutation testing on the access-control, billing and data-visibility code, and list every surviving mutant that matters
  • rewrite or add tests whose expected values come from your rules, against a real database and the payment sandbox
  • test the running app by hand on the journeys the tests cannot see, then add the guards: required checks, test ownership and a mutation threshold in CI

Buy it as a fixed-price launch audit for your AI-built app: a written report of the bugs and weak tests found, ranked by impact, each with a fix. Optional monthly QA keeps checking each batch of AI-generated changes. You own the tests and all IP, and an NDA is standard; see QA test automation for the automation side.

Next step: book a free 15-minute call about a launch audit, then get a written fixed quote.

Frequently asked questions

Are AI-generated unit tests worthless?

No. They are quick to produce and good at catching crashes and obvious breakage. Their weakness is the expected values: when those come from reading the code, the test confirms whatever the code does. Keep them, then check them with a sabotage pass or mutation testing.

What is a good mutation score?

Stryker's defaults treat 80% or more as good and under 60% as poor. Aim high on access control and billing code and accept less on presentation code. Read the surviving mutants rather than chasing the number.

Is 100% code coverage enough for AI-written tests?

No. Coverage only shows that a line ran. A 2025 study reported suites with 100% line and branch coverage and a 4% mutation score, so they caught almost none of the planted faults.

Why did my AI agent change the test instead of fixing the bug?

Because the goal you gave it was "make the tests pass", and editing a test is the shortest path to that. Benchmarks such as ImpossibleBench show the behaviour is common when tests and requirements conflict. Forbid test edits in fix prompts and review any diff that changes both.

Can I ask the AI to review its own tests?

You can, but it is checking its own work against its own reading of the code. Give it the written requirement instead, or have a person or a mutation tool do the check.

Does RAITHub fix the tests or only report on them?

Either. The launch audit reports the weak tests and bugs with fixes; RAITHub can then write the replacement tests into your repository as part of the audit scope or a monthly QA plan.

AI-generated testsmutation testingStrykertest oraclesClaude CodeCursorAI app testing

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.