Founder & Lead Engineer, RAITHub
To QA a codebase written with Cursor or Claude Code, make tests the gate rather than the agent's judgement: give the agent a test command it must pass before it stops, run the same tests in CI on every pull request, use AI review as a first pass, have someone who did not write the change review it, and test roles, payments and real devices by hand before release.
If you would rather have your Cursor or Claude Code project tested for you, see how RAITHub would test this below.
Cursor and Claude Code are different from app builders such as Lovable or Bolt. They work inside a repository you own, on whatever stack you chose, and they can run your commands. That is good news for QA: everything a professional team uses, from unit tests to CI gates, is available. The catch is that the agent writes code faster than anyone reviews it, and everyone can build with AI now while almost no one has a tester. This post is about closing that gap without slowing the agent down.
Why does code from Cursor or Claude Code need its own QA process?
Because the agent decides when it is done, and "done" means "looks done" unless you give it something stricter. Anthropic's own Claude Code best practices say it directly: "Claude stops when the work looks done. Without a check it can run, 'looks done' is the only signal available, and you become the verification loop." The same page lists "the trust-then-verify gap", where the agent "produces a plausible-looking implementation that doesn't handle edge cases", and its fix: "If you can't verify it, don't ship it."
The wider evidence points the same way. Veracode tested code from more than 100 models and found an OWASP Top 10 weakness in 45% of cases (Veracode 2025 GenAI Code Security Report). The model is not careless; it is solving the prompt in front of it, and the prompt rarely says "and make sure user B cannot see user A's invoices".
The general method, including how to check that AI-written tests actually assert something, is in how to QA an app built with AI coding tools. This post covers the parts specific to agents that live in your repository.
What QA layers does a Cursor or Claude Code project need?
| Layer | What it is | What it catches | What it misses |
|---|---|---|---|
| Instructions | Rules files: .cursor/rules, AGENTS.md, CLAUDE.md | Repeated style and pattern mistakes | Anything the agent decides to ignore; rules are advice |
| Deterministic gate in the agent | Claude Code hooks or Cursor hooks that run tests | Broken builds and failing tests before the agent stops | Behaviour no test describes |
| CI on every pull request | Typecheck, lint, tests, dependency audit | The same, for every change from every tool and person | Same as above |
| AI review | Cursor Bugbot, Claude Code /security-review and its GitHub Action | Common bugs and vulnerability patterns in the diff | Your business rules and roles |
| Independent review | A person, or a fresh agent session, that did not write the change | Assumptions the author made | Runtime behaviour on devices |
| Human testing | Exploratory, role-based, payment and device testing | What users hit first | Nothing, if scoped well; it is just slower |
How do I make tests a gate the agent cannot skip?
Use hooks, not instructions. Anthropic's docs distinguish the two: "Unlike CLAUDE.md instructions which are advisory, hooks are deterministic and guarantee the action happens." A Stop hook runs when Claude finishes responding, and if it exits with code 2, Claude Code "prevents Claude from stopping" and feeds the error back so the agent keeps working (Claude Code hooks reference). In .claude/settings.json:
{
"hooks": {
"Stop": [
{
"hooks": [
{
"type": "command",
"command": "npm run typecheck 1>&2 && npm test --silent 1>&2 || exit 2"
}
]
}
]
}
}
The output is sent to stderr because that is what Claude Code shows the agent as the reason it cannot stop yet. Keep the command fast; a five-minute suite on every turn will be resented and switched off. Run the unit tests and the access-control tests here, and leave end-to-end tests to CI. The hooks docs also describe a cap on consecutive blocks, so a suite that can never pass will not loop forever.
Cursor has its own hooks, and its documentation says it can load hooks configured for Claude Code once third-party configs are enabled (Cursor third-party hooks). Either way, the same commands run again in CI, because a gate on one developer's machine is not a gate on the repository.
What should the rules files say about testing?
Rules are advisory, but good ones reduce the number of times the gate has to fire. Cursor's rules docs describe project rules in .cursor/rules, plus AGENTS.md; Claude Code reads CLAUDE.md at the start of every session. Keep them short and specific. A testing section that works for both:
# Testing
- Run: npm run typecheck && npm test before saying a task is done
- Every API route and server action checks the session AND that the
record belongs to the caller. Add a test where a second user is refused.
- Never weaken or delete an existing test to make a change pass.
- New dependencies need a reason in the PR description.
- Do not edit .env files, migrations already applied, or CI config.
Anthropic's guidance is to prune: for each line, ask "Would removing this cause Claude to make mistakes?" A rule file of 400 lines is a rule file the agent skims.
Do Bugbot and /security-review replace a human reviewer?
No, and their makers say so. Cursor's Bugbot "analyzes PR diffs and leaves comments with explanations and fix suggestions", runs on every PR update, and reads project context from .cursor/BUGBOT.md. Claude Code's /security-review command and GitHub Action look for issues such as SQL injection, XSS, authentication flaws and insecure data handling, with the caveat that they "should complement, not replace, your existing security practices and manual code reviews" (Anthropic support).
Use them; they are cheap and fast and catch real bugs. Then add the review they cannot do. Anthropic's best practices recommend a writer and reviewer split, because "a fresh context improves code review since Claude won't be biased toward code it just wrote". A second session is good. A person with the role map in hand is better for the changes that touch data access or money.
What does a reviewer check in an AI-written pull request?
- Tests changed alongside code. A diff that edits an assertion to match new behaviour needs a reason. Agents do this to make a red test green.
- Authorisation on every new route or action. Session checked, and the record's owner checked, on the server.
- New dependencies. Real, maintained and needed; models sometimes name packages that do not exist (USENIX Security 2025).
- Agent configuration files. Changes to
.cursor/,.claude/, MCP configs and CI files deserve the same review as code. CVE-2025-54135 in Cursor showed how a prompt injection could write an MCP config that ran commands without approval; it was fixed in Cursor 1.3.9 (Tenable FAQ). Keep your tools updated, and treat config diffs as security-relevant. - Scope creep. Files changed that the task did not mention. Ask why.
- Migrations. Destructive changes need a plan; see when a database migration breaks production.
What still needs a human tester?
Everything that happens outside the repository. Tests describe what someone thought to write down; users do what nobody wrote down. Before each release, a human should walk the core journeys per role, try another user's data through the API, run payment failures in test mode, and use the app on real phones and in Safari. Exploratory testing, described in the exploratory testing guide, is how testers find the bugs that no prompt and no test anticipated.
How long does it take to set this up yourself?
For an existing repository with some tests, about one to two days: a few hours for hooks and CI, an hour for a short rules file, half a day turning on AI review and agreeing the review checklist, and a day writing the first access-control tests. A first human test pass of a small app takes another two to four days. These are estimates, not a quote. The main risk of doing it alone is that one person, working with one agent, is both author and reviewer; the gate catches what the tests know about and nothing else.
Buy, build or hire?
| Route | Choose this when |
|---|---|
| A tool or SaaS testing platform (Bugbot, /security-review, scanners) | You want automatic review on every pull request. Turn these on whatever else you do. |
| Freelancers or crowdtesting | You need exploratory and device testing before a release and can manage the testers and the triage. |
| An in-house QA hire | Several developers are merging agent-written changes daily and you need someone who knows the product deeply. |
| A managed QAaaS team | You want the gates set up, a launch audit, and monthly retesting of each release, without hiring. |
Why RAITHub for this
- Repositories are home ground. RAITHub builds and tests in TypeScript, Next.js, Node and PostgreSQL with CI on every pull request. Its builds carry large suites: 750+ tests across 217 endpoints on TheSkinProof (the founder's own venture, not a client), 530+ on Sundor Skin, 1,024 on PropDesk.
- Tests left in your repository, wired into your CI and your agent's hooks, not kept in a separate tool.
- Honest scope. There is no published case study yet of testing a Cursor or Claude Code project; the numbers above are from platforms RAITHub built.
When you don't need RAITHub
- You are an experienced developer and the project is pre-launch with no real users. The hook, CI and rules above may be enough for now.
- You need a certified penetration test or a compliance attestation. RAITHub's security testing is application-level, against OWASP guidance.
- You want a tester placed in your team under your management. RAITHub does not offer staff augmentation.
How RAITHub would test this
- Repository review: the existing tests, how much they actually assert, the CI pipeline, hooks and rules files, and recent agent-written changes to data access.
- Gates: a fast test command for the agent's Stop hook, a CI workflow on every pull request, and access-control tests for each data-holding route.
- Human testing: role-based and exploratory testing of the core journeys, payments in test mode, and real iOS and Android devices and browsers.
- Security pass: OWASP-based checks of authorisation, secrets, validation and dependencies; see the OWASP testing checklist.
What you buy: a fixed-price launch audit, then an optional monthly QA plan that retests each release. You receive: a ranked bug report with reproduction steps and a suggested fix for each issue, the tests and CI changes in your repository, and full IP under NDA. See AI-built app testing, QA test automation and QA as a service. Next step: a free 15-minute call, then a written fixed quote.
Shipping a Cursor or Claude Code project? Ask for a launch audit quote.
Frequently asked questions
Can Claude Code or Cursor test their own code?
They can write and run tests, and they should. But a check the agent wrote for its own code tends to confirm what the code does. Make the tests a deterministic gate, review them, and have someone who did not write the change test the app.
What is the difference between CLAUDE.md rules and hooks?
Rules in CLAUDE.md or .cursor/rules are instructions the agent may follow. Hooks are commands that run every time at a fixed point, such as when the agent tries to stop. Anthropic's docs describe hooks as deterministic and rules as advisory.
Is Cursor Bugbot enough code review?
It is a useful first pass on every pull request. It reviews the diff for bugs and security issues, but it does not know your roles and business rules, so changes to data access or payments still need a human reviewer.
Should agent configuration files be code-reviewed?
Yes. Rules, hooks, MCP configs and CI files change what the agent can do. A 2025 Cursor vulnerability, CVE-2025-54135, involved an MCP config written through prompt injection, so treat these diffs as security-relevant.
What should I test first in an agent-written codebase?
Authorisation: whether one user can read or change another's data through the API. Then payments and webhooks, then the main journeys on real phones. Those are the failures that hurt users.
Can RAITHub fix the issues it finds?
Yes, under a separate fixed quote, or your team can fix them, with or without the agent. Each issue in the report includes reproduction steps and a suggested fix.
Related posts
Ready to discuss your project?
Book a free 15-minute technical audit with our engineering team.