Founder & Lead Engineer, RAITHub
RAITHub tests software with a gated pyramid: type-checks, unit, integration, contract and end-to-end tests run as gates in CI, and a failure at any layer stops the release. Performance budgets, synthetic monitoring and human code review add further layers, because green CI only proves what someone thought to test.
This post describes how RAITHub's own quality system works, with numbers counted in the repositories of platforms we built. It is the practical version: what runs, what blocks, what broke, and what you receive.
What is QA-first development?
QA-first development means the test suite and its CI gates are part of the scope from the first commit, not a phase that starts after features are "done". Quality is built into how code merges, not inspected at the end.
In practice that changes three things:
- Tests ship with the feature. The pull request that adds a feature also adds the tests that prove it works.
- The pipeline decides. Whether code can merge is decided by automated gates, not by someone remembering to check.
- Review targets the gaps. Human review focuses on the bugs automated tests are structurally bad at finding, such as data flowing across boundaries.
At RAITHub, a real test suite with CI gates is included in every engagement, alongside founder-reviewed architecture, weekly demos, runbooks and handover.
What does RAITHub's test pyramid look like?
Five layers, fastest and least expensive first: type-checks, unit tests, integration tests, contract tests and end-to-end tests. Each layer is a CI gate, and performance budgets and synthetic monitoring sit alongside them.
The order matters. A type error is found in seconds and costs almost nothing to fix. A broken checkout found by an end-to-end test needs a full browser run to detect. Pushing each check as low in the pyramid as it can go keeps the pipeline fast enough that nobody is tempted to skip it.
| Layer | What it catches | Blocks the release? |
|---|---|---|
| Type-checks | Wrong data shapes passed between modules, missing null handling, a renamed field not updated everywhere | Yes |
| Unit tests | Logic errors inside one function: validation rules, pricing, permission checks | Yes |
| Integration tests | Code that works alone but fails against a real database: queries, migrations, transactions | Yes |
| Contract tests | An API response that changed shape and would break the frontend or an integration | Yes, on every pull request |
| End-to-end tests (Playwright) | A user journey that no longer completes: sign-up, login, checkout, publishing from the admin | Yes |
| Performance budgets (k6 for load) | A page or endpoint slower than its budget for LCP, INP or P95 latency | Yes, enforced in CI |
| Synthetic monitoring and error tracking (Sentry) | A journey that fails in production after release, for reasons outside the code | Alerts after release and triggers the incident playbook |
The tools matter less than the rule. If a layer fails, the change waits until it passes.
What blocks a release at RAITHub?
Any failing gate blocks a release: a type error, a failing unit, integration, contract or end-to-end test, a blown performance budget, or a project-specific safety check. The project-specific checks are where the pyramid earns its keep.
Two examples from platforms we built:
- Sundor Skin (B2B wholesale platform, deployed September 2026). CI replays all 76 database migrations on an empty database and fails the build if any buyer-scoped table lacks a row-level-security policy. A forgotten policy on a new table is exactly the mistake that shows one buyer another buyer's data, so it is a build failure, not a review comment.
- TheSkinProof (multi-vendor ecommerce, live since Q2 2025; the founder's own marketplace, built and run by RAITHub). Coverage thresholds are enforced, so a change that drops coverage below the threshold fails the pipeline instead of quietly eroding the suite.
The pattern behind both: when a class of mistake is expensive and a machine can detect it, turn it into a gate.
How many automated tests do RAITHub's platforms have?
From 400+ to 1,024 per platform, counted in each project's repository. The count is not the goal; it is evidence that the tests grew with the features instead of being added at the end.
| Platform | Scope | Automated tests | Notable safeguard |
|---|---|---|---|
| RAITHub website | This site, its admin and content management | 400+ (September 2026) | Regression test per update schema (see below) |
| TheSkinProof (the founder's own marketplace) | 225K+ lines of code, 217 API endpoints, 5 role-based portals | 750+ | Enforced coverage thresholds |
| Sundor Skin | 146 tables, 12 staff roles, 88 permission codes | 530+ across unit, integration and an IDOR security suite, plus 20 SQL assertion suites | Row-level-security check across 76 migrations; hash-chained, append-only audit log |
| PropDesk | 130+ API endpoints, 60+ pages, 4 roles, 5 daily jobs | 1,024 | Details in the case study |
IDOR stands for insecure direct object reference: changing an ID in a request to read or edit someone else's record. On a wholesale platform where each buyer must only see their own prices and orders, a suite that deliberately tries this is far cheaper than hearing about it from a buyer.
The details are in the Sundor Skin case study, the TheSkinProof case study and the PropDesk case study.
How does RAITHub find bugs before users see them?
Automated gates catch the bugs someone anticipated. Review catches the rest, by tracing data across boundaries, reproducing the deploy target's exact commands, and confirming that anything the framework could silently ignore is actually loaded.
Four examples from preparing this website for launch in September 2026. Each was small and easy to miss.
1. A silent data wipe that passed every check
On 27 September 2026, during pre-launch code review, a reviewer traced what the admin's "reorder" and "publish toggle" controls actually sent to the database. Both sent partial updates such as {id, order}. The update schema was built with zod 4's .partial(), which still applies .default() values for fields that were left out:
const caseStudy = z.object({
id: z.string(),
order: z.number().int(),
published: z.boolean().default(false),
techStack: z.array(z.string()).default([]),
images: z.array(z.string()).default([]),
})
caseStudy.partial().parse({ id: 'x', order: 2 })
// zod 4 returns:
// { id: 'x', order: 2, published: false, techStack: [], images: [] }
So moving a case study to a new position would also unpublish it and wipe its tech stack and images. Tests, type-checks and lint were all green, because no test had asked what a partial update leaves out.
The fix was a small helper that strips defaults before calling .partial(), plus a regression test for every update schema asserting that a partial update returns exactly the fields sent. Simplified:
for (const [name, schema, input] of cases) {
it(name + ' returns exactly the fields sent', () => {
const result = schema.safeParse(input)
expect(result.success).toBe(true)
if (result.success) expect(result.data).toEqual(input)
})
}
That work added 19 tests and took the suite from 381 to 400; the fix shipped in the same pre-launch review. The full write-up is in zod 4 .partial() keeps default values.
2. A deploy that passed locally and failed on Vercel
Also on 27 September 2026, a security upgrade moved nodemailer from v7 to v10 to clear a high-severity advisory. next-auth declares nodemailer as an optional peer dependency, and no next-auth 4.x release accepts nodemailer 10; the latest, 4.24.15, wants ^7.0.7. Vercel's npm install therefore failed with ERESOLVE. Local checks had used npm ci, which installs from the lockfile without that re-check, so everything passed locally.
Downgrading would have brought the advisory back. The app only uses next-auth's credentials login, so next-auth never loads nodemailer. The fix was an npm overrides entry pointing next-auth's nodemailer at the app's own version, verified with a fresh npm install, npm ls showing no invalid dependencies, and npm audit reporting 0 vulnerabilities. The lesson is now a rule: reproduce the exact install command the deploy target runs. Details are in the next-auth and nodemailer ERESOLVE deep dive.
3. A renamed file that could have dropped auth without an error
In Next.js 16, middleware.ts became proxy.ts. A middleware file the framework silently ignores produces no error; it just stops protecting routes. So the review did not stop at "the build passed". It confirmed that the build output lists "Proxy (Middleware)", that admin pages still redirect to login, and that the admin API returns 401 without a session.
4. A rate limit that let anyone lock the admin out
Before launch, login was rate-limited only per email address. That slows password guessing against one account, but it also lets anyone who knows the admin's email keep the real admin locked out. Login now uses three buckets (email plus IP, per IP, and per email across IPs) and hashes the keys to a fixed length. This is an abuse-case question, and abuse cases are part of the security review.
Why is a green CI pipeline not enough?
Because green CI proves what you tested, not what you did not. Each example above sat in a gap between layers: between the UI and the database, between local and deploy installs, between a file and the framework that reads it.
So RAITHub's reviews ask a short list of questions that tests rarely answer:
- What does this request actually write? Follow one real payload from the button to the database row.
- Does the check run what production runs? Same install command, same build, same environment.
- Would a failure here be loud? If a file, flag or policy can be ignored silently, prove it is loaded.
- Who could abuse this? Rate limits, IDs in URLs, and anything an unauthenticated visitor can trigger.
When one of these questions finds a bug, the fix includes a test that would have caught it, so the same gap cannot reopen. That is how the RAITHub website suite went from 381 to 400 tests on 27 September 2026. For related checklists, see the Next.js production checklist and the API security checklist.
What do clients receive from RAITHub's QA work?
Every engagement includes a real test suite with CI gates in your own repository, weekly demos, runbooks and handover. The QA and Reliability retainer adds performance budgets, synthetic monitoring, incident playbooks and a weekly reliability report.
Included in every engagement:
- The test suite and pipeline, in your repository. With full IP assignment, the tests belong to you and keep protecting you if you ever change vendors.
- Working software every Friday. Weekly demos, with branch previews so you can click through a change before it merges.
- Contract tests on every pull request. An API change that would break a consumer is caught at review time, not in production.
- Runbooks and handover. Written operating knowledge, so running the system does not depend on one person's memory.
Added by the QA and Reliability retainer (monthly, tiered):
- The full pyramid gated in CI: type-checks, unit, integration, contract and end-to-end.
- Performance budgets in CI for LCP, INP and P95 latency.
- Synthetic monitoring of key journeys and incident playbooks for when one fails.
- A weekly reliability report, using Playwright, k6 and Sentry as the core tooling.
An NDA is signed before any detailed discussion, so you can share the real codebase early. Certification status is on the security page.
Does QA-first make software slower or more expensive?
It moves cost earlier, where it is smaller. RAITHub's SaaS and MVP builds (4 to 6 weeks) and Backend and API builds (6 to 12 weeks) are fixed scope with the test suite included, so testing is never a line item you have to negotiate away.
The cost of skipping tests is invisible until it lands: a record silently wiped, a failed deploy, an admin locked out. Each example above was caught by a check or a review, and each fix left behind a test or a rule so it stays fixed.
The same gates run on client work: RAITHub has applied them across 8 client projects, including a fintech dashboard, a healthcare scheduling app and four SaaS products ranging from early MVPs to platforms with 10,000+ users. For how the weekly rhythm works, read how RAITHub delivers software, and for engagement options, see pricing.
How can you check whether a software vendor really tests?
Ask to see the pipeline, not a slide. A vendor with a real QA-first practice can show you what fails a build and a bug their process caught before users did.
- Which checks block a merge, and can I see a red build?
- How many automated tests does a comparable project have, and how were they counted?
- Show me a bug your review caught that the tests missed, and the regression test you added afterwards.
- Do your checks run the same install and build commands as the deploy target?
- Will the tests live in my repository, with IP assigned to me, when we finish?
These are the questions we would want asked of us, and this post is our answer. If you want the same answers about your own project, start with the free 15-minute technical audit on the contact page, or browse the platforms we built and our services.
Last reviewed: 28 September 2026
To have RAITHub build or audit a test suite for your product, see the QA and test automation service, or run through the pre-launch QA checklist yourself first.
Frequently asked questions
How does RAITHub test software?
Every change passes a CI pyramid of type-checks, unit, integration, contract and end-to-end tests, and a failure at any layer blocks the release. Human review then targets the gaps tests miss, such as data flowing from the UI to the database and differences between local and deploy environments.
What is the difference between QA-first development and adding tests later?
In QA-first development, tests and CI gates are written with each feature from the first commit, so the suite grows with the product. Adding tests later usually means testing what already exists, after the bugs have shipped and the design has set around them.
What blocks a release in RAITHub's pipeline?
Any failing test layer, a blown performance budget for LCP, INP or P95 latency, or a project-specific gate. On Sundor Skin, for example, the build fails if any buyer-scoped table lacks a row-level-security policy after replaying all 76 migrations.
Which testing and monitoring tools does RAITHub use?
Playwright for end-to-end tests, k6 for load and latency, and Sentry for production error tracking.
Are tests included in the price or charged separately?
A real test suite with CI gates is included in every RAITHub engagement. The QA and Reliability retainer is a separate monthly, tiered option for ongoing monitoring, performance budgets, incident playbooks and a weekly reliability report.
Is RAITHub SOC 2 or ISO 27001 certified?
No. RAITHub follows practices aligned with those standards, signs an NDA before detailed discussion, and assigns full IP to the client. The security page describes the practices in detail.
Can RAITHub test a codebase it did not build?
Yes. The Code Rescue engagement starts with a 2-week diagnostic of an existing codebase and puts tests around its critical paths first. The Code Rescue playbook explains the process step by step.
How do I get started with RAITHub?
Book the free 15-minute technical audit through the contact page, or email hello@raithub.com. RAITHub is founder-led by Rupak Amin, Founder and Lead Engineer, and architecture on every project is founder-reviewed.
Related posts
Ready to discuss your project?
Book a free 15-minute technical audit with our engineering team.