The Code Rescue Playbook: How to Take Over a Codebase You Didn't Write
Founder & Lead Engineer, RAITHub
To take over a codebase you didn't write, secure access first, then diagnose before changing anything: map how it deploys, what it depends on and where it is fragile. Stabilise the paths that move money and data, decide module by module what to keep or replace, and end with a documented handover.
This playbook is for founders and product owners whose developer has left, whose agency has gone quiet, or who have inherited a system nobody understands. It works in-house or with outside help. It also covers when you should not hire a rescue team, and how RAITHub's Code Rescue engagement follows these steps.
What is software project rescue?
Software project rescue is taking responsibility for a codebase that has lost its owner or its momentum, and making it safe to change again. The goal is a system your team can ship on, not a rewrite for its own sake.
A rescue is different from normal maintenance because the knowledge is missing. Nobody can say with confidence how it deploys, why a module exists, or what breaks if you touch it. The work is mostly recovering that knowledge, then protecting what matters while you fix what hurts.
What are the warning signs a codebase needs rescue?
The clearest sign is fear: people avoid changing parts of the system because they cannot predict what will break. The others are usually symptoms of the same missing knowledge.
- Only one person can deploy, or deploys happen from someone's laptop rather than a pipeline.
- Fixed bugs come back. There are no regression tests, so nothing stops old mistakes returning.
- Small changes take weeks. Estimates keep growing because every change touches something unexpected.
- Nobody knows where the secrets live. API keys, database passwords and hosting accounts sit in personal accounts or chat messages.
- Dependencies are years behind and security advisories pile up because upgrading feels too risky.
- The developer or agency has stopped responding, or the relationship has ended on bad terms.
- Production incidents are found by customers, not by monitoring.
One sign on its own is normal technical debt. Three or more, and especially the first and fourth together, mean the business is carrying risk it cannot see.
What are the phases of a code rescue?
Five phases: secure and freeze, diagnose, stabilise critical paths, decide keep or replace, and hand back. Each phase produces something you can hold, so you always know what you have paid for.
| Phase | When | What happens | Output |
|---|---|---|---|
| 1. Secure and freeze | First 48 hours | Take ownership of accounts, back up data, stop risky changes, try a clean build | Access inventory, a verified backup, a build that works or a list of why it does not |
| 2. Diagnose | 2 weeks | Codebase audit, infrastructure audit, risk register | An audit memo and a plan with a fixed scope |
| 3. Stabilise critical paths | After the diagnostic, as planned | Tests, monitoring and fixes around checkout, auth, payments and data | Critical paths under test and gated in CI |
| 4. Keep or replace | As planned, module by module | Replace what is costing you, keep what works | Replaced modules, documented decisions for the rest |
| 5. Hand back | End of the engagement | Knowledge transfer, paired reviews, runbooks | A team that can deploy, change and recover the system without the rescuer |
My developer left. What should I do in the first 48 hours?
Secure access and protect the data before anyone changes code. In the first 48 hours, the goal is to make sure the business, not a person, controls every account the system depends on.
Work through this list in order:
- Take ownership of accounts. Code repository, hosting, domain registrar and DNS, database, email delivery, payment provider, error tracking, analytics and any app store accounts. Move each to a company-owned login with at least two admins.
- Back up production data and prove the backup restores. A backup you have never restored is a hope, not a backup.
- Inventory secrets before rotating them. List every key and password the system uses and where it is configured. Then rotate the ones the departed person held. Rotating without an inventory is an easy way to take production down by accident.
- Freeze non-urgent changes. Only security fixes and outages go out until the diagnostic is done.
- Find out what is actually deployed. Which commit is running in production? Does it match the main branch, or were fixes made directly on a server?
- Try a clean build using the exact command the deploy target runs. Not the command that works on someone's machine. We explain why below.
None of this needs deep knowledge of the code. It needs someone methodical with admin rights. If you can only do one thing today, do step 1.
What should you prepare before calling a rescue team?
Prepare access, context and a short list of what hurts. The more of this you have ready, the less of the diagnostic is spent on archaeology.
- Access: read access to the repository and hosting, or at least a list of which accounts exist and who controls them.
- Ownership documents: the contract with the previous developer or agency, and whether it assigns IP to you. If it does not, sort that out first; you cannot safely build on code you do not own.
- What hurts: the three things costing you most right now, in business terms. "Checkout fails for some card types" is more useful than "the code is bad".
- What must not break: the journeys that make money or hold customer data.
- Any history: old tickets, chat exports and documentation, even if outdated.
- Deadlines: a launch, fundraise or audit that sets the real timeline.
Expect a serious rescue team to ask for an NDA before you share the code. RAITHub signs one before any detailed discussion.
What does a 2-week diagnostic cover?
A codebase audit, an infrastructure audit and a risk register. Together they answer three questions: what do we have, what could hurt us, and what should we fix first.
Codebase audit
How the code is structured, how data moves through it, what tests exist and whether they run, which dependencies are outdated or carry known advisories, and where logic is duplicated or dead. The most valuable part is tracing real data flows end to end: what a button actually writes to the database, not what the function name suggests.
Infrastructure audit
How the system is hosted, built and deployed; where secrets live; whether backups exist and restore; what monitoring exists; and which accounts are single points of failure.
Risk register
Every finding goes into one list with its likelihood, impact and the cost to fix. This turns "the code is a mess" into a ranked plan. It is also the document that stops a rescue sliding into an open-ended rewrite, because every piece of work has to point back to a risk on the list.
What does diagnosis look like on a real codebase?
It looks like tracing data and reproducing environments until the story is consistent. The four examples below are real, dated and from RAITHub's own website repository in September 2026; we do not publish details of other people's codebases.
Each shows a diagnostic habit that matters even more in an unfamiliar codebase.
Example 1: every check was green, and the data was still being wiped
Symptom (found in pre-launch code review, 27 September 2026): the admin's "reorder" and "publish toggle" controls sent partial updates such as {id, order}.
What the checks said: tests, type-checks and lint were all green.
What tracing found: the update schema used zod 4's .partial(), which still applies .default() values for omitted fields. Parsing {id, order: 2} also produced published: false, techStack: [] and images: [], so the database update would unpublish the case study and wipe its tech stack and images.
const update = caseStudySchema.partial()
update.parse({ id: 'x', order: 2 })
// zod 4: { id: 'x', order: 2, published: false, techStack: [], images: [] }
// expected: { id: 'x', order: 2 }
Fix: a small helper that strips defaults before .partial(), plus a regression test per schema asserting a partial update returns exactly the fields sent. That added 19 tests, taking the suite from 381 to 400, and the fix shipped in the same pre-launch review. Deep dive: zod 4 .partial() keeps default values.
Rescue lesson: green CI proves what was tested, not what was not. In an inherited codebase, trace data from the UI to the database for every write path that touches money or customer records.
Example 2: it builds on my machine, and fails on the server
Symptom (27 September 2026): the Vercel deploy failed with ERESOLVE after a security upgrade moved nodemailer from v7 to v10 to clear a high-severity advisory.
What the checks said: local checks passed, because they used npm ci, which installs from the lockfile without re-checking peer dependencies the way npm install does on Vercel.
What tracing found: next-auth declares nodemailer as an optional peer, and no next-auth 4.x release accepts nodemailer 10; the latest, 4.24.15, wants ^7.0.7. Downgrading would reintroduce the advisory. But the app only used next-auth's credentials login, so next-auth never loads nodemailer.
Fix: an npm overrides entry pointing next-auth's nodemailer to the app's version:
"overrides": {
"next-auth": {
"nodemailer": "$nodemailer"
}
}
Verified with a fresh npm install, npm ls (no invalid dependencies) and npm audit (0 vulnerabilities). Deep dive: the next-auth and nodemailer ERESOLVE conflict.
Rescue lesson: reproduce the exact install and build command the deploy target runs. "It works on the old developer's laptop" is false comfort.
Example 3: a missing error is not proof that something works
Symptom: in Next.js 16, middleware.ts was renamed to proxy.ts.
Why it mattered: if the framework silently ignored the file, auth gating on admin routes would disappear with no error at all.
What verification looked like: the build output lists "Proxy (Middleware)", admin pages still redirect to login, and the admin API returns 401 without a session.
Rescue lesson: after any framework upgrade in an inherited system, prove that security-critical configuration is loaded. Do not infer it from a quiet build.
Example 4: a safeguard that could be turned against the owner
Symptom: login was rate-limited only per email address.
Why it mattered: anyone who knew the admin email could keep the admin locked out indefinitely.
Fix, before launch: three buckets (email plus IP, per IP, and per email across IPs), with keys hashed to a fixed length.
Rescue lesson: audit safeguards for abuse, not just presence. A rate limit that exists can still be the wrong rate limit.
How do you stabilise the critical paths?
Put tests and monitoring around checkout, auth, payments and data before fixing anything else in them. You cannot safely change behaviour you have not pinned down.
- Write characterisation tests first. These record what the system does today, even where it is wrong, so any change in behaviour is visible and deliberate.
- Add CI gates. A pipeline that runs type-checks and those tests on every change, and blocks the merge when they fail. Contract tests on APIs catch changes that would break a frontend or an integration.
- Add error tracking and monitoring so production problems reach you before customers do.
- Fix in risk-register order. Highest impact first, each fix with a regression test, as in Example 1.
- Make deploys repeatable and reversible. One documented way to deploy, and one to roll back.
This is the same QA-first approach RAITHub uses on new builds, applied to existing code. The QA-first development post describes the test pyramid and what blocks a release.
Should you keep the existing code or replace it?
Usually keep most of it. Replace the specific modules that cost you money or time every month, and keep what works, even if it is not how you would have written it.
| What you find | Lean towards | Why |
|---|---|---|
| Code that works and has users, but is untidy | Keep, add tests | Working code encodes business rules that are easy to lose in a rewrite |
| A module that causes bugs or delays every sprint | Replace that module | The cost is recurring and measurable; the scope is contained |
| An outdated dependency with known advisories | Upgrade or replace the dependency | A dependency problem is rarely a reason to rewrite the application |
| A data model that fits the business | Keep, even if the code around it goes | Data migrations are the riskiest part of any replacement |
| A data model that fundamentally contradicts how the business works | Plan a staged replacement | Every feature built on it will keep fighting it |
When you do replace, replace in slices behind stable interfaces, with the old path still running until the new one is proven. A slice you can roll back is a decision; a big-bang rewrite is a bet.
When should you not hire a rescue team?
When the code is not the problem, when the code is not worth saving, or when you need a permanent team rather than a project. Saying this plainly saves you money and saves everyone time.
- The product is changing direction. If the next version will do something substantially different, stabilising the current one may be wasted effort.
- It is a prototype with no real users or data. If nothing depends on it, rebuilding cleanly can cost less than understanding it.
- The platform itself is at a dead end. If the framework or language version no longer gets security fixes and has no realistic upgrade path, a planned rebuild may be the responsible choice.
- The real problem is unclear requirements. No codebase survives a spec that changes every week. Fix the product process first.
- You need a full-time owner for years. A rescue can hand back to an in-house team; it is not a substitute for hiring one.
- You do not own the code. Resolve IP and access with the previous vendor before paying anyone to work on it.
When is a rewrite the right answer?
A rewrite is right when the data model fights the business, when the platform is at end of life, or when the existing code is a prototype nobody depends on. Even then, do it in stages, keep the data, and run the diagnostic first so the new system does not repeat the old one's mistakes. A good diagnostic is useful whichever way you decide.
How do you hand the codebase back safely?
Through knowledge transfer, paired reviews and runbooks, so your team can deploy, change and recover the system without the rescuer. A rescue that leaves you dependent on a new single person has only moved the risk.
- Runbooks: how to deploy, roll back, restore from backup and respond to the alerts that monitoring raises.
- Paired reviews: your developers review and ship real changes alongside the rescue team before it steps back.
- Tests in your repository, gated in CI: the safety net stays after the people leave.
- Ownership: every account in company hands, and full IP assignment for any code written during the rescue.
How does RAITHub's Code Rescue engagement work?
It starts with a 2-week diagnostic, runs 2 to 4 weeks with a fixed audit and a fixed scope, and works on any stack. The deliverable is an audit memo and a plan, and the default is to keep what works, not to rewrite everything.
- 2-week diagnostic: codebase audit, infrastructure audit and risk register.
- Stabilise: checkout, auth, payments and data first.
- Replace what is costing you, keep what works.
- Knowledge transfer and paired reviews, with an optional ongoing engagement if you want one.
- Stack-agnostic: we work in the stack you have rather than moving you to ours.
- Always included: founder-reviewed architecture, a real test suite with CI gates, full IP assignment, an NDA before detailed discussion, weekly demos with working software every Friday, and runbooks at handover.
Why this matters for budget: a fixed audit and scope means you know what the work covers before you commit, and a keep-what-works default means you do not pay to rebuild parts that were never the problem.
The same gates run on the platforms we built. Sundor Skin, deployed in September 2026, has 530+ automated tests plus 20 SQL assertion suites, and its CI replays all 76 migrations on an empty database and fails if any buyer-scoped table lacks a row-level-security policy. TheSkinProof, the founder's own marketplace, built and run by RAITHub and live since Q2 2025, has 750+ automated tests with enforced coverage thresholds. These were built, not rescued, but they show the standard a rescued codebase is brought towards.
See services and pricing for engagement details, what RAITHub is for background, or book a Code Rescue audit call (free, 15 minutes).
Last reviewed: 27 September 2026
To hand the takeover to RAITHub instead, see the code rescue service.
Frequently asked questions
My developer left and I cannot access the code. What now?
Start with account ownership: repository, hosting, domain, database and payment provider. Check your contract for IP assignment, request access in writing, and back up production data as soon as you can reach it. Do not let anyone start changing code until access is secured.
How long does a software project rescue take?
A RAITHub Code Rescue runs 2 to 4 weeks, starting with a 2-week diagnostic. The diagnostic produces an audit memo and a plan with a fixed scope, so the rest of the timeline is agreed before the work starts.
Can you fix an abandoned codebase without rewriting it?
Usually, yes. Most abandoned codebases contain working business logic worth keeping. The practical approach is to put tests around critical paths, fix the highest risks first, and replace only the modules that keep costing you.
What does a code audit deliver?
RAITHub's diagnostic delivers an audit memo and a plan, built from a codebase audit, an infrastructure audit and a risk register that ranks each finding by likelihood, impact and cost to fix.
Does RAITHub work with any tech stack?
Yes. Code Rescue is stack-agnostic: the work happens in the stack you already have, and any replacement is decided module by module in the plan.
When is a full rewrite the right choice?
When the data model fundamentally contradicts the business, when the platform no longer receives security fixes and cannot be upgraded, or when the code is a prototype nothing depends on. Even then, rebuild in stages and keep the data.
Who owns the code after a rescue?
You do. RAITHub assigns full IP for all work, signs an NDA before detailed discussion, and hands over runbooks so your team can run the system independently.
How do I start a Code Rescue with RAITHub?
Book the free 15-minute technical audit through the contact page or email hello@raithub.com. Have your list of accounts, your contract and the three problems costing you most ready.
Related posts
Ready to discuss your project?
Book a free 15-minute technical audit with our engineering team.