Architecture

Rebuild or refactor? How to decide when software is failing

A worker standing on scaffolding erected against a building under renovation
Photo by matreding on Unsplash

Every engineer who inherits a system wants to rebuild it, and the argument is always persuasive: the code is a mess, progress is slow, and a clean start with what we know now would take a few months.

It will not take a few months. More importantly, the system you want to replace is the only honest record of everything your customers turned out to need — including the parts nobody wrote down.

Why rebuilds overrun so reliably

Not because engineers are optimistic, though they are. Because the estimate is made against the visible feature set, and a system that has been in production for years is mostly invisible.

  • The undocumented behaviour customers depend on. A rounding rule, a tolerated edge case, an export format someone's accountant relies on. You find these when they break, which is after launch.
  • The bugs that became features. Every mature system has them, and nobody can list them in advance.
  • The integrations. Each one was a negotiation with someone else's API, their rate limits and their failure modes. Rebuilding means re-learning all of it.
  • Two systems at once. The old one still needs running, patching and supporting throughout, so your capacity is halved exactly when you need it most.
  • Feature freeze. Competitors keep shipping while you reproduce what you already had. This is the cost that is never on the slide and is frequently the largest.

The question that actually decides it

Not "is the code bad". Bad code that works, is rarely touched and sits away from where you need to change things is not a business problem; it is an aesthetic one.

The real question is: does the current architecture prevent the specific thing you need to do next?

SymptomUsually meansUsual answer
It is slowA missing index or an N+1 queryFix it — measure first
Deploys are terrifyingNo tests, no pipelineBuild the pipeline
Nobody understands itNo documentation, no testsAudit and characterise
One feature takes a monthCoupling in one areaRefactor that area
It cannot scale past a pointGenuine architectural ceilingRe-architect incrementally
The platform is unsupportedA real deadline imposed on youPlan a staged migration

Only the last two rows describe architecture. The first four describe neglect, and neglect is much cheaper to fix than to replace. Why is my app slow covers the first row, which is the one most often mistaken for an architecture problem.

When a rebuild genuinely is right

Four conditions. One is rarely enough; two or more and the case becomes serious:

  1. The platform is end-of-life. A framework or runtime with no security updates is a deadline somebody else set for you.
  2. The data model cannot represent the business any more. Schema is the expensive thing to change. If the fundamental shape is wrong — single-tenant where you now need multi-tenant, no concept of an entity the business now runs on — that propagates into everything.
  3. You are rebuilding a part, not the whole. Replacing one bounded service is a project. Replacing everything is a bet.
  4. The product is genuinely changing. If what you sell is becoming a different thing, the old system is solving the old problem well. That is a legitimate reason and quite different from disliking the code.

Note what is absent: the language, the framework, and the quality of the code. Those are reasons engineers want a rebuild and rarely reasons a business should pay for one.

The middle path that usually wins

Replace it piece by piece while it keeps running. Route traffic through a new implementation for one area at a time, leaving the rest untouched — Martin Fowler's strangler fig pattern, and the approach that survives contact with a system nobody fully understands.

  • Value arrives continuously rather than at the end of a long freeze.
  • You can stop. Four months in, with three areas migrated and priorities changed, you have three improved areas rather than half a rewrite.
  • Each piece is estimated against something you have now read, so estimates get better as you go instead of being one large guess made at the start.
  • The business keeps shipping, which is usually what the rebuild was supposed to enable.

Monolith to microservices covers the architectural version of this, including when splitting is not the answer either.

Before you decide, measure

Both options are expensive, and the decision is usually made on feeling. Three things make it a decision instead:

  1. Where does engineering time actually go? Roughly, over a month. If 60% goes to one subsystem, that subsystem is the project — not the whole application.
  2. Which parts of the code change most? High-churn files are where requirements were least understood and where refactoring pays back fastest. The repository history tells you this directly.
  3. What does the schema say about usage? Row counts reveal which features are genuinely used. Rebuilding features nobody uses is a surprisingly common expense — inheriting a codebase covers reading a system this way.

What we do

A short, paid assessment first, because the honest answer is often smaller than either side expects. Then scale and re-architecture work done incrementally behind the running product, so you are never holding two systems and a feature freeze at the same time.

We have told clients their system was fine and needed three specific fixes. That is a smaller engagement than a rebuild and it is the right answer often enough that we start by looking. If you are weighing this alongside changing suppliers, when your agency is not working out covers getting an independent view first.

Frequently asked questions

Should we rebuild our software or refactor it?

Refactor, unless the architecture prevents the specific thing you need to do next. Slowness, frightening deploys and code nobody understands are neglect rather than architecture, and all three are far cheaper to fix than to replace. A rebuild becomes serious when the platform is end-of-life, when the data model can no longer represent the business, when you are replacing one bounded part rather than everything, or when the product itself is genuinely becoming something different.

Why do software rebuilds take longer than expected?

Because the estimate covers the visible feature set and a mature system is mostly invisible: undocumented behaviour customers depend on, bugs that became features, integrations whose failure modes have to be re-learned, and running the old system throughout so capacity is halved. The largest cost is usually the feature freeze, where competitors ship while you reproduce what you already had.

What is the strangler fig pattern?

Replacing a system piece by piece while it keeps running — routing traffic through a new implementation for one area at a time and leaving the rest untouched. It means value arrives continuously, you can stop part-way and still keep what you have migrated, each piece is estimated against code you have actually read, and the business keeps shipping throughout.

How do we decide objectively between rebuilding and fixing?

Measure three things first. Where engineering time actually goes over a month — if most of it goes to one subsystem, that subsystem is the project rather than the whole application. Which files change most, since high churn marks where refactoring pays back fastest. And what row counts say about which features are genuinely used, because rebuilding unused features is a common and avoidable expense.

References

  1. StranglerFigApplication — Martin Fowler

Keep reading

An aerial view of a motorway backed up with traffic
Scale

Why MVPs Break Under Real Load

Success is the failure mode. The shortcuts that got you to launch are exactly the ones that break when launch works.

An open filing cabinet drawer packed with index cards
Architecture

Inheriting an undocumented codebase

The instinct is to rewrite it. The first two weeks should be spent making it legible instead — and the database will tell you more than the code does.

Let's put it into production.

Book a 30-minute call — you'll walk away with a scope, a timeline and a fixed price.

Book a call