
Why MVPs Break Under Real Load
Success is the failure mode. The shortcuts that got you to launch are exactly the ones that break when launch works.

Every engineer who inherits a system wants to rebuild it, and the argument is always persuasive: the code is a mess, progress is slow, and a clean start with what we know now would take a few months.
It will not take a few months. More importantly, the system you want to replace is the only honest record of everything your customers turned out to need — including the parts nobody wrote down.
Not because engineers are optimistic, though they are. Because the estimate is made against the visible feature set, and a system that has been in production for years is mostly invisible.
Not "is the code bad". Bad code that works, is rarely touched and sits away from where you need to change things is not a business problem; it is an aesthetic one.
The real question is: does the current architecture prevent the specific thing you need to do next?
| Symptom | Usually means | Usual answer |
|---|---|---|
| It is slow | A missing index or an N+1 query | Fix it — measure first |
| Deploys are terrifying | No tests, no pipeline | Build the pipeline |
| Nobody understands it | No documentation, no tests | Audit and characterise |
| One feature takes a month | Coupling in one area | Refactor that area |
| It cannot scale past a point | Genuine architectural ceiling | Re-architect incrementally |
| The platform is unsupported | A real deadline imposed on you | Plan a staged migration |
Only the last two rows describe architecture. The first four describe neglect, and neglect is much cheaper to fix than to replace. Why is my app slow covers the first row, which is the one most often mistaken for an architecture problem.
Four conditions. One is rarely enough; two or more and the case becomes serious:
Note what is absent: the language, the framework, and the quality of the code. Those are reasons engineers want a rebuild and rarely reasons a business should pay for one.
Replace it piece by piece while it keeps running. Route traffic through a new implementation for one area at a time, leaving the rest untouched — Martin Fowler's strangler fig pattern, and the approach that survives contact with a system nobody fully understands.
Monolith to microservices covers the architectural version of this, including when splitting is not the answer either.
Both options are expensive, and the decision is usually made on feeling. Three things make it a decision instead:
A short, paid assessment first, because the honest answer is often smaller than either side expects. Then scale and re-architecture work done incrementally behind the running product, so you are never holding two systems and a feature freeze at the same time.
We have told clients their system was fine and needed three specific fixes. That is a smaller engagement than a rebuild and it is the right answer often enough that we start by looking. If you are weighing this alongside changing suppliers, when your agency is not working out covers getting an independent view first.
Refactor, unless the architecture prevents the specific thing you need to do next. Slowness, frightening deploys and code nobody understands are neglect rather than architecture, and all three are far cheaper to fix than to replace. A rebuild becomes serious when the platform is end-of-life, when the data model can no longer represent the business, when you are replacing one bounded part rather than everything, or when the product itself is genuinely becoming something different.
Because the estimate covers the visible feature set and a mature system is mostly invisible: undocumented behaviour customers depend on, bugs that became features, integrations whose failure modes have to be re-learned, and running the old system throughout so capacity is halved. The largest cost is usually the feature freeze, where competitors ship while you reproduce what you already had.
Replacing a system piece by piece while it keeps running — routing traffic through a new implementation for one area at a time and leaving the rest untouched. It means value arrives continuously, you can stop part-way and still keep what you have migrated, each piece is estimated against code you have actually read, and the business keeps shipping throughout.
Measure three things first. Where engineering time actually goes over a month — if most of it goes to one subsystem, that subsystem is the project rather than the whole application. Which files change most, since high churn marks where refactoring pays back fastest. And what row counts say about which features are genuinely used, because rebuilding unused features is a common and avoidable expense.