Assess the system as it is, not as it was documented
Documentation describes an earlier version of the system, and the people who know the current version are busy running it. The first deliverable of a modernisation is an accurate map, and getting it right costs less than any wrong architectural decision that follows.
Use three sources in this order: observe what the system actually does, read what it claims to do, then interview the people who operate it. The differences between those three are where the risk lives.
Look for the parts that change often. High change frequency is a reasonable proxy for business criticality and for areas where nobody holds a clear model, and it tells you where to put characterisation tests first.
- A functional map built from observation rather than from design documents
- Dependency inventory including batch jobs, file drops and manual steps
- Change frequency per module, from version control where it exists
- A written list of the things nobody is currently certain about
Why rewrites run long
A rewrite inherits only the requirements somebody already understood. Everything the current system does that nobody wrote down gets dropped, then discovered in production, then rebuilt under time pressure. That is the mechanism behind most stalled programmes, and it is not a discipline problem.
Rewrites also freeze the product. While the new system is being written, the old one still needs changes, and each one has to be implemented twice or explicitly deferred. The business experiences the programme as a long pause.
None of this argues against replacing code. It argues against replacing everything at once, before the replacement has a contract, a test suite and one live slice.
A rewrite is a bet that the requirements are already fully understood. Test that bet before funding it.
Pin the current behaviour before you change it
You cannot safely refactor code you have not characterised. Write tests that assert what the system does today, including the parts that look like bugs, because downstream consumers may depend on them.
Record-and-replay is often the fastest route. Capture inputs and outputs from production traffic, add the captured pairs as fixtures, and let the suite grow from real usage rather than from imagination. Where you cannot test at a boundary, at least make current behaviour observable so the first change is not blind.
Keep the tests at the boundaries that matter: the module you intend to extract, and the interfaces other teams already call. Coverage percentage is not the objective here.
- Characterisation tests around the module about to be extracted
- Record-and-replay fixtures captured from real traffic
- Explicit tests for the odd behaviours, not only the intended ones
- Contract tests on interfaces other teams already call
Put a facade in front and leave the legacy behind it
The facade is what makes modernisation incremental. It exposes a small, clean interface over a large, awkward system and absorbs the quirks so nothing downstream has to know about them. New capability goes behind the facade, and the legacy implementation shrinks behind it.
Keep the legacy contract unchanged while it is still in use. Improving the old interface at the same time as wrapping it produces two sources of truth and a long debugging exercise.
Shift traffic gradually and make reversal a configuration change. If switching back requires a deployment, that step was not reversible and the plan was not ready.
- Expose only what callers need, not everything the legacy system supports
- Translate legacy quirks into the new model at the boundary
- Traffic shifted per capability, with instant reversal
- Log every call so remaining legacy usage stays visible
routes:
- path: /api/v1/claims/{id}
upstream: legacy.claims
mode: proxy
- path: /api/v1/claims/{id}/assessment
upstream: modern.assessment
mode: shadow
Check which callers still hit the legacy path every month. That list is the migration plan.
Move the data in stages, never in one event
Data is where modernisation programmes usually break, because the migration is treated as the final step rather than the longest one. Treat it as a series of states with reconciliation at each transition.
Read from the new store while the old one remains authoritative. Then introduce write-through or change capture so both stay in step, verify continuously, and only then move ownership. Every stage should be safe to stop, because the previous stage is still complete.
Write migration code that is safe to run twice. A migration that assumes a clean run is one that will eventually be the reason for an incident.
- New-store reads shadowed against legacy before anything depends on them
- Write-through or change capture keeping both stores consistent
- Scheduled reconciliation with a stated tolerance
- Idempotent migrations that can be re-run safely
- Archive and retention decided before the old copy is dropped
If reconciliation can only run once, it is not a check. It is an event somebody will skip.
Decommissioning is a phase, not a cleanup task
The migration is not finished when the new path serves all traffic. It is finished when the old system is switched off, and that requires proving nothing depends on it, including the batch jobs and the exports nobody remembers.
Track remaining usage explicitly, per caller, per endpoint and per table. It produces a truthful remaining-effort number, which is the thing modernisation business cases usually get wrong.
Two systems running in parallel cost money in licences, infrastructure and attention. Set a target date for the old path from the beginning, even if the date moves.
- Every consumer inventoried, including scripts and manual exports
- Per-endpoint usage tracking to prove the legacy path is unused
- A decommissioning date established at the start of the programme
- Legacy credentials and infrastructure removed, not merely disconnected
If the old system is still reachable six months later, the strangler pattern has become a second system.
Takeaways
- Build the map from observation before planning anything
- Characterise the behaviour you are about to change
- Put a facade in front so every step stays reversible
- Migrate data incrementally with reconciliation at each stage
- The programme is not finished until the legacy path is off