Forty services, one payments ledger, zero acceptable minutes of downtime. That was the brief, and we hit it — not with courage, but with a checklist boring enough to trust at 6am on a Sunday. Here is the cutover discipline I now use on every migration.
Weeks out: prove the rollback first
Nobody rehearses the rollback because everyone assumes forward motion. We built the rollback before the migration path: one command, tested monthly, timed. When the team knows retreat is safe, advance is calm. Our rollback was exercised four times in staging. It was never used in production — that is the point.
- Rollback is a single documented command with a measured runtime.
- Data can flow backward for at least 24 hours after cutover.
- Every stakeholder knows the abort criteria before the window opens.
Days out: dual-write and diff everything
For three weeks, every write hit the old and new systems while a differ job compared reads. We fixed eleven discrepancies before cutover — eleven incidents we never had. Replication lag dashboards stayed on a wall monitor the whole time, with paging thresholds tighter than production.
Cutover morning: the runbook is the boss
The runbook had 47 steps, each with an owner, an expected duration and a verification query. Nobody improvised; deviations went through the incident commander. Total cutover took 34 minutes against a 4-hour window, and support received three tickets — all password resets unrelated to the move.
After: keep the old system warm
We kept the legacy stack in read-only mode for 30 days with automated parity checks, then archived it. Decommissioning is a ceremony: snapshot, export, verify the export, then — and only then — terminate. Future-you, facing the auditor, says thanks.
Downtime is not a technical inevitability. It is a planning deficit with a invoice attached. Plan like the rollback matters and the cutover takes care of itself.
- Migration
- Databases
- Runbooks
- SRE