Infrastructure9 min read
Seven minutes of downtime
Moving a live game off managed AWS onto three servers I own. The rehearsal took weeks; the switch itself took seven minutes. Here is what the rehearsal was actually for.
What we were running
The game ran on EKS with Helm and Traefik in front and Aurora Serverless v2 behind. It worked. It also cost about $546 a month, and behind it sat a backend of 27,378 lines of Python and FastAPI that I was replacing piece by piece anyway.
The first honest question was not how to move it. It was whether anything in the managed stack was doing work that three ordinary machines could not do. The answer was no, and that turned the decision from architecture into arithmetic.
Rehearsing the cutover
A migration is not a clever plan. It is a boring list, executed twice: once as a rehearsal on real data, once for real. The rehearsal is where you find that a step you wrote as one line is actually four, and that the fourth one asks for a password nobody has.
The part I refused to assume was the restore. Nightly archives are encrypted and kept outside the cloud account, which is only useful if the archive turns back into a running game. So the drill was measured from zero: empty machine, archive, running API with matching row counts.
A backup you have never restored is a rumour. Ours came back with matching row counts seventeen minutes from zero, and that number is the only reason the switch itself was boring.
The seven minutes
On 2 August 2026 the switch was four things in order: flip the old side to read-only, run the final delta, move DNS, watch the health check go green. Every step had a rollback that had also been rehearsed, and the whole thing was written down so that a tired person could follow it at the wrong hour.
pg_dump --format=custom "$SOURCE" | pv | pg_restore -d "$TARGET"
psql -f row-counts.sql # 17 min from zero, every count matchedSeven minutes is not a brag. It is what is left when the hard parts have already happened in rehearsal.
What it cost
The monthly bill went from about $546 to roughly $140, and the AWS account was drained and closed rather than left humming with a forgotten resource in a region I never open.
The money is the least interesting part. The change that mattered is that the failure modes became ones I can reason about without a console: a machine is up or it is not, a disk is full or it is not, a service answers or it does not.
What I would do again
Rehearse the restore, not the backup. Write the cutover as a list, not as a plan. Measure the drill from zero rather than from a snapshot that happened to be warm. And put the monitoring somewhere the thing it monitors cannot take down with it.