Cloud
Multi-Region Failover
Picking a region is O(k) for k regions in the priority list — you just walk the list until you find a healthy one. The real thing being managed here is blast radius: this contains an entire region's outage, instead of letting it take your whole service down.
The idea, in plain English
Imagine a company with three regional warehouses — one in the east, one in the west, one overseas — with a standing rule: always ship from the east warehouse if it's open, otherwise the west one, otherwise the overseas one, otherwise nothing goes out at all. Multi-region failover applies that same rule to cloud regions: separate copies of your infrastructure running in different geographic data centers. Instead of just one primary and one backup, you keep an ordered priority list of regions. Traffic goes to the highest-priority region that's currently healthy. It automatically shifts down the list the moment something above it goes down, then shifts back up once that region recovers.
How it works
- 1Keep an ordered priority list of regions, from most preferred to least. This order is your failover plan, decided ahead of time.
- 2On each check, go through the list in order and pick the first region that's currently healthy. That region serves traffic right now.
- 3If a higher-priority region that was down comes back healthy, traffic moves back up to it automatically. This is called 'failback,' and you don't have to undo anything by hand.
- 4If every region in the list is down at once, there's nowhere left to send traffic. That's a full outage, not just a failover.
When you'd use it
Use this when you run a service that truly cannot go down, even if an entire data center or geographic region has a bad day, such as a power outage, a natural disaster, or a cloud provider's regional failure. A single backup server only survives a server failing. Multi-region failover survives a whole region failing.
Common beginner mistakes
- Putting all regions physically close together, like three data centers in the same city. That defeats the purpose if a single regional event, such as a storm or a power grid failure, can take all of them out together.
- Failing back to a recovering region the instant it answers once. This is the same trap as a circuit breaker's half-open test: one healthy check doesn't prove it's stable and won't flap again.
- Not testing what actually happens when every region is down. 'That'll never happen' is exactly the assumption that turns a bad day into a total outage with no plan.
Try it — edit and run
Click the code to edit · press ⌘/Ctrl+↵ to run
Editable code. Tab and Shift+Tab indent. Press Escape, then Tab, to move focus out of the editor.
Tick 1: all regions healthy -> serving from us-east
Tick 2: down: us-east -> serving from us-west (failover)
Tick 3: down: us-east, us-west -> serving from eu-west (failover)
Tick 4: all regions down -> OUTAGE, no region available
Tick 5: all regions healthy -> serving from us-east (failback)
Summary: 2 failovers, 1 failback, 1 outage tick out of 5 total ticksNot sure this is the right topic? See the learning paths → or where this leads →