Skip to content

Cloud

Health Check & Failover

Each health check is O(1) — one quick status read. The value here isn't speed. It's staying available: a system with failover can survive a server dying; one without it goes down along with that server.

The idea, in plain English

A health check is like a manager knocking on an employee's door every few minutes, just to see if they're still there. Failover is what happens if nobody answers: the manager hands the work to the backup person instead, until the first one shows back up. In cloud systems, a small watcher process 'knocks' on a server, called a health check, on a regular schedule. If the main server (the 'primary') doesn't answer, the app switches to a backup server. This is called a 'failover,' and it means users don't notice anything went wrong. When the primary answers again, the app can switch back. This is called 'failing back.'

How it works

  1. 1On every tick, a fixed check-in point that stands in for a moment in time (not a real clock), ask the primary server: are you healthy?
  2. 2If it answers healthy, keep serving traffic from the primary. Nothing changes.
  3. 3If it doesn't answer, it's down. Fail over right away: start serving traffic from the backup server instead.
  4. 4Keep checking the primary in the background. As soon as it's healthy again, fail back: switch traffic back to the primary.

When you'd use it

Use this for any service where downtime is costly — checkout pages, login systems, anything users expect to 'just work.' A single server can crash, lose its network connection, or need a restart at any moment. Health checks catch that fast, and failover means users barely notice it happened.

Common beginner mistakes

  • Checking too rarely. A dead server can then keep getting real traffic for a long time before anyone notices.
  • Failing back to the primary the instant it answers once. Like a circuit breaker's half-open test, one lucky reply doesn't prove it's stable again.
  • Having no real backup at all. Failover only helps if the backup can actually handle the traffic it gets.

Try it — edit and run

Click the code to edit · press ⌘/Ctrl+↵ to run

Editable code. Tab and Shift+Tab indent. Press Escape, then Tab, to move focus out of the editor.

Expected output — hit Run to try it
Tick 1: primary check -> healthy (serving from primary)
Tick 2: primary check -> healthy (serving from primary)
Tick 3: primary check -> down, failing over to backup (serving from backup)
Tick 4: primary check -> down (serving from backup)
Tick 5: primary check -> healthy, failing back to primary (serving from primary)
Tick 6: primary check -> healthy (serving from primary)
Summary: 2 of 6 ticks served by backup (failover working)

Not sure this is the right topic? See the learning paths → or where this leads →