Skip to content

System Design

Heartbeat Failure Detection

O(1) time to record one heartbeat · O(n) time to check the status of n servers.

The idea, in plain English

Imagine calling a friend every few minutes to check in, while they do something risky alone. As long as they keep answering, you assume they're fine. But if a few calls in a row go unanswered, you start to worry. This is exactly how servers watch each other. Every server sends a tiny 'I'm alive' ping — a 'heartbeat' — on a regular schedule. A monitor tracks the last time it heard from each one. If too much time passes with no ping, the monitor marks that server 'down.' It assumes the server has failed.

How it works

  1. 1Every server sends a heartbeat — just its ID and the current time — to a monitor, at a steady interval.
  2. 2The monitor remembers, for each server, the time of the last heartbeat it received.
  3. 3Whenever you check a server's status, compare the current time to its last heartbeat time. If more time has passed than the allowed timeout, mark it 'down.' Otherwise, it's 'up.'

When you'd use it

Use this once your app is popular enough that you run many servers, and you need to automatically notice when one crashes. Then a load balancer can stop sending it traffic, or an alert can fire. This is how real clusters — like Kubernetes watching its nodes — know something has failed, without a human checking by hand.

Common beginner mistakes

  • Setting the timeout so short that a server that's just briefly slow — a network hiccup — gets wrongly marked 'down.' A good timeout allows a little slack.
  • Trusting each server's own clock for timestamps in a real system. Clocks can drift between machines, so production systems must account for that. In these examples, we always pass in a fixed, agreed-upon time, so the result stays predictable.

Try it — edit and run

Click the code to edit · press ⌘/Ctrl+↵ to run

Editable code. Tab and Shift+Tab indent. Press Escape, then Tab, to move focus out of the editor.

Expected output — hit Run to try it
t=5: node-1=up, node-2=up
t=9: node-1=up, node-2=up
t=12: node-1=up, node-2=down
t=20: node-1=down, node-2=down

Not sure this is the right topic? See the learning paths → or where this leads →