Skip to content

Cloud

Serverless & Cold Starts

Each call is O(1) to classify — one subtraction and a comparison. The real-world cost is the cold-start penalty itself, which can be many times longer than the actual work. That's why keeping functions 'warm' is a common way to speed things up.

The idea, in plain English

Think of a pop-up food stall that packs up and goes home if nobody's ordering for a while. It then has to unfold the tent, light the grill, and set everything up again before it can serve the next customer. That slow reopening is a 'cold start.' Serverless computing works the same way. Your code doesn't run on a server that's always on. It runs in a container that starts up on demand, then shuts down after sitting idle for a while. If a request arrives while the container is already up and running, it's 'warm,' and the request is served right away. If the container had shut down, it's 'cold,' and there's extra delay to start it back up before your code can even run. This lesson uses fixed numbers to stand in for time, not a real clock or timer, so it gives the same result every time.

How it works

  1. 1Track the tick, a fixed point in a sequence that stands in for a moment in time, of the last request that was served.
  2. 2When a new request arrives, check the gap since that last request. If the gap is bigger than the idle timeout, the container has already shut down.
  3. 3If the container is gone, this call is a COLD start. It pays a fixed startup cost first, then does the actual work.
  4. 4If the container is still around, meaning the gap is small, this call is WARM. It skips the startup cost and goes straight to the work.

When you'd use it

Use this for any 'serverless' function, like AWS Lambda or Cloud Functions, that only runs when it's called instead of sitting on all the time. Understanding cold starts matters whenever you need fast, steady responses. The first request after a quiet stretch will always be slower than the ones right after it.

Common beginner mistakes

  • Assuming every call is equally fast. Capacity planning and latency budgets both need to account for the slower, occasional cold ones.
  • Setting the idle timeout so short that a function barely stays warm between real, spaced-out requests.
  • Using a real timer or sleep in an example just to simulate time passing. That makes it slow and unpredictable. Use a fixed number to stand in for time instead.

Try it — edit and run

Click the code to edit · press ⌘/Ctrl+↵ to run

Editable code. Tab and Shift+Tab indent. Press Escape, then Tab, to move focus out of the editor.

Expected output — hit Run to try it
Call at tick 1: COLD start (container was down) -> startup 100 + work 20 = 120ms
Call at tick 2: WARM (container already running) -> work 20 = 20ms
Call at tick 3: WARM (container already running) -> work 20 = 20ms
Call at tick 10: COLD start (container was down) -> startup 100 + work 20 = 120ms
Call at tick 11: WARM (container already running) -> work 20 = 20ms
Call at tick 20: COLD start (container was down) -> startup 100 + work 20 = 120ms
Summary: 3 cold starts, 3 warm calls, total time 420ms

Not sure this is the right topic? See the learning paths → or where this leads →