Server
Warm start
Warm start is a request served by a function instance that is already loaded in memory, so it skips download, boot and init. It is the fast path of a serverless platform.
How it is measured
Separate warm from cold in logs. Most platforms report init duration only for cold invocations, so warm requests are those without it. Compare their durations and look at the share of warm hits.
A warm invocation includes only the handler time. Expect single-digit to tens of milliseconds for light code, against hundreds or seconds for a cold path.
Worked example
A thumbnail function receives 5,000 requests during a morning peak. 4,940 are warm and run in 38 ms, and 60 are cold at 1.4 seconds, a 1.2 percent cold ratio. Overnight, with one request an hour, almost every call is cold.
Because warm instances keep globals, the team caches the image library handle outside the handler. That cut warm duration from 38 ms to 21 ms and kept stale credentials out of reach by refreshing them every 10 minutes.
How it differs
A warm start reuses an instance already in memory. A cold start has to build one first. Warm excludes boot and init cost, and cold excludes any benefit of earlier work. The same function can show both in one hour, depending on idle gaps and how many parallel requests arrive, since each extra concurrent request may need a new instance.
Common errors
Relying on warm state between calls, such as a global variable that holds user data. Benchmarking only warm invocations. Keeping warm with pings and paying for it. Expecting warmth across regions or versions. Assuming instance lifetime is guaranteed.
In practice
Report the warm ratio and the warm p95 separately from cold. Initialise clients outside the handler, keep per-request data inside it, and accept a few cold starts unless a path is truly latency-sensitive.