Server
Cold start
Cold start is the extra time the first request pays when no copy of the function or container is ready. The platform has to fetch code, boot a runtime and run your init before it can answer.
How it is measured
Measure it as the difference between the first request after idle and the median of requests afterwards, ideally using the platform's own init-duration field in logs. Report it in milliseconds and always alongside how long the function had been idle.
Language and package size dominate the number. A 2 MB bundle on a V8 isolate can start in under 10 ms, while a 150 MB container with a JVM and a database handshake can take several seconds.
Worked example
A booking form on a Lambda-backed Node API gets about 30 submissions a day, spread out. Median response is 140 ms, but submissions after a quiet hour take 2.8 seconds because the function loads an ORM and opens a database connection during init.
Trimming the bundle from 38 MB to 6 MB and moving the connection behind a lazy call brought the cold path to 900 ms. Provisioned concurrency of 1 removed it for 11 dollars a month.
How it differs
A cold start is the first slow request when nothing is loaded. A warm start is the later fast request that reuses an instance already in memory. Cold start excludes any reused state. Warm start excludes the boot and init. Both are the same code, and the only difference is what had to happen first.
Common errors
Averaging cold and warm requests into one latency number. Testing right after a deploy and believing the result is typical. Pinging every five minutes to hide the problem instead of shrinking init. Importing whole SDKs when you need one call. Blaming the database for time spent loading modules.
In practice
Pull init duration from your function logs and split p95 into cold and warm this week. Cut the biggest import, defer the database connection until needed, and consider one provisioned instance only for endpoints where a 3 second wait loses a sale.