Server
Graceful shutdown
Graceful shutdown is a process stopping in an orderly way: it refuses new work, finishes what it already accepted, and then exits. The alternative is a hard kill that drops requests midway.
How it is measured
Test it by sending SIGTERM to a process under load and counting what happens. The numbers are requests in flight at signal time, how many completed, how many returned errors, and the seconds until exit.
Platforms give a grace period, 30 seconds by default in Kubernetes, before sending SIGKILL. Your shutdown must finish inside that window, so the longest allowed request has to be shorter than the grace period.
Worked example
An order API receives SIGTERM during a rolling deploy while 12 requests are in flight, 2 of them payment captures taking about 3 seconds. The old build stops its listener, waits for all 12 to finish, closes its database pool, and exits after 3.4 seconds.
A previous version exited on SIGTERM at once and dropped both captures, leaving two customers charged without an order row. Adding a 25 second shutdown handler fixed it.
How it differs
Graceful shutdown is what the app does when told to stop. Connection drain is what the load balancer does so no new traffic arrives. Shutdown excludes routing choices, drain excludes the app's internal cleanup. If you stop the process before the balancer stops routing, visitors still get connection errors.
Common errors
Ignoring SIGTERM so the platform has to kill the process. Calling process.exit in a handler. Closing the database pool before in-flight requests finish. Setting a grace period shorter than the longest request. Not stopping background jobs and leaving them half done.
In practice
Send SIGTERM to a staging instance with a load generator running and count failed requests. Add a preStop delay of a few seconds so the balancer notices first, then make the app stop accepting, finish, and close in that order.