Server

Saturation

Saturation is the point at which a resource is fully used, so more work does not finish faster and only waits in line. Latency rises sharply past it.

How it is measured

Watch the resource that is likely to fill first: CPU run queue, worker pool busy count, connection pool wait, disk queue depth or memory. The sign is utilization near its limit with a queue that keeps growing.

Plot throughput against load. The curve climbs, flattens, and sometimes drops. The knee where it flattens is where latency starts to explode.

Worked example

A Node API on 2 vCPUs serves 180 requests per second at p95 of 120 ms. At 240 it gives p95 of 340 ms, and at 300 it gives 4.8 seconds with throughput still stuck near 250 per second. The extra 50 requests per second just sit in a queue.

The saturated resource was CPU at 100 percent on both cores. Adding a third instance moved the knee from about 250 to about 370 requests per second.

How it differs

Saturation is a resource being full and work waiting. Backpressure is the response to it, telling callers to slow down. Saturation excludes any control action, it is a state. Backpressure excludes the measurement, it is a policy. Concurrency is how many requests are in flight, which rises when a system saturates.

Common errors

Looking at average CPU and missing queueing in a pool. Testing only at low load. Assuming latency grows linearly with load. Scaling a service while its database is the saturated part. Ignoring the knee and running at 95 percent utilization.

In practice

Find the saturation point for your app with a stepped load test and write it down. Set alerts on queue depth or wait time, and keep normal peak at about 60 to 70 percent of that knee.

See also

Backpressure, Concurrency

Sources

Count this on a real site.

Watch my website