Server

Load balancer

Load balancer is the component that receives requests and decides which origin handles each one. It spreads load and stops sending traffic to origins that fail health checks.

How it is measured

Look at requests per target, the distribution across targets, the healthy-host count, and the 5xx rates it reports separately from your app's. Also check the algorithm: round robin, least connections or hash on IP.

Balancers log two latencies worth separating: the time to connect to the target, and the time the target took to respond. A big gap between balancer-reported time and app-reported time points to queuing.

Worked example

Three origins sit behind a balancer using round robin. One of them is on a slower instance type with 2 vCPU against 4. Each receives a third of requests, about 120 per second, and the small one runs at 92 percent CPU while the others sit at 48.

Switching to least-outstanding-requests shifts load away from the slow node and p95 falls from 1.3 seconds to 410 ms, with no new hardware.

How it differs

A load balancer chooses where each request goes. A reverse proxy accepts requests on behalf of the app and may also cache, terminate TLS or rewrite. A balancer excludes content decisions beyond routing, a reverse proxy may do only one backend. In practice one tool, such as Nginx or an ALB, often plays both roles.

Common errors

Using round robin across servers of different sizes. Not enabling health checks. Terminating TLS at the balancer and forgetting the origin sees plain HTTP. Letting the balancer become a single point of failure. Sticking sessions when sessions could be shared.

In practice

Check the per-target request counts on your balancer for the past day. If they are uneven, change the algorithm or fix the sticky rule, and make sure a failing node actually leaves the pool within a minute.

See also

Health check, Sticky session

Sources

Count this on a real site.

Watch my website