Uptime

Timeout

Timeout is the point at which a client gives up waiting. After it, the request counts as failed even if the server would have answered later.

How it is measured

Separate the stages: connect timeout, TLS handshake, time to first byte, and total. A probe with 10 seconds total and 3 seconds connect will report which stage tripped. Log which limit was hit, not just timeout.

Choose from your latency distribution: above p99 of normal, below the point where a user leaves. In chains of calls, each downstream timeout must be shorter than its caller's.

Worked example

A PHP storefront calls a shipping-rate API with the default cURL timeout of 30 seconds, while nginx's fastcgi_read_timeout is 60 and the uptime probe's is 10. When the rate API stalls at 14:02, probes fail at 10 seconds but PHP workers hang for 30, filling a pool of 20 workers in minutes.

Setting the rate call to 3 seconds with a flat-rate fallback keeps the pool free. Failed probes drop from 100 percent to 0 during the next stall.

How it differs

Timeout is the client giving up. HTTP 504 is a gateway reporting that its upstream took too long. A timeout can happen at any layer and may show no status code. A 504 is a response that did arrive.

Common errors

Leaving defaults at 30 to 60 seconds. Setting longer timeouts downstream than upstream. Retrying immediately on timeout. Counting timeouts as success with no log. Setting a probe timeout shorter than real page times. Not distinguishing connect from read.

In practice

List every timeout in the request path of your key page, from browser to database. Make each inner value smaller than its outer one. Add a fallback where a slow third party can block you.

See also

HTTP 504, Probe

Sources

Count this on a real site.

Watch my website