Server
Readiness probe
Readiness probe is a check that tells the platform whether a running process can accept traffic at this moment. Failing it removes the instance from the load without killing it.
How it is measured
Define the endpoint, period, timeout and thresholds, and decide what 'ready' means: warm caches loaded, migrations done, database reachable. Observe the ready count of pods or targets over time during deploys.
The key number is how long an instance takes from start to ready. Compare it with how long you wait before sending traffic.
Worked example
A Java service loads a 400 MB model on start and takes 70 seconds to be useful. Its readiness probe at /ready returns 503 until the model is loaded, then 200. During a rolling deploy Kubernetes sends no traffic to the new pod until second 71.
Before the probe, the pod got traffic at second 5 and answered 500 to about 1,400 requests across the rollout.
How it differs
A readiness probe answers 'can it serve right now?' A liveness probe answers 'is it hung?' Readiness excludes the restart decision, and an instance can be alive but not ready. Liveness excludes temporary conditions like a warm-up. Failing readiness is routine, failing liveness is a kill.
Common errors
Reporting ready before caches are loaded. Making readiness depend on a shared dependency, so all pods leave at once. Using the same endpoint as liveness. Leaving out the probe and relying on the port being open. Setting the period so long that failures take minutes to notice.
In practice
Make /ready return 200 only when the app is truly able to answer a real request, and test a deploy while watching error counts. Let the probe fail during shutdown so traffic stops before the process does.