Health check
Also called: liveness probe, readiness probe, /healthz.
A small request the platform sends to each copy of your app every few seconds to ask "are you alive?" and "can you take traffic?". A copy that is not ready gets no traffic, and a copy that stops answering is restarted. In Kubernetes these checks are called probes.
Every 10 s the kubelet asks each pod GET /ready (can it take traffic?) and GET /healthz (is it alive?). Here 10 s is shortened to half a second.
In the load balancer: pods 1, 2, 3 · Failed requests: 0 · Restarts: 0
3 pods are ready and share the requests. Freeze pod 3 to see the probes find it.
Say it in a prompt
Add health checks to the orders service on Kubernetes: a readiness probe on GET /ready (200 only after the database pool and caches are warm) and a liveness probe on GET /healthz (200 if the process can answer), both every 10 seconds with a 1-second timeout and failureThreshold 3. /healthz must not call the database, so a database outage does not restart every pod. Vague vs precise prompt
Vague prompt
make Kubernetes restart the app if it breaks Typical resultAdds one liveness probe that calls the database. When the database is slow, every pod fails it at once and Kubernetes restarts all of them, and new pods get traffic before they have started.
Precise prompt
Add a readiness probe on GET /ready (200 only when warm) and a liveness probe on GET /healthz (no database calls), every 10 s, 1 s timeout, failureThreshold 3. Typical resultNew pods get traffic only once they are ready, a hung pod is taken out and restarted after about 30 seconds, and a slow database does not restart the whole fleet.
Seen on
- Kubernetes docs: When a liveness probe fails more times than allowed, the kubelet restarts the container; when a readiness probe fails, the pod's IP is removed from its Services, so it gets no traffic. Probes run every 10 seconds by default.
- Kubernetes docs: A step-by-step task page: configure an HTTP liveness probe on /healthz, and a readiness probe so a pod that is still starting gets no traffic from Services.
You might describe it as
- restart the server when it hangs
- don't send traffic until the app has started
- an endpoint that says the service is up
Not to be confused with
- Circuit breaker
A health check is the platform testing each copy of your app and sending it no traffic while it fails; a circuit breaker is the caller that stops calling a dependency that keeps failing.
- Canary deploy
A health check asks "is this pod alive and ready?"; a canary deploy asks "is the new version as good as the old one for real users?".
- Load balancer
A health check decides which copies of your app may get traffic; a load balancer spreads the requests over the copies that pass.