Step 01 of 06
Learn the concept
Health endpoints look like tiny routes until they take down a fleet. A readiness check that fails because a dependency hiccups can remove every pod at once. A lie is bad. A dramatic truth can be worse.
The ideas this is made of
Liveness is about restarting
A liveness endpoint should fail only when restarting the process is likely to help: deadlock, unrecoverable internal state, or a wedged main loop. It should not fail because example.com is down. Restarting every monitor because the internet is having a moment is slapstick with YAML.
Readiness is about receiving traffic
Readiness can fail when the service is temporarily unable to serve useful responses, such as before migrations complete or while the database file cannot be opened. The platform may stop sending traffic but keep the process alive. That distinction lets a recovering instance warm up without being killed.
Dependency checks need blast-radius thinking
If every instance marks itself unready when one shared dependency is slow, the load balancer may remove all instances and create a total outage. Sometimes readiness should report degraded state in JSON while still returning 200, especially for dependencies not required to serve the checked route.
Health responses are for machines first
Keep health bodies small and stable. 200 OK with {"status":"ok"} is enough for liveness. Readiness can include named checks such as sqlite and scheduler. Avoid expensive probes; the platform may call these endpoints every few seconds forever.
package main
import (
"fmt"
"net/http"
"net/http/httptest"
)
func main() {
mux := http.NewServeMux()
mux.HandleFunc("GET /healthz", func(w http.ResponseWriter, r *http.Request) { fmt.Fprint(w, "ok") })
mux.HandleFunc("GET /readyz", func(w http.ResponseWriter, r *http.Request) { fmt.Fprint(w, "ready") })
w := httptest.NewRecorder()
mux.ServeHTTP(w, httptest.NewRequest("GET", "/healthz", nil))
fmt.Println(w.Code, w.Body.String())
}The handlers are intentionally cheap. Beacon will add local dependency state to readiness, not a fresh internet probe.
Liveness versus readiness
| Endpoint | Question | Failure action |
|---|---|---|
/healthz | Should restart? | Kill and restart |
/readyz | Should receive traffic? | Remove from load balancer |
/metrics | How is it behaving? | Observe, not decide |
What these are called on the job
Liveness — Whether a process should be restarted.
Readiness — Whether an instance should receive traffic.
Cascading outage — A failure pattern where one dependency problem causes many healthy components to withdraw or fail.
