Project challenges / verified progress
Beacon: turn a program into a service

The engineering notebook

Tell health truthfully

What should `/healthz` and `/readyz` actually answer?

Loading statusStage 9 of 10

  • Workspace not ready
  • Agent not ready
Focus25:00
A small focus ritual

0 focus sessions completed. Every fourth session offers a longer break. Start each phase when you are ready.

Study time never unlocks verified lesson progress.

Loading...

Loading verified progress...

Loading GitHub account...
Phase 4 — Process behaviour

Step 01 of 06

Learn the concept

Health endpoints look like tiny routes until they take down a fleet. A readiness check that fails because a dependency hiccups can remove every pod at once. A lie is bad. A dramatic truth can be worse.

TWO HEALTH QUESTIONSService probeplatform asks/healthzshould it liveliveness/readyzshould get trafficreadiness
The endpoints are separate because the platform actions are separate. Restarting a process and removing it from traffic are not the same remedy.
Step 01

The ideas this is made of

Liveness is about restarting

A liveness endpoint should fail only when restarting the process is likely to help: deadlock, unrecoverable internal state, or a wedged main loop. It should not fail because example.com is down. Restarting every monitor because the internet is having a moment is slapstick with YAML.

Readiness is about receiving traffic

Readiness can fail when the service is temporarily unable to serve useful responses, such as before migrations complete or while the database file cannot be opened. The platform may stop sending traffic but keep the process alive. That distinction lets a recovering instance warm up without being killed.

Dependency checks need blast-radius thinking

If every instance marks itself unready when one shared dependency is slow, the load balancer may remove all instances and create a total outage. Sometimes readiness should report degraded state in JSON while still returning 200, especially for dependencies not required to serve the checked route.

Health responses are for machines first

Keep health bodies small and stable. 200 OK with {"status":"ok"} is enough for liveness. Readiness can include named checks such as sqlite and scheduler. Avoid expensive probes; the platform may call these endpoints every few seconds forever.

Two tiny health handlers
package main

import (
	"fmt"
	"net/http"
	"net/http/httptest"
)

func main() {
	mux := http.NewServeMux()
	mux.HandleFunc("GET /healthz", func(w http.ResponseWriter, r *http.Request) { fmt.Fprint(w, "ok") })
	mux.HandleFunc("GET /readyz", func(w http.ResponseWriter, r *http.Request) { fmt.Fprint(w, "ready") })
	w := httptest.NewRecorder()
	mux.ServeHTTP(w, httptest.NewRequest("GET", "/healthz", nil))
	fmt.Println(w.Code, w.Body.String())
}

The handlers are intentionally cheap. Beacon will add local dependency state to readiness, not a fresh internet probe.

Liveness versus readiness

EndpointQuestionFailure action

/healthz

Should restart?

Kill and restart

/readyz

Should receive traffic?

Remove from load balancer

/metrics

How is it behaving?

Observe, not decide

What these are called on the job

  • Liveness — Whether a process should be restarted.

  • Readiness — Whether an instance should receive traffic.

  • Cascading outage — A failure pattern where one dependency problem causes many healthy components to withdraw or fail.