k3sm licensed apache-2.0 (DCO)

Liveness and readiness

Attach probes to a native Pod and watch readiness gate Service endpoints while liveness restarts the container.


This one runs about 15 minutes and ends with a Pod carrying both probe kinds. A readiness verdict visibly gates its membership in a Service’s endpoints, and a liveness failure restarts the container.

A Virtual Kubelet provider replaces the kubelet, so k3sm runs container probes itself and reproduces the kubelet prober’s semantics. A verdict is committed only after successThreshold consecutive successes or failureThreshold consecutive failures. All four handler kinds work: httpGet, tcpSocket, exec, and grpc.

1. Two Different Questions#

The probes look similar and answer opposite questions. Getting them the wrong way round is the most common self-inflicted outage in any Kubernetes, and k3sm is no exception.

  • Readiness asks whether traffic should go here right now. Failing it removes the Pod from Service endpoints and does nothing else. Make it cheap, with a short period.
  • Liveness asks whether the process is wedged and unrecoverable. Failing it kills and re-execs the container. Make it slower, more tolerant, and dumber than the readiness probe. A liveness probe that checks a downstream dependency converts that dependency’s outage into a restart loop.

2. The Manifest#

apiVersion: v1
kind: Pod
metadata:
  name: probes
  namespace: default
  labels:
    app: probes
spec:
  nodeSelector:
    kubernetes.io/os: darwin
  tolerations:
    - key: k3sm.io/provider
      operator: Exists
      effect: NoSchedule
  containers:
    - name: app
      image: native
      command:
        - /bin/sh
        - -c
        - |
          while :; do
            printf 'HTTP/1.1 200 OK\r\nContent-Length: 3\r\nConnection: close\r\n\r\nok\n' \
              | /usr/bin/nc -l 8080 >/dev/null 2>&1 || sleep 1
          done
      ports:
        - name: http
          containerPort: 8080
      readinessProbe:
        httpGet:
          path: /
          port: http
        initialDelaySeconds: 2
        periodSeconds: 5
        timeoutSeconds: 2
        failureThreshold: 3
      livenessProbe:
        tcpSocket:
          port: http
        initialDelaySeconds: 15
        periodSeconds: 20
        timeoutSeconds: 5
        failureThreshold: 3
      resources:
        requests:
          memory: 32Mi
        limits:
          memory: 128Mi

This is examples/probes.yaml.

k3sm kubectl apply -f probes.yaml
k3sm kubectl get pod probes -o wide
k3sm kubectl describe pod probes | sed -n '/Conditions/,/Events/p'

The Ready condition flips once the readiness probe has committed a success. describe is where the probe history shows up.

3. Watch Readiness Gate Endpoints#

Put a Service in front of it and the readiness verdict becomes visible as endpoint membership:

k3sm kubectl expose pod probes --name probes --port 80 --target-port http
k3sm kubectl get endpointslices -l kubernetes.io/service-name=probes -o yaml

While the Pod is ready, its address is in the slice and marked serving. Break readiness (stop the responder, for instance with k3sm kubectl exec probes -- /usr/bin/pkill nc) and after failureThreshold consecutive failures the verdict commits, the Ready condition drops, and the address stops being a candidate for traffic. The demo responder is a loop, so it comes back and the Pod recovers; watch it with k3sm kubectl get pod probes -w.

4. How Restarts Behave Here#

On the default runtime, restarts behave the way they do upstream. An exited container is restarted in place per its restartPolicy, with the restart count and a CrashLoopBackOff waiting reason shown as upstream shows them. Always restarts on any exit, including a clean exit 0; OnFailure restarts on a non-zero exit or signal. A committed liveness failure restarts the container through the same path, so the probe above recovers a wedged process and an ordinary crash comes back on its own.

The --runtime hostprocess opt-out is the exception. There an exited container is reaped and never respawned, and only a controller replacing the Pod brings the workload back. On that opt-out, put anything you expect to exit under a controller.

5. Clean Up#

k3sm kubectl delete service probes
k3sm kubectl delete pod probes

Next#

  • statefulset covers identity that outlives the process.
  • limitations puts the restartPolicy note in context with the rest.