Statefulsets and stable identity
Give a workload a stable name, its own per-replica volume, and a headless Service, and see exactly which parts of that identity resolve today.
Give it 20 minutes. You come away with a StatefulSet holding a per-replica local-path volume and a headless Service, and you see how much of StatefulSet identity (name, storage, DNS) is live on k3sm, along with the one caveat on the DNS half.
Read storage first. The node pinning explained there is what makes a StatefulSet on k3sm behave the way it does.
1. What Works and What Does Not#
StatefulSet identity has two halves.
- Stable names and stable storage work. The replica is
counter-0, it keeps that name across restarts, and itsvolumeClaimTemplatesclaim follows it. The claim binds to a local-path volume on one Mac, and the Pod is scheduled back to that Mac for as long as the claim exists. - Stable DNS resolves. The cluster resolver answers
counter-0.counter.default.svcwith the replica’s per-endpoint identity A record, and serves SRV records per named port and the matching PTR. In-Pod resolution is wired on the default runtime, with one caveat. The resolver shim cannot load into SIP platform binaries (/bin/sh,/usr/bin/*), so a shell script’s lookups fall back to the host resolver. Ship a compiled binary, the same constraint the volume mount below carries. On--runtime hostprocess, in-Pod cluster DNS is not wired.
The headless Service below is what gives the replica that name. The bare Service name resolves to
the all-backends A set, and the per-replica name to counter-0 itself. The full DNS inventory is
on limitations.
2. The Manifest#
apiVersion: v1
kind: Service
metadata:
name: counter
namespace: default
labels:
app: counter
spec:
clusterIP: None
selector:
app: counter
ports:
- name: http
port: 8080
protocol: TCP
---
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: counter
namespace: default
spec:
serviceName: counter
replicas: 1
selector:
matchLabels:
app: counter
template:
metadata:
labels:
app: counter
spec:
nodeSelector:
kubernetes.io/os: darwin
tolerations:
- key: k3sm.io/provider
operator: Exists
effect: NoSchedule
containers:
- name: counter
image: native
command:
- /opt/counter/bin/counter
- --data-dir=/var/lib/counter
- --listen=0.0.0.0:8080
ports:
- name: http
containerPort: 8080
volumeMounts:
- name: data
mountPath: /var/lib/counter
resources:
requests:
memory: 32Mi
limits:
memory: 128Mi
volumeClaimTemplates:
- metadata:
name: data
spec:
accessModes: ["ReadWriteOnce"]
storageClassName: local-path
resources:
requests:
storage: 1Gi
This is examples/statefulset-local-path.yaml. As in the storage tutorial, command[0] must point
at your own compiled darwin/arm64 binary. A /bin/sh command would run but would not see
/var/lib/counter, because macOS strips the path-rebase shim from SIP platform binaries.
The manifest runs one replica. Each replica would claim its own local-path volume on whichever node the scheduler picked for it, and on a single-node cluster the replicas would additionally contend for the same host port space.
k3sm kubectl apply -f statefulset-local-path.yaml
k3sm kubectl get pods -l app=counter
k3sm kubectl get pvc,pv
3. Inspect the Identity You Have#
k3sm kubectl get pv -o jsonpath='{.items[0].spec.nodeAffinity}'
k3sm kubectl get endpointslices -l kubernetes.io/service-name=counter
The node affinity is the pin. The endpoint slice records the replica’s address, and
counter-0.counter.default.svc resolves to that same address from inside a Pod. Delete the Pod
and watch it come back as counter-0 with the same claim:
k3sm kubectl delete pod counter-0
k3sm kubectl get pods,pvc -l app=counter
4. Clean Up#
k3sm kubectl delete -f statefulset-local-path.yaml
k3sm kubectl get pvc
The claims survive the StatefulSet, and the local-path volumes survive the claims. Remove them by hand when you are done with the data.
Next#
- two macs covers what node pinning means once there is more than one node.
- limitations has the DNS and NetworkPolicy notes that bear on headless traffic.