Troubleshooting
When a k3sm cluster does not behave. Start with k3sm status, then the common failure modes below.
Start with k3sm status#
k3sm status is one screen covering the two LaunchDaemons, the apiserver, the node, the workloads, the
data root, the datastore and your kubeconfig. A one-line verdict sits at the top, and the command that
fixes whatever is down sits at the bottom.
A healthy cluster looks like this:
k3sm running: 1/1 nodes ready v0.1.1 · kube v1.36.2 · macOS 26.1
install ok /Library/k3sm/k3sm · launcher /usr/local/bin/k3sm · 2 LaunchDaemons in /Library/LaunchDaemons
netd running pid 512 · socket /var/lib/k3sm/run/netd.sock
server running pid 840
apiserver ready https://127.0.0.1:6444 readyz ok
node ready 1/1 nodes ready · my-mac kubelet v1.36.2
workloads ok 3 running
data-root ok /var/lib/k3sm (apfs volume k3sm, 29.9G of 100G used, mounted)
datastore ok kine sqlite, wal, user_version 1
kubeconfig ok ~/.kube/config context "k3sm"
Next: k3sm kubectl get pods -A
A control plane that will not start looks like this. The verdict names the cause, and the server row quotes the last line of its log:
k3sm stopped: io.k3sm.server is crash-looping (12 runs, last exit 1) v0.1.1 · kube v1.36.2 · macOS 26.1
install ok /Library/k3sm/k3sm · launcher /usr/local/bin/k3sm · 2 LaunchDaemons in /Library/LaunchDaemons
netd running pid 512 · socket /var/lib/k3sm/run/netd.sock
server crash-loop 12 runs, last exit 1
log: level=ERROR msg="control plane exited" err="bind: address already in use" (×12)
apiserver down https://127.0.0.1:6444: connection refused
node unknown apiserver unreachable
workloads unknown apiserver unreachable
data-root ok /var/lib/k3sm (apfs volume k3sm, 29.9G of 100G used, mounted)
datastore ok kine sqlite, wal, user_version 1
kubeconfig ok ~/.kube/config context "k3sm"
Next: sudo launchctl kickstart -k system/io.k3sm.server
k3sm status logs server
Three more views go deeper when the overview is not enough:
k3sm status daemons # per-daemon detail: run count, last exit, plist and log paths
k3sm status cluster # readyz, node, workloads by phase, datastore, runtime daemon
k3sm status logs server # tail a daemon's log (netd or server; both when omitted)
And two flags for scripts and for watching a restart:
k3sm status -o json # the machine interface; the text screen is rendered from it
k3sm status --wait # poll until the cluster is running, or --timeout elapses
The exit code is the verdict, so a script can branch on it. k3sm status --help lists the codes.
Rows say unknown when your account cannot read something the daemons own: the run directory, the
datastore, the runtime socket. Nothing has failed there; re-run with sudo for those rows.
Logs#
k3sm daemons run as launchd jobs under the io.k3sm.* labels and log to the unified log system:
log show --predicate 'subsystem BEGINSWITH "io.k3sm"' --last 10m
launchctl print system/io.k3sm.server # daemon state (label may differ per role)
Not everything lands there. The server process’s own structured logs (the LoadBalancer controller, the
ingress host, the control-plane supervisor) go to stderr, which launchd routes to
/var/log/k3sm/server.log. The unified-log predicate above shows none of them.
The /var/log/k3sm directory and the daemon log files in it (server.log, netd.log,
datavol.log) are readable by root and by members of the admin group only, because they carry
daemon arguments, cluster endpoints and failure detail. Read them with sudo if your account is not
an administrator of the Mac. k3sm status prints “not readable by this account” in place of a log
tail it cannot read.
k3sm: command not found After Install#
Install links /usr/local/bin/k3sm → /Library/k3sm/k3sm. If the shell cannot find it:
ls -l /usr/local/bin/k3sm # should be a symlink into /Library/k3sm
echo $PATH # should contain /usr/local/bin
/Library/k3sm/k3sm version # always works, the binary itself
The link is there but the shell does not see it. A terminal opened before the install may have cached the old lookup. Run
hash -r, or open a new window./usr/local/binis missing fromPATH. It is the first entry in/etc/paths, so a login shell normally has it; aPATHthat is overwritten (rather than appended to) in a shell profile can drop it. Add it back, or invoke/Library/k3sm/k3smdirectly.A regular file sits at
/usr/local/bin/k3sm. Install refuses to replace anything that is not a symlink, because that file belongs to something else. Move it aside and re-runsudo k3sm install.Install failed with
refusing to link into /usr/local/bin. The directory is not root-owned, or it is group/other-writable, which would let another account swap the launcher out from under you. Fix it and re-run the install:sudo chown root:wheel /usr/local/bin && sudo chmod 755 /usr/local/bin sudo k3sm install
kubeconfig … not found From k3sm kubectl#
k3sm kubectl and k3sm kubeconfig say this when they were pointed at a specific server work
directory that holds no admin kubeconfig, or when they found no credentials at all. Which one it is
depends on how you ran them:
K3SM_WORK_DIRis set. That names one server and is never second-guessed, so if the directory holds no kubeconfig the command stops there. Point it at the right work directory, or unset it.Nothing was found anywhere. The message names both places it looked. If you are root or the
_k3smservice user, that means the control plane has not written its kubeconfig. Check the server is actually up (see below) and read its log.Your own kubeconfig has no
k3smcontext. For an ordinary account that context is the way in, because the installed server’s work directory belongs to the service user and is not readable by you:kubectl config get-contexts # is there a k3sm context? sudo k3sm install # idempotent; re-merges it into ~/.kube/config k3sm kubeconfig --write # with the context in place, refreshes it (--path writes elsewhere)
See kubectl access for the full resolution order.
Control Plane Startup Failure#
- Run
k3sm statusfirst. It says whether the server daemon is running, crash-looping, stopped or not loaded, and it quotes the last line of the server log for the two failure cases. - Confirm install completed:
k3sm versionandsudo k3sm install(idempotent). - Check the datastore is present and not locked. The kine/SQLite DB lives under the server work directory (see Backup & restore).
- Restart the daemon:
launchctl kickstart -k system/io.k3sm.server(label per your role).
The Data Root Is Declared But Not Mounted#
If you keep /var/lib/k3sm on its own APFS volume, declared either in /etc/fstab or by sudo k3sm install --data-volume, that volume must be mounted before the daemons start. When it is not,
writes land on the boot disk in the bare mountpoint and shadow the real volume, which looks exactly
like an empty cluster. k3sm status reports the data root as not-mounted and the verdict as
stopped, because the daemons refuse to run against an unmounted mountpoint rather than build a
second, empty datastore on top of your data.
sudo k3sm datavol mount
sudo launchctl kickstart -k system/io.k3sm.netd
sudo launchctl kickstart -k system/io.k3sm.server
If the volume was never declared with a k3sm record (an /etc/fstab line you wrote yourself, with
no sudo k3sm install --data-volume run since), datavol mount has nothing to work from and you
mount it by hand instead:
diskutil info /var/lib/k3sm # the Device Node line names the volume
sudo rm -r /var/lib/k3sm/run # the shadow holds only the netd socket
sudo diskutil mount -mountPoint /var/lib/k3sm /dev/diskNsM
sudo launchctl kickstart -k system/io.k3sm.netd
sudo launchctl kickstart -k system/io.k3sm.server
sudo k3sm install refuses to run until the volume is mounted, so mount it first and then re-run the
install if you were in the middle of one.
The Data Root Has The Wrong Owner#
The data root itself, /var/lib/k3sm, belongs to root (root:wheel, mode 0755), the same way
k3s keeps its data dir. The directories the daemons write under it (run, server, agent,
pods, storage, the image store) belong to the unprivileged _k3sm service user, and the
node’s mesh keys live in /var/lib/k3sm/keys, which only root can read or change. If the root is
owned by anyone else, k3sm status reports the data root as wrong-owner with the uid it found.
A root owned by _k3sm is the layout older releases left behind. Restarting the netd helper
realigns the ownership:
sudo launchctl kickstart -k system/io.k3sm.netd
sudo k3sm install # the authoritative repair, and idempotent
Stuck or Crash-Looping Pods#
On the default runtime (what every installed cluster runs), restartPolicy is honored. The
container is restarted in place with an upstream-shaped CrashLoopBackOff backoff, and
kubectl get pod shows the restart count climbing. If yours is not restarting, work through these:
- Is it a plain init container? A non-sidecar init container that fails is not re-run in place. That is the one remaining restart gap on the default runtime, and only a controller replacing the Pod unsticks it.
- Did you start the node with
--runtime hostprocess? That explicit rootless-dev opt-out honors norestartPolicyat all. An exited container is reaped once and never respawned. Drop the flag to get the default runtime back. - Is the Pod’s policy
Never, or is it aJobthat has already succeeded? Both are working as specified. - Stuck rather than exiting? Check the Pod was adapted to the native image model, because a raw Linux image cannot run as a Darwin process (see Images).
See Limitations for the precise scope.
A Pod’s Image Will Not Pull#
A container whose image cannot be resolved does not fail the rest of the Pod. The Pod is
created, the containers that can start do start, and the one that cannot sits Waiting with the
reason kubectl get pod prints in its STATUS column. The Pod’s phase stays Pending while any
container waits, and the restart count stays at 0, because a container that never started has not
restarted.
Read the reason first, because each one has a different recovery:
| Reason | What it means | How it recovers |
|---|---|---|
ErrImagePull | The attempt that just failed: the registry pull, the imagePullSecret, the platform match, or the signature gate. | Retried automatically, 10s then 20s, doubling to a 300s ceiling. |
ImagePullBackOff | The same container between attempts. | Nothing to do; the next attempt is already scheduled. |
ErrImageNeverPull | imagePullPolicy: Never and the image is not in this node’s store. Nothing is pulled and no registry is contacted. | Load it (k3sm image load), then wait for the node to re-check: up to two resync intervals, about 20s. |
CreateContainerConfigError | The container’s run spec could not be built: a missing ConfigMap or Secret key, an env reference that does not resolve, an entrypoint the image config does not supply. | Fix the referenced object, then wait for the node to re-check: up to two resync intervals, about 20s. |
InvalidImageName | The reference does not parse. | Only a spec change helps: kubectl set image, or delete and recreate. Parsing the same string again cannot give a different answer. |
kubectl describe pod carries the same story as Events. Pulling opens every attempt and names
the image; Pulled closes a successful one, saying either how long the fetch took or that the image
was already on the machine. On the failing side: Failed on each failed attempt, BackOff when a
retry is scheduled, InspectFailed for an unparseable reference, and ErrImageNeverPull each time
the node re-checks for an image that policy forbids it to fetch. A container whose image is a host
binary gets no Pulling or Pulled at all, because nothing is fetched for it: that is the native
sentinel, or an absolute path on a container that sets neither command nor args. Give such a
container a command and the reference is treated as an image again.
One attempt announces Pulling for every container of the Pod, not just the one that is behind.
The node resolves a Pod’s images in order inside a single start attempt, so a later container’s
Pulled reports a wait that includes the earlier containers’ fetches. That is what the
(… including waiting) figure means here, and it is why the two durations in a Pulled message can
be far apart. The message also omits the image-size clause upstream Kubernetes appends: the runtime
reports no size for a resolution, and a fabricated byte count would be worse than a shorter sentence.
Four things worth knowing before you debug further:
- On the
vmRuntimeClass, one bad image holds the whole Pod. AvmPod is a single guest, so an image that will not resolve fails the guest build rather than leaving one container behind. Every container reportsPending: the container the failure names carries the reason and the message, and the others readContainerCreating, because they never started for a reason of their own. The Pod is kept and the create is retried on the same schedule described above, so nothing goesFailedand no controller churns replacements. A host-process Pod behaves as the paragraph at the top of this section says: its other containers start and keep running. - The message is bounded. The waiting message and the
Failedevent carry a truncated copy of the underlying error, where upstream Kubernetes prints it whole. The full text, including the registry response, is in the node log (k3sm logs, orlog show --predicate 'subsystem == "io.k3sm.node"'). - The schedule is per Pod and per image. Ten Pods referencing one bad image are ten independent retry schedules, exactly as they would be on a kubelet. Restarting the node daemon resets every schedule to the 10s base, so a restart is the fastest way to retry everything at once after you fix a registry.
- Deleting and recreating the Pod resets its backoff. The schedule lives in the running daemon and is keyed to the Pod, so a fresh Pod starts at 10s again rather than inheriting a 300s wait.
DNS From Inside a Pod Does Not Resolve Cluster Names#
On the default runtime this should work. Service A records, headless Services, StatefulSet per-Pod names, SRV and PTR all resolve. Check these in order:
- Is the lookup coming from
/bin/shor another/usr/bintool? macOS stripsDYLD_INSERT_LIBRARIESfrom SIP platform binaries, so k3sm’sgetaddrinfoshim never loads into them and their lookups go to the host resolver. Do the lookup from your compiled workload instead. macOS imposes that limit; it is not a misconfiguration. - Is the Pod’s
dnsPolicyDefaultorNone? Neither selects cluster DNS, so nothing is injected. UseClusterFirst(the default when the field is unset). - Is it an AAAA lookup? k3sm’s CIDRs are IPv4; AAAA is never answered.
- Are you on
--runtime hostprocess? In-pod cluster DNS is not wired there, so every lookup goes to the host resolver. ThevmRuntimeClass does not share this gap; cluster DNS works there.
See Limitations for the full per-path picture.
UDP Service Failures#
Only cluster DNS on :53 uses UDP today; general UDP Services (ClusterIP and NodePort) are deferred.
See Limitations.
LoadBalancer Stuck <pending>#
Do not look in the unified log for this one. The LoadBalancer controller and the ingress host log
through slog to stderr, which the launchd job routes to a file, so the
log show --predicate 'subsystem BEGINSWITH "io.k3sm"' command at the top of this page shows nothing
for them. Read the file instead:
grep svclb /var/log/k3sm/server.log | tail -50
grep ingress /var/log/k3sm/server.log | tail -50
The controller emits one line at start carrying both addresses, which are different:
svclb: loadbalancer controller starting bind=0.0.0.0 advertise=100.64.0.1
Then look for one of these:
svclb: loadbalancer port is RESERVED by a k3sm wildcard listener. The Service declares a port k3sm’s own listeners own: the NodePort range30000-32767, or the kubelet API port10250. No listener is bound and the status stays empty, because taking10250would breakkubectl logs/exec/topon this node. Pick a differentspec.ports[].port. Normally the API rejects such a Service atkubectl applywith a message naming the port; if it did not, the admission policy failed to provision, and the log carries a matchingprovision reserved-loadbalancer-port DENY policyerror.svclb: listener bind failed. Something else already holds that wildcard port. The log line carries the exact diagnostic command; run it:lsof -nP -iTCP:<port> -sTCP:LISTENThe usual culprits are another process on the Mac, a pod (macOS has no network namespaces, so pods share the node’s port space; see Limitations), or a second LoadBalancer Service declaring the same port. k3sm never picks a different port for you.
loadbalancer/ingress status will stay EMPTY: no advertisable node address could be derived. The listeners are bound and serving, but there is no address that would be correct to publish, so nothing is written rather than advertising an unreachable one. Checkkubectl get node -o wideshows a non-loopbackINTERNAL-IP, and that you did not start the server with--network none.ingress bind retries exhausted. The ingress listeners gave up (bounded retry, no port fallback). Free the port and restart the daemon.
Restart the control plane after fixing the conflict:
sudo launchctl kickstart -k system/io.k3sm.server
A <pending> Service is not visible in kubectl describe, because k3sm has no EventRecorder for
this path yet (the event pipeline is planned). The log file is the only place the reason appears.
kubectl top Returns No Metrics#
k3sm ships no metrics-server and has no CPU accounting; install a metrics-server operator if you need the
metrics.k8s.io verb. See kubectl access and Limitations.
kubectl exec -it or attach Into a Linux Pod Is Refused#
Both verbs are served by the Pod’s guest, so the refusal message says which half is missing.
- "…does not advertise the … capability … recreate the pod". The Pod is running a guest image older than the one this node pins, which can only happen under a development override. Delete and recreate the Pod so it boots the pinned image.
FailedPreconditionon attach. The container kept no stdin to write to. Setstdin: true(andtty: trueif you want a terminal) on the container and recreate the Pod; stdio is decided when the container starts, not when you attach.Unimplementedonkubectl attach -i/-t. That Pod is on the default native path, where attach is output-only. Usekubectl exec -it, or thevmRuntimeClass. See Limitations.
A garbled first screen after attaching is not a failure, because attach replays a bounded buffer and
can start mid-escape-sequence. Redraw with Ctrl-L.
A Pod Cannot Pull From the Node-Local Registry#
- Connection refused. The registry is off by default. Start the server with
--registry-port <port>(or usek3sm dev, which enables it), and read the published port back from thelocal-registry-hostingConfigMap inkube-publicrather than remembering it. - A Pod under
runtimeClassName: vmcannot reach it at all. The guest has its own loopback, and the registry listens only on the node’s. That is a documented limit, not a misconfiguration. Use a registry the guest can reach, or load the image into the node’s store directly. - Push denied. Pushes need the per-boot credential, pulls do not. If the control plane runs as
a different user than the one you are pushing from, point push at its state directory with
--work-dir.
See Node-local registry for the full model.
Certificate Nearing Expiry#
Component certificates are re-issued on every control-plane boot, so a restart renews them.
sudo k3sm certificate rotate reports what a restart would re-issue (and both CA pins, which
never change); --restart performs it. Rotation does not revoke anything and does not cover
worker-node certs, so read Certificates before you rely on it.
Rosetta-Selecting Pods Stuck Pending#
The node advertises a capability label only when its start-time probe said yes.
- Check what the node advertises:
kubectl get nodes -L k3sm.io/virtualization,k3sm.io/rosetta,k3sm.io/rosetta-linux. - If the key is missing, the node’s log carries two lines for it: one naming the
conditionand thereason(NotInstalled,TranslationFailed,NotSupported,QueryFailed,VMBackendUnavailable) explaining why the capability was not advertised, and one naming thelabelkey that was therefore left absent. Grep the log for the key itself (k3sm.io/rosetta-linux) and you land on it. k3sm.io/rosetta-linuxneeds both virtualization and guest Rosetta. A node withk3sm.io/rosettabut nok3sm.io/virtualizationwill never carry it.- Installed Rosetta after the node came up? The probes run once at daemon start, so restart it:
sudo launchctl kickstart -k system/io.k3sm.server. - Your Pod must also keep
kubernetes.io/os: darwinin itsnodeSelector; a Pod with only the capability key is rejected with a422. SeevmRuntimeClass. - Scheduled, but stuck
ProviderFailedwith aProviderCreateFailedevent naming a platform mismatch? (kubectl describe podshows it, and it is neverImagePullBackOff.) The Rosetta labels are advertised but not yet honored at pull. Multi-arch selection itself works, and k3sm reads the manifest list and picks a platform, butlinux/amd64is not among the platforms it will accept on either path today. The guest path itself (runtimeClassName: vm,linux/arm64) is built and works; the node’s VM host attaches no Rosetta directory share to its guests, so nothing inside a guest could translate an amd64 payload yet. An amd64-only image is therefore refused at pull rather than started and left to crash. That is a documented gap, not a broken node; seevmRuntimeClass.
Multi-Node Join Fails#
- Re-mint the token (
sudo k3sm token create); tokens expire. - Confirm the agent can reach the server on
6443and the wireguard mesh is up. See Multi-node.
A Node Went Offline and Pods Are Stuck Terminating#
A Mac that goes away while it is running Pods leaves them stuck: kubectl get pod shows them
Terminating indefinitely, because nothing on that Mac can confirm the container is actually gone.
Work through this in order.
Confirm the Mac is genuinely off, not asleep or off-network. Ping its LAN address, and check how long
kubectl get node <name>has shown itNotReady(the age of that condition, not the node). A Mac that is asleep or briefly disconnected heals itself on wake and needs nothing below; do not tell a node that is coming back that it is out of service.If the Mac is coming back, stop after applying the taint below and remove it once the node rejoins. Deleting the node is only for a Mac that is actually leaving the cluster; see Removing a worker for that flow instead.
Apply the out-of-service taint once you are sure the Mac is not coming back on its own:
kubectl taint node <name> node.kubernetes.io/out-of-service=nodeshutdown:NoExecutek3sm statusprints this exact command in the workloads row once a Pod has been stuck terminating on aNotReadynode for more than two minutes, so you rarely have to type it from memory. The taint tells the pod garbage collector it may finish deleting Pods bound to that node. It rechecks every 20 seconds (the pinned v1.36 podgc interval), so a stuck Pod clears within about that long of the taint landing.Remove the taint once the node is back:
kubectl taint node <name> node.kubernetes.io/out-of-service:NoExecute-Force-delete as a last resort, only if a Pod is still stuck after the taint, or you cannot wait for it:
kubectl delete pod <name> --grace-period=0 --forceA force-deleted Pod may still be running on the partitioned Mac. The cluster starts a replacement, and two copies can run at once until the Mac returns and is cleaned up. For a Pod with attached storage the stakes are higher: check the volume’s node binding first, because two writers on one volume corrupt data where two stateless copies merely waste cycles.
Prevent one offline node from freezing a rollout. A DaemonSet’s default
RollingUpdatestrategy waits for every existing Pod to be replaced in place, so a node that never reports back holds up the whole rollout. SetmaxSurge: 1andmaxUnavailable: 0on the DaemonSet’supdateStrategy.rollingUpdateso a new Pod comes up beside the old one instead of waiting for a slot an offline node will never free.
Next#
- FAQ has quick answers.
- Limitations answers whether something is a bug or a documented gap.
- Backup & restore covers recovering the datastore.
- Certificates covers the PKI, rotation, and its limits.