Architecture
How a kubectl apply becomes a sandboxed Darwin process, through the embedded control plane, the Virtual Kubelet node, the native runtime daemon, and the userspace network.
k3sm is one binary. Started as a server it supervises an upstream Kubernetes control plane, joins itself to that control plane as a node, and runs Pods as native macOS processes. This page follows that path from the top.
Shape of a Running Node#
Two pieces of existing code make this feasible. Virtual Kubelet reimplements the kubelet’s
API-facing half in portable Go and delegates execution to a provider, so the parts of a kubelet that
are pure Kubernetes come for free, and only the parts that are Linux get rebuilt. The control-plane
components build from source for darwin/arm64 essentially unchanged, with kine supplying a SQLite
datastore where etcd would be.
Control Plane#
The binary runs the apiserver, scheduler, and controller-manager as supervised child processes, pointed at a kine instance speaking the etcd API over a unix socket to a WAL-mode SQLite file. It is upstream Kubernetes, so RBAC, admission, CustomResourceDefinitions, server-side apply, and every workload controller behave as on any other distribution. The control plane runs on one server; a second server cannot join it yet.
The controller-manager’s controller set is narrowed. The node-side controllers that assume a Linux kubelet (attach/detach, cloud node lifecycle) are switched off, because there is no Linux node underneath for them to act on.
Authorization is Node,RBAC with the NodeRestriction admission plugin, provisioned fail-closed at
start-up. If the RBAC graph cannot be laid down, bring-up halts rather than running a control plane
in an unknown authorization state. The server’s own node runs as system:node:<name> under the Node
authorizer, as any kubelet does, so a bug in the node cannot act as cluster admin. Secrets sit
unencrypted in the datastore unless you opt in to encryption at rest with
sudo k3sm install --secrets-encryption on the install that creates the cluster; the key is then
generated on the Mac.
A supervisor HTTP endpoint on the same binary serves the join flow a new node uses to enroll: a bootstrap token exchange, a per-node password, CSR signing for that node’s certificates, and mesh enrollment.
Node#
A Virtual Kubelet with a Darwin provider registers the Mac as a Node and carries a provider taint.
The provider rebuilds the kubelet’s user-visible surface on processes: Pod phases and
conditions, readiness and liveness probes, graceful stop, and (because they are not inherited from
anywhere) kubectl logs, exec, top, and port-forward are all implemented in the provider
itself. When memory runs low, the node reports memory pressure and evicts one Pod at a time, ranked
as the kubelet ranks them.
Upstream Kubernetes does not accept Pod.spec.os.name: darwin, so k3sm ships a
ValidatingAdmissionPolicy instead, and workloads must select kubernetes.io/os=darwin. Combined
with the provider taint, that stops a stray Linux manifest from being scheduled onto a Mac and
stranding.
Runtime#
k3sm-runtimed is the containerd analog, and is where the Linux-to-Darwin translation happens.
- Each Pod gets a per-Pod directory whose root is materialized with APFS
clonefile(copy-on-write, so nearly free), one generated default-deny Seatbelt profile, one process group, and one Pod IP. - Containers run in place at host paths. There is no chroot and no mount namespace, so processes
see the system’s own
/System,dyldresolves every Apple framework with no plumbing, and the whole arrangement is SIP-safe. - Isolation is a generated SBPL profile that grants read access to
/System,/usr/lib, the dyld cryptex and the Pod’s own directory, and write access only to the Pod’s APFS data volume, while/Usersand other Pods are denied. Networking is opt-in, and all-or-nothing once it is, because macOS accepts no per-address filters in a sandbox profile. - Resources are enforced in userspace. A sampler polls
proc_pid_rusageroughly once a second, and a memory breach ends in a SIGKILL surfaced asOOMKilled. CPU is QoS and priority, not CFS millicores, sokubectl topand CPU-driven autoscaling are servable, but a CPU limit is not enforceable. - The backend is swappable and version-gated. Seatbelt confines every native Pod; a Pod naming the
vmRuntimeClass instead runs inside its own micro-VM,linux/arm64only, behind a build-time check that re-verifies the macOS symbols each release.
Each native container runs under a resident shim that holds its output and exit status, and container logs go to disk in the CRI format. The provider talks to the daemon over a gRPC contract on a unix socket, so the two restart independently, and a restarted node daemon re-attaches to live Pods instead of recreating them.
Networking#
Pod networking is built from loopback aliases, a userspace proxy, and a wireguard mesh. There is no CNI, no veth, no iptables, and no kube-proxy.
- Pod IPs are
lo0aliases from the shared-address range, a/24per node. Same-node Pod to Pod traffic is loopback, and XNU preserves the bound source address, so no NAT is needed to keep identity. - Services are a userspace proxy that owns each ClusterIP socket, watches Services and EndpointSlices, and load-balances to local Pods or to remote Pods across the mesh. NodePort and LoadBalancer listeners bind the wildcard address, as those types do upstream, so a Pod and a LoadBalancer can collide on a port, with no network namespace to separate them.
- DNS is a per-node resolver on the cluster DNS address. Pods find it through a
DYLD_INSERT_LIBRARIESgetaddrinfoshim, because macOS resolves through mDNSResponder andconfigdand never reads/etc/resolv.conf. - Multi-Mac traffic runs over a wireguard-go mesh on a root-created
utun, with each peer’s allowed range set to its own Pod range, so traffic is routed rather than translated. Public keys and endpoints are distributed through the join flow and aMeshPeercustom resource; private keys never leave the node they were generated on. Multi-node is experimental and preview-quality until v0.3.
Privilege#
k3sm needs no sudo for daily commands and does not run the whole distribution as root.
One long-running component runs as root. k3sm-netd is a small helper that owns lo0 aliases, the
packet-filter anchor, the utun device, and wireguard. Everything else (the control plane, the node,
the runtime daemon, the Service proxy, and every Pod) runs as the unprivileged _k3sm user and
reaches the helper over a uid-authenticated unix socket carrying a closed, typed RPC. The helper
renders every system-command argument itself and re-validates every parameter; it never executes
client-supplied text. If you put the data root on its own APFS volume, a second root LaunchDaemon,
io.k3sm.datavol, mounts that volume at boot.
Install takes one administrator step; after that, day-to-day use needs no sudo.
limitations covers what this leaves open: there is no per-Pod uid isolation,
so untrusted workloads belong in the vm RuntimeClass, which gives each Pod a micro-VM boundary.
Packaging#
k3sm installs as LaunchDaemons, set up and removed with launchctl bootstrap / bootout. The one
binary supplies each launchd identity: the root networking helper, and the server running as
_k3sm (on a worker Mac, the agent takes the server’s place). Both survive boot, so a headless
Mac’s cluster comes back after a reboot.
The code-running entitlements that let the runtime execute foreign binaries belong to the
Pod-running path only, and the root helper gets none of them, because a root process must not load
foreign code. The distribution avoids restricted Network Extension capability and stays on raw
utun and packet filter. Releases so far are ad-hoc signed, carry no Developer ID, and are not
notarized, so the published checksum shows a download is intact but not who built it.
Repositories#
The code is split along the seams above, in four public repositories:
| repository | role |
|---|---|
| k3sm | the distribution (control-plane embedding, the Virtual Kubelet node, the CLI, packaging) |
| runtimed | the native runtime daemon (Seatbelt, APFS, posix_spawn, resource limits) |
| darwin-net | pod networking (lo0 address management, the Service proxy, the wireguard mesh, DNS) |
| apis | the shared gRPC, CRD, and Go contracts everything else imports |
Contributing guides: k3sm, runtimed, darwin-net, and apis.
The design document, with the red-team findings that shaped these decisions, is in the k3sm repository.