k3sm licensed apache-2.0 (DCO)

Architecture

How a kubectl apply becomes a sandboxed Darwin process, through the embedded control plane, the Virtual Kubelet node, the native runtime daemon, and the userspace network.


k3sm is one binary. Started as a server it supervises an upstream Kubernetes control plane, joins itself to that control plane as a node, and runs Pods as native macOS processes. This page follows that path from the top.

Shape of a Running Node#

The shape of a running k3sm nodeKubernetes API clients talk to one signed k3sm binary. The binary runs the upstream apiserver, scheduler and controller-manager over kine and SQLite, a Virtual Kubelet with a Darwin provider that registers the Mac as a node, and darwin-net for pod IPs on lo0, the userspace Service proxy, DNS and the wireguard-go mesh to other Macs. The provider talks over a unix-socket gRPC contract to the root k3sm-runtimed daemon, which spawns each Pod as a Seatbelt-confined native Darwin process on an APFS clonefile root. A second Mac joins the mesh as a k3sm agent: the same provider, proxy and mesh, a join client only, with no kubelet and no kube-proxy.kubectl · client-go · any Kubernetes API clientk3sm: one signed binarykube-apiserverkube-schedulerkube-controller-managerupstream Kubernetes, built from source for darwin/arm64kine → SQLite (WAL)the etcd replacementVirtual Kubelet+ the Darwin providerregisters the Mac as a Nodedarwin-netlo0 pod IPs · Service proxyDNS · wireguard-go meshk3sm agentworker Macjoin client + the sameprovider, proxy, and meshno kubelet, no kube-proxyits own pod range, routedgRPC over a unix socketk3sm-runtimed (root)Seatbelt profile · APFS clonefile root · posix_spawn · userspace limitsPoda native Darwin processPoda native Darwin processPoda native Darwin processno container, no VM, no shim: each one shows up in ps
The control plane, the node, and the runtime inside one binary, plus a worker Mac joined to it.

Two pieces of existing code make this feasible. Virtual Kubelet reimplements the kubelet’s API-facing half in portable Go and delegates execution to a provider, so the parts of a kubelet that are pure Kubernetes come for free, and only the parts that are Linux get rebuilt. The control-plane components build from source for darwin/arm64 essentially unchanged, with kine supplying a SQLite datastore where etcd would be.

Control Plane#

The binary runs the apiserver, scheduler, and controller-manager as supervised child processes, pointed at a kine instance speaking the etcd API over a unix socket to a WAL-mode SQLite file. It is upstream Kubernetes, so RBAC, admission, CustomResourceDefinitions, server-side apply, and every workload controller behave as on any other distribution. The control plane runs on one server; a second server cannot join it yet.

The controller-manager’s controller set is narrowed. The node-side controllers that assume a Linux kubelet (attach/detach, cloud node lifecycle) are switched off, because there is no Linux node underneath for them to act on.

Authorization is Node,RBAC with the NodeRestriction admission plugin, provisioned fail-closed at start-up. If the RBAC graph cannot be laid down, bring-up halts rather than running a control plane in an unknown authorization state. The server’s own node runs as system:node:<name> under the Node authorizer, as any kubelet does, so a bug in the node cannot act as cluster admin. Secrets sit unencrypted in the datastore unless you opt in to encryption at rest with sudo k3sm install --secrets-encryption on the install that creates the cluster; the key is then generated on the Mac.

A supervisor HTTP endpoint on the same binary serves the join flow a new node uses to enroll: a bootstrap token exchange, a per-node password, CSR signing for that node’s certificates, and mesh enrollment.

Node#

A Virtual Kubelet with a Darwin provider registers the Mac as a Node and carries a provider taint. The provider rebuilds the kubelet’s user-visible surface on processes: Pod phases and conditions, readiness and liveness probes, graceful stop, and (because they are not inherited from anywhere) kubectl logs, exec, top, and port-forward are all implemented in the provider itself. When memory runs low, the node reports memory pressure and evicts one Pod at a time, ranked as the kubelet ranks them.

Upstream Kubernetes does not accept Pod.spec.os.name: darwin, so k3sm ships a ValidatingAdmissionPolicy instead, and workloads must select kubernetes.io/os=darwin. Combined with the provider taint, that stops a stray Linux manifest from being scheduled onto a Mac and stranding.

Runtime#

k3sm-runtimed is the containerd analog, and is where the Linux-to-Darwin translation happens.

  • Each Pod gets a per-Pod directory whose root is materialized with APFS clonefile (copy-on-write, so nearly free), one generated default-deny Seatbelt profile, one process group, and one Pod IP.
  • Containers run in place at host paths. There is no chroot and no mount namespace, so processes see the system’s own /System, dyld resolves every Apple framework with no plumbing, and the whole arrangement is SIP-safe.
  • Isolation is a generated SBPL profile that grants read access to /System, /usr/lib, the dyld cryptex and the Pod’s own directory, and write access only to the Pod’s APFS data volume, while /Users and other Pods are denied. Networking is opt-in, and all-or-nothing once it is, because macOS accepts no per-address filters in a sandbox profile.
  • Resources are enforced in userspace. A sampler polls proc_pid_rusage roughly once a second, and a memory breach ends in a SIGKILL surfaced as OOMKilled. CPU is QoS and priority, not CFS millicores, so kubectl top and CPU-driven autoscaling are servable, but a CPU limit is not enforceable.
  • The backend is swappable and version-gated. Seatbelt confines every native Pod; a Pod naming the vm RuntimeClass instead runs inside its own micro-VM, linux/arm64 only, behind a build-time check that re-verifies the macOS symbols each release.

Each native container runs under a resident shim that holds its output and exit status, and container logs go to disk in the CRI format. The provider talks to the daemon over a gRPC contract on a unix socket, so the two restart independently, and a restarted node daemon re-attaches to live Pods instead of recreating them.

Networking#

Pod networking is built from loopback aliases, a userspace proxy, and a wireguard mesh. There is no CNI, no veth, no iptables, and no kube-proxy.

  • Pod IPs are lo0 aliases from the shared-address range, a /24 per node. Same-node Pod to Pod traffic is loopback, and XNU preserves the bound source address, so no NAT is needed to keep identity.
  • Services are a userspace proxy that owns each ClusterIP socket, watches Services and EndpointSlices, and load-balances to local Pods or to remote Pods across the mesh. NodePort and LoadBalancer listeners bind the wildcard address, as those types do upstream, so a Pod and a LoadBalancer can collide on a port, with no network namespace to separate them.
  • DNS is a per-node resolver on the cluster DNS address. Pods find it through a DYLD_INSERT_LIBRARIES getaddrinfo shim, because macOS resolves through mDNSResponder and configd and never reads /etc/resolv.conf.
  • Multi-Mac traffic runs over a wireguard-go mesh on a root-created utun, with each peer’s allowed range set to its own Pod range, so traffic is routed rather than translated. Public keys and endpoints are distributed through the join flow and a MeshPeer custom resource; private keys never leave the node they were generated on. Multi-node is experimental and preview-quality until v0.3.

Privilege#

k3sm needs no sudo for daily commands and does not run the whole distribution as root.

One long-running component runs as root. k3sm-netd is a small helper that owns lo0 aliases, the packet-filter anchor, the utun device, and wireguard. Everything else (the control plane, the node, the runtime daemon, the Service proxy, and every Pod) runs as the unprivileged _k3sm user and reaches the helper over a uid-authenticated unix socket carrying a closed, typed RPC. The helper renders every system-command argument itself and re-validates every parameter; it never executes client-supplied text. If you put the data root on its own APFS volume, a second root LaunchDaemon, io.k3sm.datavol, mounts that volume at boot.

Install takes one administrator step; after that, day-to-day use needs no sudo. limitations covers what this leaves open: there is no per-Pod uid isolation, so untrusted workloads belong in the vm RuntimeClass, which gives each Pod a micro-VM boundary.

Packaging#

k3sm installs as LaunchDaemons, set up and removed with launchctl bootstrap / bootout. The one binary supplies each launchd identity: the root networking helper, and the server running as _k3sm (on a worker Mac, the agent takes the server’s place). Both survive boot, so a headless Mac’s cluster comes back after a reboot.

The code-running entitlements that let the runtime execute foreign binaries belong to the Pod-running path only, and the root helper gets none of them, because a root process must not load foreign code. The distribution avoids restricted Network Extension capability and stays on raw utun and packet filter. Releases so far are ad-hoc signed, carry no Developer ID, and are not notarized, so the published checksum shows a download is intact but not who built it.

Repositories#

The code is split along the seams above, in four public repositories:

repositoryrole
k3smthe distribution (control-plane embedding, the Virtual Kubelet node, the CLI, packaging)
runtimedthe native runtime daemon (Seatbelt, APFS, posix_spawn, resource limits)
darwin-netpod networking (lo0 address management, the Service proxy, the wireguard mesh, DNS)
apisthe shared gRPC, CRD, and Go contracts everything else imports

Contributing guides: k3sm, runtimed, darwin-net, and apis.

The design document, with the red-team findings that shaped these decisions, is in the k3sm repository.