k3sm licensed apache-2.0 (DCO)

The vm RuntimeClass

Opt a Pod into a Virtualization.framework isolation boundary, check the node capability labels behind it, and see exactly where the Linux-image path stops today.


About 15 minutes. You leave with the manifest that opts a Pod out of the shared native trust domain and into a Virtualization.framework boundary, the node labels that make such a Pod schedulable, and a clear line around what the Linux side of this path can and cannot do yet.

The vm RuntimeClass runs a standard linux/arm64 OCI image inside a Virtualization.framework guest . The path is single-node, and linux/amd64 is not part of this release: it needs in-guest translation, so such an image is refused at pull.

1. Why It Exists#

Every default Pod on a node runs as the same unprivileged _k3sm user, Seatbelt-confined. There is no per-Pod uid isolation, so same-node Pods share one OS trust domain. That works for one operator’s own workloads and fails for anything untrusted. A Pod that names the vm RuntimeClass gets an isolation boundary backed by Virtualization.framework instead of a sandbox profile inside the shared domain.

The same reasoning applies to two other things the native path does not give you. Per-pod IPs are addressing and identity, never network isolation; any same-node process can dial any pod IP, because Seatbelt cannot express per-IP network filters. NetworkPolicy is enforced only on Service-VIP-mediated ingress on the server node, so direct pod-IP traffic bypasses it, and a joined worker enforces none. In all three cases the boundary is the vm RuntimeClass.

2. Opting In#

apiVersion: v1
kind: Pod
metadata:
  name: untrusted-job
  namespace: default
spec:
  runtimeClassName: vm
  nodeSelector:
    kubernetes.io/os: darwin
  tolerations:
    - key: k3sm.io/provider
      operator: Exists
      effect: NoSchedule
  containers:
    - name: app
      image: native
      command: ["/opt/demo/bin/app"]

A Pod without runtimeClassName: vm uses the default native-process runtime and shares the _k3sm trust domain with its neighbours. The RuntimeClass also merges k3sm.io/virtualization: "true" into the Pod’s node selector, so it only lands on a Mac that can host a guest.

A vm Pod also gets its own network stack behind the guest NAT. The native path now gives same-node Pods separate per-IP port spaces for ordinary ports (≥1024). Ports below 1024, workloads that override the DYLD injection, and image: native host binaries (like the example above) still share the host’s one port space. The guest sidesteps all of those.

3. Check the Node Can Host It#

A k3sm node advertises what the host machine is capable of as k3sm.io/* node labels, each stamped from a probe of the host at node start. A label is present with the value "true" or absent, never "false".

k3sm kubectl get nodes -L k3sm.io/virtualization,k3sm.io/rosetta,k3sm.io/rosetta-linux
labelpresent when the host canhonored today
k3sm.io/virtualizationrun a Virtualization.framework guestyes, schedules and boots a linux/arm64 guest
k3sm.io/rosettatranslate darwin/amd64 Mach-O payloads nativelyno, advertised only
k3sm.io/rosetta-linuxtranslate linux/amd64 payloads inside a guestno, never set in this release

Three properties matter before you build selectors on these labels.

  • k3sm.io/virtualization already runs a guest, and neither Rosetta label leads to translated execution. A vm Pod with an untranslated linux/arm64 payload boots today; translation is not wired on either path.
  • k3sm.io/rosetta-linux is a conjunction rather than an implication. It needs both the guest backend and guest Rosetta, because translation happens inside a guest. A Mac with Rosetta 2 and no virtualization capability carries k3sm.io/rosetta and never k3sm.io/rosetta-linux.
  • Rosetta never changes the node’s architecture. kubernetes.io/arch stays arm64 and kubernetes.io/os stays darwin. Translation is an additional capability, advertised only through the k3sm.io/* keys.

The probes run once, at daemon start. If you install Rosetta 2 on a Mac that is already a node, the node keeps reporting the old answer until you restart the daemon:

softwareupdate --install-rosetta --agree-to-license
sudo launchctl kickstart -k system/io.k3sm.server
k3sm kubectl get nodes -L k3sm.io/rosetta,k3sm.io/rosetta-linux

A node that loses a capability keeps advertising it until restart, and Pods that select it are bound to a node that can no longer honour them. Withdraw the claim by hand with k3sm kubectl label node <node> k3sm.io/rosetta- and then restart the daemon.

4. What Runs, and Where It Still Stops#

A Pod that names runtimeClassName: vm with a linux/arm64 OCI image boots and reaches Running. A container restart, which recreates the VM, costs a median 165 ms (p95 171 ms); kernel start to init exec is a median 50 ms, both on an M1 Ultra Mac Studio.

  • The guest kernel and initramfs are fetched from a pinned release and digest-verified on every start, and a mismatch fails closed for vm Pods only.
  • The host Rosetta label is advertised and not honored, and the guest one is never set, because the VM host attaches no Rosetta share to its guests. Where the probe found Rosetta the node is selectable, but image pull still asks only for the node’s native architecture. An amd64-only image is refused at pull time. The Pod lands in ProviderFailed with a ProviderCreateFailed event reading no image manifest matches a runnable platform, never ImagePullBackOff. The node is healthy; the image has no platform this release can run.
  • The Linux guest payload path runs linux/arm64 only. A rootfs assembled from OCI layers is what lets a vm Pod run a Linux image, and only for linux/arm64. An amd64-only image is refused at pull with a message naming both platforms, rather than started and left to crash.
  • Translated execution shares the node’s trust domain. Rosetta is served by a system helper outside the Pod’s sandbox profile, with a node-global ahead-of-time cache. A translated Pod adds no isolation, so untrusted work still needs the vm boundary itself.

For an arm64 Linux image, the vm RuntimeClass is the way to run it. For an amd64-only image, what will not work walks the failure and the options.

Next#