The vm RuntimeClass
Opt a Pod into a Virtualization.framework isolation boundary, check the node capability labels behind it, and see exactly where the Linux-image path stops today.
About 15 minutes. You leave with the manifest that opts a Pod out of the shared native trust domain and into a Virtualization.framework boundary, the node labels that make such a Pod schedulable, and a clear line around what the Linux side of this path can and cannot do yet.
The vm RuntimeClass runs a standard linux/arm64 OCI image inside a Virtualization.framework
guest . The path is single-node, and linux/amd64 is not part of this
release: it needs in-guest translation, so such an image is refused at pull.
1. Why It Exists#
Every default Pod on a node runs as the same unprivileged _k3sm user, Seatbelt-confined. There is
no per-Pod uid isolation, so same-node Pods share one OS trust domain. That works for one operator’s
own workloads and fails for anything untrusted. A Pod that names the vm RuntimeClass gets
an isolation boundary backed by Virtualization.framework instead of a sandbox
profile inside the shared domain.
The same reasoning applies to two other things the native path does not give you.
Per-pod IPs are addressing and identity, never network isolation; any same-node process can dial
any pod IP, because Seatbelt cannot express per-IP network filters. NetworkPolicy is enforced
only on Service-VIP-mediated ingress on the server node, so direct pod-IP traffic bypasses it, and a
joined worker enforces none. In all three
cases the boundary is the vm RuntimeClass.
2. Opting In#
apiVersion: v1
kind: Pod
metadata:
name: untrusted-job
namespace: default
spec:
runtimeClassName: vm
nodeSelector:
kubernetes.io/os: darwin
tolerations:
- key: k3sm.io/provider
operator: Exists
effect: NoSchedule
containers:
- name: app
image: native
command: ["/opt/demo/bin/app"]
A Pod without runtimeClassName: vm uses the default native-process runtime and shares the _k3sm
trust domain with its neighbours. The RuntimeClass also merges k3sm.io/virtualization: "true" into
the Pod’s node selector, so it only lands on a Mac that can host a guest.
A vm Pod also gets its own network stack behind the guest NAT. The native path now gives same-node
Pods separate per-IP port spaces for ordinary ports (≥1024). Ports below 1024, workloads that
override the DYLD injection, and image: native host binaries (like the example above) still share
the host’s one port space. The guest sidesteps all of those.
3. Check the Node Can Host It#
A k3sm node advertises what the host machine is capable of as k3sm.io/* node labels, each stamped
from a probe of the host at node start. A label is present with the value "true" or absent,
never "false".
k3sm kubectl get nodes -L k3sm.io/virtualization,k3sm.io/rosetta,k3sm.io/rosetta-linux
| label | present when the host can | honored today |
|---|---|---|
k3sm.io/virtualization | run a Virtualization.framework guest | yes, schedules and boots a linux/arm64 guest |
k3sm.io/rosetta | translate darwin/amd64 Mach-O payloads natively | no, advertised only |
k3sm.io/rosetta-linux | translate linux/amd64 payloads inside a guest | no, never set in this release |
Three properties matter before you build selectors on these labels.
k3sm.io/virtualizationalready runs a guest, and neither Rosetta label leads to translated execution. AvmPod with an untranslatedlinux/arm64payload boots today; translation is not wired on either path.k3sm.io/rosetta-linuxis a conjunction rather than an implication. It needs both the guest backend and guest Rosetta, because translation happens inside a guest. A Mac with Rosetta 2 and no virtualization capability carriesk3sm.io/rosettaand neverk3sm.io/rosetta-linux.- Rosetta never changes the node’s architecture.
kubernetes.io/archstaysarm64andkubernetes.io/osstaysdarwin. Translation is an additional capability, advertised only through thek3sm.io/*keys.
The probes run once, at daemon start. If you install Rosetta 2 on a Mac that is already a node, the node keeps reporting the old answer until you restart the daemon:
softwareupdate --install-rosetta --agree-to-license
sudo launchctl kickstart -k system/io.k3sm.server
k3sm kubectl get nodes -L k3sm.io/rosetta,k3sm.io/rosetta-linux
A node that loses a capability keeps advertising it until restart, and Pods that select it are
bound to a node that can no longer honour them. Withdraw the
claim by hand with k3sm kubectl label node <node> k3sm.io/rosetta- and then restart the daemon.
4. What Runs, and Where It Still Stops#
A Pod that names runtimeClassName: vm with a linux/arm64 OCI image boots and reaches Running.
A container restart, which recreates the VM, costs a median 165 ms (p95 171 ms); kernel start
to init exec is a median 50 ms, both on an M1 Ultra Mac Studio.
- The guest kernel and initramfs are fetched from a pinned release and digest-verified on every
start, and a mismatch fails closed for
vmPods only. - The host Rosetta label is advertised and not honored, and the guest one is never set, because the
VM host attaches no Rosetta share to its guests. Where the probe found Rosetta the node is
selectable, but image pull still asks only for the node’s native architecture. An amd64-only image
is refused at pull time. The Pod lands in
ProviderFailedwith aProviderCreateFailedevent readingno image manifest matches a runnable platform, neverImagePullBackOff. The node is healthy; the image has no platform this release can run. - The Linux guest payload path runs
linux/arm64only. A rootfs assembled from OCI layers is what lets avmPod run a Linux image, and only forlinux/arm64. Anamd64-only image is refused at pull with a message naming both platforms, rather than started and left to crash. - Translated execution shares the node’s trust domain. Rosetta is served by a system helper outside
the Pod’s sandbox profile, with a node-global ahead-of-time cache. A translated Pod adds no
isolation, so untrusted work still needs the
vmboundary itself.
For an arm64 Linux image, the vm RuntimeClass is the way to run it. For an amd64-only image,
what will not work walks the failure and the options.
Next#
- building an image packages a native binary as an OCI image.
- the vm runtimeclass is the reference page, including the full label semantics.