Linux Images
k3sm runs Pods as native Darwin processes under a single _k3sm user, so there is no per-pod uid
isolation, and same-node Pods share one OS trust domain. The vm RuntimeClass is the intended
answer for untrusted or multi-tenant workloads, an isolation boundary backed by
Virtualization.framework.
Single-node. A Pod that sets
runtimeClassName: vmboots alinux/arm64container image in its own micro-VM. It supports boot and restart,kubectl logswith every option (including--tail,-fand--previous),kubectl execwith exit-code propagation,CrashLoopBackOffand restart backoff, PersistentVolumeClaim storage that survives a hard hypervisor kill, per-container CPU and memory accounting, and in-guest networking. The guest leases an address on the node’s NAT segment, resolves cluster DNS, reaches ClusterIP Services, and is itself reachable through its own Service’s ClusterIP on the same node, and at its pod IP on the ports it declares, through a relay on its node. Guest-to-guest reachability is still not a documented boundary, and the relay does not promise one. See Limitations for everything that was measured, including what is still not wired.It ships
linux/arm64only (linux/amd64needs in-guest translation and is held for a later release); it passes against the release build. This path is single-node.
When to Use It#
Use vm when a workload must not share the _k3sm trust domain with its neighbors: untrusted code,
tenant isolation, or anything you would isolate with a strong boundary on Linux. See
Limitations for what is measured and what is not yet wired.
Concepts frames it the same way. The default native path is not a security
boundary between Pods; vm is. The rationale and the trust-domain analysis live in
the privilege model.
Using It#
apiVersion: v1
kind: Pod
metadata:
name: untrusted-job
spec:
runtimeClassName: vm
nodeSelector:
kubernetes.io/os: darwin
tolerations:
- key: k3sm.io/provider
operator: Exists
effect: NoSchedule
containers:
- name: app
image: myapp # must be linux/arm64
The tolerations entry is required on any k3sm Pod, not only a vm one, because every k3sm node
carries the k3sm.io/provider:NoSchedule taint. A Pod that does not tolerate it stays Pending with
untolerated taint rather than running. Admission warns about a missing toleration but only injects
one for DaemonSet Pods. The same stanza appears in Quickstart.
The image must be linux/arm64. An amd64-only image is refused at pull with a message naming the
mismatch. It is never started and left to crash:
no image manifest matches a runnable platform: want [linux/arm64/v8], image provides [linux/amd64]
Pods without runtimeClassName: vm use the default native-process runtime and therefore share the
_k3sm trust domain with other default Pods on the same node.
Interactive Sessions#
A vm Pod answers the interactive kubectl verbs, and the guest gives every container a real
terminal to run them on.
kubectl exec -it#
kubectl exec -it untrusted-job -- /bin/sh
The guest allocates a pseudo-terminal for the session, so the command runs on a terminal. test -t 0
is true, shell job control and line editing work, stty size reports a real window, and resizing
your own terminal delivers SIGWINCH to the process. tty names a device that exists inside the
container, because the terminal is allocated from the container’s own devpts instance rather than
the guest’s. ps, w, and any program that reopens its controlling terminal therefore agree with
each other instead of naming something that is not there.
Exec without a terminal is unchanged. The command runs on pipes, stdout and stderr stay
separate, and the exit code propagates.
kubectl attach#
kubectl attach connects to the process the container is already running instead of starting a
new one. Declare on the container what stdio it should keep:
spec:
runtimeClassName: vm
containers:
- name: app
image: myapp
tty: true # give the container a terminal as its stdio
stdin: true # retain a writable stdin
A container with tty: true is started on its own pseudo-terminal, sized 24x80 until a client
resizes it, held as all three descriptors, and made the session’s controlling terminal. That is the
shape docker run -t gives you, and it has one visible consequence for logs. The terminal’s line
discipline merges stdout and stderr before either leaves the container, so kubectl logs shows
one merged stream for a tty container, exactly as it already does for a tty exec.
A container with stdin: true and no terminal keeps a writable pipe instead. A container that
declares neither is started exactly as before.
Logs#
A vm Pod’s logs are the same files a native Pod’s are. The guest’s output is relayed to the node
and written to /var/log/pods/<namespace>_<pod>_<uid>/<container>/<restartCount>.log in the CRI
format, by the same writer, so there is one log tree on the node and one reader over it. Rotation,
--previous after a restart, the flat /var/log/containers symlinks and a host-level log shipper
all behave identically on both paths. Container logs is the full page.
Then:
kubectl attach -it untrusted-job
- Closing the client detaches without killing anything. It unsubscribes that client and does nothing else. The process is never signalled, its stdin is never closed, and its terminal is never hung up.
- A new attach follows output from the moment it connects, as with any CRI runtime. Use
kubectl logsfor what was written before. - Concurrent attaches are allowed. Each client gets its own copy of the output, and their keystrokes interleave in arrival order, which is left to the people at the keyboards to coordinate.
- Asking for stdin on a container that kept none fails loudly, with a
FailedPreconditionnaming the fix (stdin: true) rather than a silent discard of what you typed.
The two ceilings on this path (stdinOnce, and what a replayed screen can look like) are on
Limitations. So is what the default native runtime does instead.
Every Container Gets a Minimal /dev#
A container’s root filesystem comes out of an OCI image, whose /dev is empty by construction, and
the container is cut off from the guest’s own device tree. A container with nothing there breaks in
ordinary ways. echo x > /dev/null writes an ordinary file that grows forever, /dev/urandom is
missing under every language runtime that seeds from it, and no terminal could exist at all. So each
container is given:
| Path | What it is |
|---|---|
/dev/null, /dev/zero, /dev/full, /dev/random, /dev/urandom, /dev/tty | the OCI runtime-spec default character devices |
/dev/pts, /dev/ptmx | the container’s own devpts instance, and a relative symlink into it |
/dev/shm | a private tmpfs bounded at 64 MiB, the size other runtimes give a container that did not ask for one |
Two things follow from that list.
- The table is the whole of it. The guest’s own
/devis never re-exposed wholesale, and nothing outside the list appears. Enumerating what a container gets, rather than filtering what it does not, is what keeps the node that reaches the Pod’s guest agent out of every container’s reach, since that agent can start processes in, and read the logs of, every container in the Pod. - Your Pod wins a conflict. Anything the default
/devwould place where one of your volumes mounts is omitted rather than stacked underneath. AMemoryemptyDirat/dev/shmtherefore replaces the bounded default outright, which is how you ask for a bigger one. A Pod that mounts over/devitself gets its mount and no privatedevpts.
Trade-Offs#
- Isolation is a genuine boundary, at the cost of VM startup and overhead versus a native process.
- Fidelity on this path is measured; see Limitations for what is covered and what is not yet wired.
- When a Seatbelt SPI symbol-canary trips on the native path, the runtime degrades to
vmor refuse-to-run, never to an unconfined process (see the privilege model).
Node Capability Labels#
A k3sm node advertises what the host machine is capable of as k3sm.io/* node labels, each
stamped from a probe of the host at node start. A label is either present with value "true" or
absent, never "false", and it is removed when the capability goes away:
| Label | Present when the host can… | Gated by | k3sm honors it today? |
|---|---|---|---|
k3sm.io/virtualization | run the vm RuntimeClass (Virtualization.framework) | the vm RuntimeClass’s own nodeSelector | yes, for scheduling |
k3sm.io/rosetta | translate darwin/amd64 Mach-O payloads via host Rosetta 2, natively, with no VM | your Pod’s nodeSelector | not yet; see “advertised, not yet honored” below |
k3sm.io/rosetta-linux | translate linux/amd64 ELF payloads in a Linux guest via Rosetta for Linux | your Pod’s nodeSelector | not yet; see “advertised, not yet honored” below |
Inspect them with:
kubectl get nodes -L k3sm.io/virtualization,k3sm.io/rosetta,k3sm.io/rosetta-linux
Two properties matter before you build selectors on these labels.
k3sm.io/rosetta-linuxis a conjunction. It requires both thevmbackend and guest Rosetta, because Rosetta for Linux translates inside a guest, and with no VM there is nothing to translate in. So a Mac with Rosetta 2 installed but no virtualization capability carriesk3sm.io/rosettaand notk3sm.io/rosetta-linux, and one never implies the other.- Rosetta never changes the node’s architecture.
kubernetes.io/archand the node’s reportedArchitecturestayarm64, the machine’s actual ISA, on a Rosetta-capable node, andkubernetes.io/osstaysdarwineven forrosetta-linux. Translation is an additional capability, advertised only through thek3sm.io/*keys; nothing in the cluster is told the node is amd64 or is Linux.
Rosetta Labels, Advertised Only#
The k3sm.io/rosetta label reports what the host probe found (the probe did find Rosetta 2), and it
does make the node selectable. But k3sm does not consume it when it pulls your image yet.
The pull still asks only for the node’s native architecture (darwin/arm64), so an
amd64-only image is refused at pull time with a no-matching-platform error. That refusal shows
up on the Pod as ProviderFailed with a ProviderCreateFailed event naming the mismatch, so
look at kubectl get pods and kubectl describe pod, not ImagePullBackOff, which this path never
produces. A multi-arch image that includes darwin/arm64 is unaffected, and runs natively as it
always did.
k3sm.io/rosetta-linux is never set on any node today, on any host. The Linux-guest payload path
itself (rootfs, guest image pull, boot) is built and works, as the status note at the top of this
page says. What is missing is narrower. The node’s VM host attaches no Rosetta directory share to
the guests it builds, so nothing inside a guest could translate a linux/amd64 ELF however capable
the host is, and the guest-Rosetta probe is short-circuited to that answer before it ever asks the
framework. Advertising the label without the share would add linux/amd64 to the pull candidate set
for every vm Pod on the node and then fail inside the guest after the pull, so the label stays
off until the share is attached.
Two things must land before either label changes what k3sm will actually run:
k3sm.io/rosetta(host, darwin/amd64): spawning a translated Mach-O inside the Seatbelt sandbox is not wired. Selecting amd64 payloads before that lands would also weaken a kernel-level check k3sm relies on, because an unsigned arm64 binary is killed by the OS while an unsigned x86_64 one is not.k3sm.io/rosetta-linux(guest, linux/amd64): the VM host needs to attach a Rosetta directory share to the guests it builds. The guest path that would carry it (runtimeClassName: vm,linux/arm64) is already built and works; the share is the one piece not yet wired.
So today k3sm.io/rosetta answers “could this host translate?”, not “will k3sm run my amd64
workload here?”, and k3sm.io/rosetta-linux does not answer at all, because it is never set. Until
the paths above land, ship arm64 (or multi-arch) images. If a Pod is stuck ProviderFailed with a
ProviderCreateFailed event naming a platform mismatch, that is this gap, not a broken node.
Translated Execution Shares the Node’s Trust Domain#
Know one property before you plan on translation. Rosetta does not run entirely inside a Pod’s
sandbox. Translation is served by Apple’s oahd helper, a system daemon outside the Pod’s Seatbelt
profile running as its own user (_oahd), and translated code is cached ahead-of-time in a
node-global directory, /private/var/db/oah, shared by everything on the machine. A Pod’s execution
populates that cache but cannot read it back, and the Pod’s Seatbelt profile does not mediate
either the helper or the cache. So a translated Pod stays in the same-node shared trust domain as
every other default Pod. Translation adds no isolation, and for untrusted workloads the answer remains
the vm RuntimeClass above.
Selecting a Rosetta-Capable Node (Keep the os Key)#
Because these are plain capability labels with no RuntimeClass behind them, your Pod selects them
itself, and it must keep kubernetes.io/os: darwin alongside. This is the selector shape to write
when the paths above land. As written today the two labels behave differently at schedule time.
k3sm.io/rosetta: "true" matches a node when Rosetta 2 is installed, so the Pod schedules and then
fails at pull (see the previous section). k3sm.io/rosetta-linux: "true" matches no node,
because the label is never set (see above), so the Pod stays Pending, unscheduled, until a node
advertises it:
apiVersion: v1
kind: Pod
metadata:
name: legacy-amd64-job
spec:
runtimeClassName: vm # REQUIRED for rosetta-linux — translation happens in a guest
nodeSelector:
kubernetes.io/os: darwin # REQUIRED — do not drop this
k3sm.io/rosetta-linux: "true" # the capability you need
containers:
- name: app
image: myapp-linux-amd64
Three things about that manifest:
runtimeClassName: vmis required here. Rosetta for Linux translates inside a Linux guest, so without it the Pod runs on the native host-process path, where alinux/amd64payload has no meaning. ThevmRuntimeClass also mergesk3sm.io/virtualization: "true"into yournodeSelector, which is consistent, becausek3sm.io/rosetta-linuxalready implies a VZ-capable host.- Dropping
kubernetes.io/os: darwinand writing only the capability key fails admission with a422. k3sm enforces a cluster policy that every Pod declare the darwin node selector, which is what keeps Linux-assuming workloads off these nodes. The capability key adds to that selector, it does not replace it. - The host-translation variant selects
k3sm.io/rosetta: "true"instead and carries noruntimeClassName(the native path, no VM), under the same not-yet-honored caveat.
For a workload you want to run today, drop both the capability key and the RuntimeClass and ship an
arm64 (or multi-arch) image on the plain native path, with kubernetes.io/os: darwin in the selector.
Installing Rosetta After the Node Is Up (Restart Required)#
Capability probes run once, at daemon start. If you install Rosetta 2 (or grant virtualization capability) on a Mac that is already serving as a k3sm node, the node keeps reporting the old answer until the daemon restarts:
# 1. install Rosetta 2 (Apple's installer; one-time, per host)
softwareupdate --install-rosetta --agree-to-license
# 2. restart the k3sm daemon so the capability probes re-run
# (io.k3sm.server is the installed control-plane/node LaunchDaemon; use the label
# your role installed — see the Troubleshooting page)
sudo launchctl kickstart -k system/io.k3sm.server
# 3. confirm the label appeared
kubectl get nodes -L k3sm.io/rosetta,k3sm.io/rosetta-linux
Until step 2, a Pod selecting k3sm.io/rosetta stays Pending with no node to bind to. The
reverse direction, a node that loses a capability, has a documented ceiling; see
Limitations.
If the label still does not appear after a restart, k3sm logs the reason it withheld each capability
(the runtimed condition’s reason, e.g. NotInstalled / TranslationFailed / NotSupported /
VMBackendUnavailable) at node bring-up. See Troubleshooting.
Next#
- Limitations puts the no-per-pod-uid-isolation gap in context.
- Concepts covers the trust-domain model.
- Troubleshooting handles a capability label that will not appear.