MLX Serving
Serve a language model on your Mac’s GPU and call it from any client of the common /v1/chat/completions HTTP API. k3sm models the
workload as an MLXModel object. You declare the model and the memory it needs, and k3sm renders
the serving workload, its Services, and its weight cache.
Requirements: Apple Silicon (arm64), macOS 26+, a running k3sm cluster (Quickstart), and network access to fetch model weights the first time.
1. Check the Node Offers a GPU#
k3sm advertises the Mac’s GPU as the extended resource mlx.k3sm.io/gpu, alongside labels describing the
chip. A node that does not advertise it cannot serve a model, and an MLXModel scheduled there stays
Pending.
kubectl get nodes -L mlx.k3sm.io/chip,mlx.k3sm.io/chip-family,mlx.k3sm.io/memory-gb
kubectl get nodes -o jsonpath='{.items[*].status.allocatable.mlx\.k3sm\.io/gpu}{"\n"}'
The count is 1 or 2. The node advertises two slots when its usable GPU memory ceiling holds two
1.5 GiB slots, and one otherwise. The slots share the Mac’s one integrated GPU; they do not isolate it,
so two models on one node contend for the same device.
A free slot is not enough on its own. A second model also has to fit in GPU memory alongside the first,
or the node refuses it and records a FailedGPUFit event on the pod naming what it wanted, what is
already committed, and the ceiling. Lower one model’s memory, or delete the other, and it starts.
A pod that requests the GPU has to declare a container memory limit. One without a limit on every
container is refused on any node that knows its ceiling, because the node cannot budget for it. An
MLXModel always carries one, derived from memory.
2. Apply a Model#
examples/mlxmodel.yaml serves a small pinned model and is the fastest
way to see the path work:
kubectl apply -f examples/mlxmodel.yaml
Three fields in it look optional but are not:
| Field | Why you must set it |
|---|---|
runtime.image | k3sm ships no built-in serving image. Use the published mlx-serve image, ideally by digest. The digest for your release is recorded in the image’s build notes. |
port | Must match the image; mlx-serve listens on 8000. |
cache.storageClassName | k3sm marks no StorageClass as default, so local-path must be named or the cache volume never binds. See Storage. |
memory is the other field to get right. On Apple Silicon the GPU shares system memory, so this
is both the scheduling constraint and the budget the serving engine’s context window is derived from.
More memory buys a longer context; too little is rejected up front with a reason, rather than failing
later on the node.
Pin revision to an exact model revision. Leaving it empty means the repository’s default branch, which
moves, and two replicas started weeks apart would then serve different weights under one object.
Two things follow from how the pin is applied. The serving engine has no option that takes a revision, so k3sm points it at that revision’s directory inside the cache volume instead:
revisionrequirescache. Without a cache volume there is no such directory, and the model is rejected up front with anInvalidSpecreason rather than served from the moving default branch.- A pinned revision is loaded from the cache, not fetched into it. On a volume that has never held this
model, leave
revisionempty for the first start (the engine downloads the default branch), or stage the weights into the volume yourself.
quantization is rejected today, because the engine has no expression for it and serving a different
variant than the one asked for would look like success. Name the quantized repository in model
instead. mlx-community/Qwen3-0.6B-4bit already does.
3. Watch It Become Ready#
The first start downloads the weights, which can take a while on a cold cache. If revision is
pinned, the weights must already be in the volume (above). Read the conditions, not the PHASE
column. PHASE is a one-word summary for humans and loses information:
kubectl get mlxmodel qwen3-06b -w
kubectl describe mlxmodel qwen3-06b | sed -n '/Conditions/,$p'
kubectl wait --for=condition=Ready mlxmodel/qwen3-06b --timeout=30m
The Ready condition’s reason names the state: Pending (no replica running yet), Downloading
(fetching weights), Loading (weights fetched, model loading), Serving (ready), PodFailed (the
replica died), ScaledToZero (you set replicas: 0).
There is no liveness probe on the serving pod, because a first start is an unbounded download and a probe that killed it would restart the download from zero.
4. Call It#
When the model is ready, status.endpoint carries the in-cluster address:
kubectl get mlxmodel qwen3-06b -o jsonpath='{.status.endpoint}{"\n"}'
The Service VIP is reachable from the Mac itself, so you can call it directly:
VIP="$(kubectl get svc qwen3-06b -o jsonpath='{.spec.clusterIP}')"
curl -sS "http://$VIP:8000/v1/chat/completions" \
-H 'Content-Type: application/json' \
-d '{
"model": "mlx-community/Qwen3-0.6B-4bit",
"messages": [{"role": "user", "content": "Name three primary colours."}],
"max_tokens": 64
}'
Any client of that API works. Point its base URL at http://$VIP:8000/v1 and give it any
non-empty API key. Concurrent requests are batched by the server; the number it will batch is derived
from memory.
5. Delete It#
kubectl delete mlxmodel qwen3-06b
That removes the serving workload, both Services, and the cache PVC. The underlying PersistentVolume is retained, so the downloaded weights stay on disk until you remove them by hand (Storage).
Things to Know#
- MLX serving runs on the default native runtime path. Do not set
runtimeClassName: vm. The serving process needs direct access to the Mac’s GPU, and thevmRuntimeClass isolates a workload from exactly that. - Because it runs natively, a served model shares the
_k3smtrust domain with the other native Pods on that Mac. Serve models and code you trust; see Limitations. - Weights never live in the image. They are downloaded on first start into the cache volume, so give
that volume room for the model you are serving. A pinned
revisionis loaded from that volume rather than downloaded into it (see step 2). - Sharded serving is experimental: an alpha, for trusted tenancy only, and not yet run on hardware.
A model may set
spec.distributed(ranks,backend,parallelism), and the operator places the ranks as gang-scheduled rank Pods with per-rank DNS and reports placement and link health in status. Theringbackend places over the existing mesh. Thejacclbackend needs RDMA-capable direct links between Macs, which are not wired into the node yet, so it reportsShardsPlaced=False. Withoutspec.distributed, one model serves from one node. - The Apple Neural Engine is not a serving target. MLX runs on the GPU; Apple publishes no stable API for scheduling ANE work, so k3sm has no ANE path planned.
- Memory accounting covers the serving process group; the context window is pinned from
memoryso the cache cannot grow past the limit mid-generation.