k3sm licensed apache-2.0 (DCO)

Serving a model with MLX

The MLXModel resource and the mlx.k3sm.io/gpu extended resource are shipped. Here is what serving a model through Kubernetes looks like on k3sm today, and what still has to be true on your own Mac to try it.


About 10 minutes to read. This page has no manifest of its own to apply; it explains what MLX serving on k3sm is, then sends you to the walkthrough that runs it.

MLX and Apple-GPU support shipped . Apply an MLXModel, watch it reach Ready, and get a completion back through the Service. What follows is the shape of that path and the limits still on it.

Why the Track Exists#

Apple Silicon’s advantage for inference is unified memory. The GPU and the CPU address the same pool, so a model does not get copied across a bus, and a hypervisor boundary either hides that or taxes it. A Pod on k3sm is a native Darwin process, so it is on the right side of that boundary already. What was missing was a way for Kubernetes to know the hardware is there.

How It Works#

Two pieces make up the path, and both are named in limitations:

  • An MLXModel custom resource, in the mlx.k3sm.io API group, describes a model to serve. Applying one renders the ordinary workload objects a serving stack needs, a StatefulSet and a ClusterIP Service in front of it, so what runs underneath is a normal Kubernetes deployment shape.
  • An mlx.k3sm.io/gpu extended resource, advertised by nodes whose hardware qualifies. That advertisement fails closed, so a node that cannot serve never claims it can. A Pod requests it the way any Pod requests an extended resource, mlx.k3sm.io/gpu: 1 under resources.limits, and the upstream scheduler does the placement.

k3sm ships no built-in serving engine, so an MLXModel names its image. The published one, ghcr.io/k3sm-io/mlx-serve, speaks the common /v1/chat/completions HTTP API. kubectl get mlxmodels lists what you have applied, and a Ready condition means the same thing it means for any other workload. The Pod is up and the endpoint is answering.

What Is Still Limited#

Two things have to be true on your Mac:

  • The node advertises the GPU. A node only advertises mlx.k3sm.io/gpu when the probe behind it confirms the hardware, and an MLXModel scheduled on a node without it stays Pending. The first start also needs network access to fetch the weights.
  • You have the root install tier, the same sudo k3sm install step every other tutorial on this site asks for.

A few design choices to know before you write a manifest:

  • A node advertises one GPU slot, mlx.k3sm.io/gpu: 1, or two when its usable GPU memory ceiling holds two 1.5 GiB slots. The slots share the Mac’s one GPU and do not isolate it, so two models on one node contend for the same device. A second model must also fit in GPU memory beside the first, or the node refuses it with a FailedGPUFit event naming the numbers.
  • GPU memory is not a separate accounted resource. spec.Memory becomes a plain memory request and limit, sized from the model’s own weight and context-window formula, enforced the same sampled way as any other Pod’s memory. A second GPU-specific meter would double-count the same bytes.
  • There is no autoscaling. Setting spec.replicas: 0 scales a served model to zero and back by hand.
  • A model can be sharded across Macs through the spec’s distributed field. In this release that is alpha, for trusted tenancy only, and has not been run on hardware.

Try It#

The hands-on walkthrough, including the manifest and the exact commands, lives at MLX quickstart. The checks this site runs against its tutorials do not cover those commands, because they need the root install tier and an Apple GPU.

Next#