vllm-mlx can serve a registry of named models behind one process and one OpenAI-compatible API surface.
This mode is designed for Apple Silicon machines where unified memory is the main constraint:
- models load lazily on first use
- idle models are evicted with an LRU policy under a memory budget
- contention can be configured to wait, fail fast, or preempt active models
/v1/modelsreflects the configured registry instead of a single default model
Use registry-backed serving when you want one server to expose multiple models such as:
- a small low-latency chat model
- a larger reasoning or coding model
- a multimodal model for image or video requests
Keep single-model serving when you want the smallest operational surface and the highest per-model simplicity.
vllm-mlx serve --models-config /etc/vllm-mlx/models.yaml --host 0.0.0.0 --port 8000Use --memory-budget-gb to override manager.memory_budget_gb (or the legacy
manager.memory_budget) for a particular launch. The CLI value takes precedence
over the YAML value and can supply the budget when the YAML field is absent.
You can still use global serve flags such as:
--api-key--rate-limit--timeout--default-temperature--default-top-p--reasoning-parser--enable-auto-tool-choice--tool-call-parser
Do not combine --models-config with:
- a positional model argument
--served-model-name
The registry is a YAML file with two top-level sections:
manager: global budget and contention behaviormodels: named model entries that clients select via the OpenAImodelfield
Example:
manager:
memory_budget_gb: 100
idle_unload_seconds: 300
contention_policy:
strategy: wait_then_preempt
wait_timeout_s: 45
preempt_after_s: 15
models:
- name: fast
path: /Users/david/ai-models/mlx_models/gemma-4-E2B-it-5bit
preload: true
continuous_batching: false
estimated_memory_gb: 4
- name: smart
path: /Users/david/ai-models/mlx_models/Qwen3.5-27B-VLM-MTP-8bit
continuous_batching: true
enable_mtp: true
estimated_memory_gb: 36
- name: vision
path: /Users/david/ai-models/mlx_models/gemma-4-31B-it-6bit
mllm: true
continuous_batching: true
estimated_memory_gb: 44Total resident-model budget for the registry manager.
This budget counts model weights only. It is the number the manager compares against when deciding whether a new model fits or an idle one must be evicted. It does not include, and does not reserve room for:
- KV cache
- activations during prefill and decode
- OS / filesystem cache
- other colocated services
On a 128 GB machine, a practical starting point is often 80-100 GB.
For different host or launch profiles, pass --memory-budget-gb instead of
maintaining duplicate registry files. The override remains a weights-only
budget and does not change the Metal allocation ceiling described below.
Automatically unload a model after it has had no active requests for this many
seconds. Values less than or equal to 0 disable idle unloading.
If this setting is omitted, registry mode inherits
--auto-unload-idle-seconds; that flag defaults to 0. A value in the YAML
file takes precedence over the CLI fallback. Preloaded models are also eligible
for idle unloading after their preload lease is released.
The manager budget and the MLX allocation ceiling are two separate numbers, and
the budget does not derive from the ceiling. The ceiling is installed at engine
start from --gpu-memory-utilization:
allocation_ceiling = gpu_memory_utilization x device_working_set_size
The weights plus the KV cache plus activations all have to fit under that ceiling, while the budget only accounts for the weights. If the budget is set above what is actually allocatable, the manager's arithmetic says N models fit, it keeps them all resident, and MLX hits the ceiling — so you get a hard out-of-memory failure instead of the graceful eviction the budget exists to provide.
The invariant to maintain is:
memory_budget_gb <= gpu_memory_utilization x device_RAM
- KV/activation headroom
- prefix cache actually resident
The server reconciles the two process-wide terms at startup and logs them together with the prefix-cache setting:
Registry memory budget: 68.0 GB of model weights; Metal allocation ceiling
64.0 GB (50% of 128.0 GB, from serve default); prefix-cache maximum
20.0 GB per memory-aware prefix-cache engine (--cache-memory-mb, 2 of 3 entries)
The entry count follows the cache path each model will use. With
--use-paged-cache, text engines use the paged cache and are excluded from this
count. MLLM engines still construct a memory-aware prefix cache, so their
per-engine --cache-memory-mb maximum remains included. For an unresolved
remote model ID, set mllm: true on the registry entry so the startup report
can include its cache without relying on a model-name heuristic.
If an entry enables continuous batching while the global serve default leaves it disabled, the report uses the same scheduler defaults as the engine, including the default 20% memory-aware cache limit.
When the weights budget alone does not fit below the ceiling, startup warns:
WARNING models-config manager.memory_budget_gb (68.0 GB) exceeds the Metal
allocation ceiling (64.0 GB). ...
This is a diagnostic, not a clamp — the server still starts with the budget you configured. It is also a necessary, not sufficient condition: passing the check does not mean you will not run out of memory, because the KV cache, prefix cache and activations all come out of the same ceiling and are workload-dependent. Treat the ceiling as an upper bound and leave real margin below it.
Notes on how the check is computed:
- The Metal limit is installed only by continuous-batching entries — that is the
one path calling
mx.set_memory_limit, and simple-mode entries are not even constructed with agpu_memory_utilization. The check therefore considers only the effective utilization of continuous-batching entries, taking the lowest, since each such load re-installs the process-wide limit. Agpu_memory_utilizationset on a simple-mode entry has no effect on the ceiling and is ignored here. - A registry with no continuous-batching entries gets no attributed ceiling: nothing installs one, so the report says so rather than deriving a figure from a value that is never applied. The serve default likewise only competes when some continuous-batching entry actually inherits it.
- The conflict check compares only the weights budget against the ceiling, because both are process-wide totals and therefore directly comparable.
--cache-memory-mbis not subtracted from the ceiling. It is a per-engine maximum: it is cloned into each resident continuous-batching engine and allocated lazily, and simple-mode entries never receive it at all. Subtracting it once would understate capacity with one resident model and overstate it with several, so it is reported next to the ceiling rather than folded into it. It is reported only when it can actually bind: for continuous-batching entries using the memory-aware prefix cache. Text entries using--use-paged-cacheare excluded, while MLLM entries remain included because they still construct the memory-aware prefix cache.- A separate warning fires when
--cache-memory-mbalone is at or above the ceiling, which is a configuration error in its own right. - On hosts where MLX cannot report a Metal working-set size, the check reports that the budget could not be reconciled and issues no warning.
Controls what happens when a request needs a model that does not currently fit.
Supported strategies:
fail: return capacity failure immediatelywait: wait for capacity to free uppreempt: cancel active requests on other models and evict themwait_then_fail: wait up towait_timeout_s, then failwait_then_preempt: wait up topreempt_after_s, then start preempting, and stop waiting atwait_timeout_s
Recommended defaults:
- shared internal service:
wait_then_preempt - user-facing low-latency API:
wait_then_fail - strict isolation / no interruption:
wait
Required:
name: request-time model id- one of
path,source, ormodel
Optional:
preload: load this model at startupcontinuous_batching: override the global mode for this modelmllm: force multimodal loading when autodetect is not enoughenable_mtp: enable native MTP for this modelprefill_step_sizespecprefillspecprefill_thresholdspecprefill_keep_pctspecprefill_draft_modelstream_intervalgpu_memory_utilizationestimated_memory_gb
For deterministic eviction behavior:
- local models should have real weight files on disk
- non-local model ids should set
estimated_memory_gb
If a registry entry points at a non-local source and no estimated_memory_gb is provided, startup will reject the config. This prevents the manager from making bad eviction decisions from guesswork.
Both sizing paths are weight estimates, not total runtime memory:
- for a local source, the estimate is the summed on-disk size of the entry's
.safetensors/.gguffiles - for a declared model id, the estimate is the operator-supplied
estimated_memory_gb
Neither includes KV cache or activations, so a model's real peak footprint is
larger than the number the manager charges against memory_budget_gb. Size the
budget with that gap in mind — see
Budget vs. the Metal allocation ceiling.
Clients select a registry entry through the normal OpenAI model field:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")
resp = client.chat.completions.create(
model="smart",
messages=[{"role": "user", "content": "Explain speculative decoding."}],
)If the requested model is not registered, the server returns 404 and lists the configured model ids.
curl http://localhost:8000/v1/modelsRegistry-backed responses include the configured model ids and current state such as:
loadedloadingunloadedpreempting
Each model entry also includes last_used_at when it is loaded. For the
effective idle timeout together with all model states, inspect /v1/status:
curl http://localhost:8000/v1/statuscurl http://localhost:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "fast",
"messages": [{"role": "user", "content": "hello"}],
"max_tokens": 32
}'Then repeat with a second model id to verify:
- lazy load works
- the memory budget is enforced
- the selected contention policy behaves as expected
- Start with local-disk model paths, not remote model ids.
- Set
estimated_memory_gbfor every large model, even when local, so your operational budget stays explicit. - Preload only the model that must be instantly available.
- Verify
/v1/modelsbefore exposing the endpoint to shared traffic. - Exercise the configured contention strategy under load before production cutover.
- Bad or missing
estimated_memory_gbon non-local sources: config load failure - Too-small
memory_budget_gb: repeated capacity failures or unnecessary preemption - Too-large
memory_budget_gbrelative to--gpu-memory-utilization: MLX out-of-memory instead of eviction (the startup log warns about this) - Over-aggressive
preemptpolicy: active requests get cancelled during model swaps - Too many
preload: trueentries: startup load storm and immediate budget pressure
Use global defaults for the common case, then override only the model-specific performance knobs that materially differ.
Good candidates for per-model overrides:
continuous_batchingenable_mtpmllmprefill_step_sizestream_interval
Keep these global unless you have a strong reason not to:
- auth
- rate limits
- request timeout
- reasoning parser selection
- tool parser selection
- manager memory budget / contention policy