Flash weight streaming for MLX: run massive models larger than your RAM on Apple Silicon.
-
Updated
Jun 13, 2026 - Python
Flash weight streaming for MLX: run massive models larger than your RAM on Apple Silicon.
An LLM inference engine written in pure Rust, designed to run large models on hardware that would normally refuse them.
Run large MLX models on Apple Silicon with flash weight streaming, using native precision beyond RAM limits
Run LLMs larger than your RAM — out-of-core local inference on consumer hardware (MoE, GGUF, llama.cpp). NVMe-as-memory with honest telemetry: real tok/s, page faults, disk traffic.
To associate your repository with the weight-streaming topic, visit your repo's landing page and select "manage topics."