Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
-
Updated
Sep 1, 2026 - C++
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Ollama based Benchmark with detail I/O token per second. Python with Deepseek R1 example.
(SMG) Shepherd Model Gateway official documentation, designed by studio-noiich.
To associate your repository with the tokenspeed topic, visit your repo's landing page and select "manage topics."