Skip to content

Latest commit

 

History

History
40 lines (31 loc) · 1.88 KB

File metadata and controls

40 lines (31 loc) · 1.88 KB

Xiaomi MiMo Runbook Hub

Use this page as the stable entry point for Xiaomi MiMo V2.5 Pro work on RTX PRO 6000 Blackwell.

Current Recommendation

Need Page
Run MiMo V2.5 Pro FP4-DFlash on vLLM Xiaomi MiMo V2.5 Pro FP4-DFlash v3
Reproduce the previous vLLM DFlash page MiMo FP4-DFlash v2
Understand the older SGLang/B12X route MiMo V2.5 Pro hub
Reproduce quantization details MiMo quantization
Reproduce SGLang launch details MiMo SGLang running notes

Current Runtime Shape

Area Current guidance
Current model XiaomiMiMo/MiMo-V2.5-Pro-FP4-DFlash
Current runtime vLLM V2 with DFlash
Attention TRITON_ATTN for the documented v3 path
MoE FlashInfer CUTLASS path in the documented launch
Long context Do not set --max-model-len unless the page explicitly says to; the checkpoint advertises its own long context length.

Version Map

Page Status Why keep it
MiMo FP4-DFlash v3 Current Fathomless validation and seq-mask fix notes.
MiMo FP4-DFlash v2 Historical Earlier vLLM DFlash runbook.
MiMo FP4-DFlash Historical Initial MiMo DFlash page.
MiMo V2.5 Pro SGLang/B12X hub Archive Older SGLang and B12X integration path.

Operational Reminders

  • Check for the fast DFlash marker in logs. A diffkv attention marker usually means the wrong path is selected for the current MiMo fast recipe.
  • MiMo has historically been sensitive to attention-mask argument ordering, so smoke-test short decode and acceptance before running a long benchmark.