sm121
Here are 23 public repositories matching this topic...
Measured SM121 compatibility recipe for MiniMax H3 FL2VA on one NVIDIA DGX Spark with vLLM-Omni and online FP8.
-
Updated
Aug 4, 2026 - Python
7.67× LoRA / 8.35× Full FT speedup for Qwen3.5 (0.8B–27B) on NVIDIA DGX Spark — wall-clock parity with rented H100. Lossless within BF16. Three-command interactive wizard handles model picker, data validator, training, and merge.
-
Updated
May 19, 2026 - Python
DeepSeek-V4-Flash DSpark speculative decoding on 2x DGX Spark (GB10/sm_121), 1M context — our fp8 recipe + independent reproduction and cross-build benchmarks of the NVFP4-KV build. Honest, apples-to-apples measurements.
-
Updated
Jul 5, 2026 - Python
Production runbook for Qwen3.5-122B hybrid INT4+FP8 on NVIDIA DGX Spark GB10 — optimization stack, PD firmware wedge diagnosis, bench results
-
Updated
Jun 18, 2026
Patches + recipe to deploy festr2/MiMo-V2.5-Pro-NVFP4-MXFP8-attn-TP8 on 8-node DGX Spark sm_121 (Ray + vLLM, TP=8). Fixes the fused-qkv loader bug that mis-slotted Q values as K/V on 7 of 8 ranks.
-
Updated
May 19, 2026 - Python
Qwen3.8-Flash-Next as Cogni-Brain on NVIDIA DGX Spark (GB10): HashK GPU PLE + SGLang NEXTN, 36.8 tok/s code, 100/100 tool-eval, 262K context.
-
Updated
Sep 1, 2026 - Python
Pre-built PyTorch wheels and build scripts for NVIDIA DGX Spark (GB10, sm_121, Blackwell, CUDA 13.0, ARM64)
-
Updated
Jun 25, 2026 - Shell
Empirical kernel scheduling characterization for NVIDIA GB10 (SM121a). Sweeps GEMM tile configurations, classifies PTX instruction paths, captures hardware telemetry
-
Updated
May 10, 2026 - C++
DGX Spark (GB10/SM121) platform support for Meta's KernelAgent — auto-detect, hardware constraints, safe Triton configs
-
Updated
Mar 14, 2026 - Python
Measured vLLM serving for 1–2 NVIDIA DGX Spark (GB10): TP=2 RoCE flagship, validation ledger, stranger-ready Path A/B onboarding.
-
Updated
Sep 2, 2026 - Python
Measured GLM-5.2-Vision NVFP4 reproduction on 8x DGX Spark: 1M context, native SM121, c1/c8 throughput, and fresh 60-minute stability.
-
Updated
Aug 3, 2026 - Python
Reproducible TP=2 deployment, container build, and experiment report for DeepSeek V4 Flash on two DGX Sparks
-
Updated
Aug 21, 2026 - HTML
LLM serving engine for NVIDIA DGX Spark (GB10/sm_121) with Bayesian auto-tuning
-
Updated
Jul 6, 2026 - Cuda
Add this topic to your repo
To associate your repository with the sm121 topic, visit your repo's landing page and select "manage topics."