Skip to content

All

    Repositories list

    • DeepSeek-V4-Flash-Vision 305B EXL3 on 8x16GB: 380K context, 4 streams, DSpark3 — Lna-Lab serving recipe
      Python
      0000Updated Sep 3, 2026Sep 3, 2026
    • 蒸留蔵 — distilled long-term memory for agents: recall by meaning, writing gated by evidence, one kura per agent mode. Ships as a DeepSeek Harness plugin and an MC…
      Python
      MIT License
      54100Updated Sep 3, 2026Sep 3, 2026
    • Run a 177B model (Qwen3.8-Flash-Next) on an 8 GB laptop GPU. Measured: 6.6 GiB VRAM, 47.8 GiB RAM, 34-35 tok/s. The n-gram table stays on disk; the experts run …
      Python
      MIT License
      2900Updated Sep 1, 2026Sep 1, 2026
    • Python
      0000Updated Aug 30, 2026Aug 30, 2026
    • Python
      Apache License 2.0
      1700Updated Aug 10, 2026Aug 10, 2026
    • C
      MIT License
      0000Updated Jun 16, 2026Jun 16, 2026
    • LNAIME

      Public
      Python
      0100Updated Jun 10, 2026Jun 10, 2026
    • Reproducible recipe: serve abliterated Gemma-4-12B (gemma4_unified) at 50-118 tok/s on no-NVLink Blackwell (SM120) via vLLM nightly + ModelOpt FP8/NVFP4 + MTP s…
      Python
      Apache License 2.0
      01400Updated Jun 7, 2026Jun 7, 2026
    • Python
      Other
      0000Updated May 29, 2026May 29, 2026
    • LnaLang4U

      Public
      1M-context DeepSeek-V4-Flash inference on NVIDIA Blackwell using sglang and SSD KV cache offload.
      Python
      1700Updated May 15, 2026May 15, 2026
    • Python
      0100Updated Apr 27, 2026Apr 27, 2026
    • NVFP4 inference on Blackwell GeForce (RTX 5090/5080/5070 Ti/RTX PRO 6000) — SM120 patches for vLLM + FlashInfer + CUTLASS. 175 tok/s on Qwen3.6-35B MoE.
      Python
      Other
      22200Updated Apr 27, 2026Apr 27, 2026
    • Cuda
      0200Updated Apr 27, 2026Apr 27, 2026
    • Lna-Lab production pipeline: GGUF -> modelopt-format NVFP4 + working MTP head for vLLM on RTX PRO 6000 Blackwell (SM120). Stages 2 (NVFP4) and 3 (MTP graft) are…
      Python
      Other
      31000Updated Apr 27, 2026Apr 27, 2026
    • Python
      0000Updated Apr 16, 2026Apr 16, 2026
    • Python
      0100Updated Apr 13, 2026Apr 13, 2026
    • Blackwell-ready TurboQuant KV cache compression for Trinity-Large-Thinking on vLLM.
      Python
      Apache License 2.0
      0000Updated Apr 8, 2026Apr 8, 2026
    • An open-source cross-platform studio for running NVFP4 models locally, featuring a chat interface, OpenAI-compatible API, performance benchmarking, and multilin…
      Python
      Apache License 2.0
      0300Updated Mar 29, 2026Mar 29, 2026
    • Recompose any story into an original world via contextual abstraction.
      Cypher
      MIT License
      0000Updated Feb 7, 2026Feb 7, 2026
    • Game-style operator UI for Gemini CLI | MGS codec-inspired cyberpunk interface with Neo4j conversation graphs
      MIT License
      0000Updated Aug 20, 2025Aug 20, 2025
    • lna-es

      Public
      あらゆるジャンルのテキストをLLMを使いNeo4Jグラフ化して、グラフのみのデータから意味的復元をするシステムのスターター(MCP対応予定)
      11600Updated Aug 20, 2025Aug 20, 2025
    ProTip! When viewing an organization's repositories, you can use the props. filter to filter by custom property.