Evidence-backed AMD Strix Halo local-AI setup and benchmarks: Qwen3.8, Ollama, llama.cpp, Vulkan/ROCm, large GGUFs, and cross-OEM results.
-
Updated
Aug 30, 2026 - Python
Evidence-backed AMD Strix Halo local-AI setup and benchmarks: Qwen3.8, Ollama, llama.cpp, Vulkan/ROCm, large GGUFs, and cross-OEM results.
Performance-tuned llama.cpp for AMD Strix Halo (gfx1151): FA + MoE-prefill fixes with a bundled current Mesa driver. Vulkan and HIP; portable dir, Docker, and distrobox.
vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-stream, vision, tool calling, 256K context, OpenAI-compatible, Docker. Matches DGX Spark FP8+DFlash+MTP at a third of the cost. No CUDA.
Production-ready, reproducible Ansible for DeepSeek V4 Flash on 128 GiB AMD Strix Halo, with two qualified Vulkan/ROCmFPX stacks, matched quality and throughput benchmarks, and 512K context validation.
A turnkey, fully-local AI workstation engineered for the AMD Ryzen AI Max+ 395. LLM inference, voice, document parsing, browser automation, agents — all on-device.
Claude Code skill for AMD Strix Halo (Ryzen AI MAX+ 395) ML setup. Handles PyTorch installation (official wheels don't work with gfx1151), GTT memory config, and environment setup. Enables 30B parameter models.
ROCmFPX llama.cpp fork for Windows 🏆 — native build, headless OpenAI-compatible server & benchmarks. Tested on AMD Strix Halo (gfx1151), runs on other GPUs too.
GNOME top-bar indicator for the AMD Strix Halo power level on the Bosgame M5 (Sixunited AXB35-02, Ryzen AI Max+ 395) — reads the hardware power button's EC register and adds an ACPI platform_profile so GNOME's own Power Mode menu drives the real TDP.
Optimized dual AMD Strix Halo (gfx1151) vLLM MoE inference: TP=2 over USB4, tuned int4 MoE kernel, ~8us interconnect, serialized serving
Talos-O (Omni): A sovereign, embodied agentic organism forged on AMD Strix Halo. Integrating the Chimera Kernel (Linux 7.0), Zero-Copy Introspection, and the Phronesis Engine. Built from First Principles.
Direct EC fan control for AMD Strix Halo (Ryzen AI Max+ 395) mini-PCs — validated on the Bosgame M5. Fixes sustained-load thermal throttling.
Reproducible local-LLM benchmark harness: llama.cpp on AMD Strix Halo (gfx1151, Ryzen AI Max+ 395) and NVIDIA DGX Spark — frozen corpora, quality gates with unit tests, sealed run bundles. Apache-2.0
Run Qwen 3.6-27B AWQ-INT4 models with DFlash speculative decoding on AMD Strix Halo hardware using vLLM for high-throughput inference.
Drop-in recipe for running faster-whisper on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151) with Ubuntu 26.04 + ROCm 7.2.2 — no source build required
Measured inference stack for Strix Halo (Ryzen AI Max+ 395, 128 GB) — llama.cpp serving plus TTS/TTI/TTV media workloads under one memory authority. Not a fork: master plus a short list of patches. Memory budget, defect registry, prefix-cache gateway.
Keeping a 128 GB unified-memory APU fleet from eating itself: the KFD restore-worker thrash case study (51-minute model loads that should take 71 seconds), cgroup v2 budget architecture, and triage procedures for AMD Strix Halo. PolyForm noncommercial.
Known-good local LLM inference on the Framework Desktop (Ryzen AI Max+ 395 / Strix Halo)
Production ROCm llama.cpp build recipe for AMD Strix Halo (gfx1151, Ryzen AI MAX+ 395) — plus seven-model serving runbooks, KV-cache sizing against a unified-memory budget, sanitized systemd units, and a pitfalls doc of the failures that cost real days. PolyForm noncommercial.
Unofficial SoulX-FlashHead ROCm Lite fork for AMD Ryzen AI Max+ 395 / Radeon 8060S (gfx1151)
Docker stack: Ollama v0.21.0 built from source against ROCm 7.2.2 with native gfx1151 (Strix Halo) — serves Gemma 4 up to 256K context on AMD Ryzen AI MAX+ 395 / Radeon 8060S. Includes a 9-layer make validate ladder for the host firmware, ROCm runtime, container, and long-context inference.
To associate your repository with the ryzen-ai-max topic, visit your repo's landing page and select "manage topics."