{"data":{"slug":"hogeheer499-commits-strix-halo-guide","name":"strix-halo-guide","tagline":"AMD Ryzen AI Halo setup guide for LLM frameworks like Ollama and llama.cpp Vulkan on Radeon hardware","github_url":"https://github.com/hogeheer499-commits/strix-halo-guide","owner":"hogeheer499-commits","repo":"strix-halo-guide","owner_avatar_url":"https://avatars.githubusercontent.com/u/267467744?v=4","primary_language":"Python","stars":336,"forks":23,"topics":["amd","beelink","benchmark","framework-desktop","gfx1151","gguf","llama-cpp","llm","local-ai","local-llm","mini-pc","ollama","qwen3","radeon-8060s","rocm","ryzen-ai-max","ryzen-ai-max-395","strix-halo","unified-memory","vulkan"],"archived":false,"github_pushed_at":"2026-09-19T21:59:03+00:00","maintenance_label":"Very active","stars_delta_30d":69,"url":"https://www.graphcanon.com/tools/hogeheer499-commits-strix-halo-guide","markdown_url":"https://www.graphcanon.com/tools/hogeheer499-commits-strix-halo-guide.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/hogeheer499-commits-strix-halo-guide","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=hogeheer499-commits-strix-halo-guide","description":"Evidence-backed AMD Strix Halo local-AI setup and benchmarks: Qwen3.8, Ollama, llama.cpp, Vulkan/ROCm, large GGUFs, and cross-OEM results.","homepage_url":"https://strixhaloguide.com/","license":"MIT","open_issues":8,"watchers":6,"ai_summary":"Guide to setting up local LLM environments using AMD Ryzen AI MAX+ 395 and Radeon 8060S, including benchmarks and ROCm support.","readme_excerpt":"## Quick Start (6 Steps)\n\nFor those who want to get running as fast as possible:\n\n1. **BIOS:** Set UMA Frame Buffer to 512MB if available; if your BIOS minimum is 2GB, leave it at 2GB. Keep IOMMU enabled/default for laptops, suspend, and NPU use. Disabling it is an optional desktop benchmark profile.\n2. **Install Ubuntu 24.04 LTS.** X11 is needed only for a desktop tool that requires it, not headless inference.\n3. **Memory profile:** The recorded 128GB Beelink profile uses `amdgpu.gttsize=131072 ttm.pages_limit=31457280`. Do not copy these limits to 64GB/96GB systems; follow the scoped manual memory section. Add `amd_iommu=off` only for the optional desktop benchmark profile after reading [Choose the IOMMU policy](#step-12-choose-the-iommu-policy).\n4. **Driver and power policy:** Follow the measured Mesa/RADV route and record the active power manager. Preserve an existing policy by default; `tuned accelerator-performance` is an opt-in reproduction profile, not a universal requirement or guaranteed speedup.\n5. **Ollama:** Install, configure Vulkan backend with `OLLAMA_VULKAN=1`, `OLLAMA_IGPU_ENABLE=1`, and `HIP_VISIBLE_DEVICES=-1`. Without `OLLAMA_IGPU_ENABLE=1`, measured builds can detect the Radeon 8060S and still fall back to CPU-only inference.\n6. **Test:** `ollama run qwen3.6:35b-a3b` -- the measured Ollama 0.31.2 system-service path reached about 60 t/s generation. Exact speed depends on runtime, model, power state, and background load.\n\nUse the setup script below for the automated path. The phases later in this README are the manual reference and fallback path if you want to inspect or reproduce each change yourself.\n\nWant the current dense multimodal model instead? Read the\n[`Qwen3.8 27B route decision`](QWEN38_STRIX_HALO.md) before changing the\nreboot-qualified default: the official model is measured here, but its runtime,\ncontext boundary, and community performance routes need different caveats.\n\n---\n\n### Hardware Context\n\nThis table is a product/context map, not an endorsement list or the source of the thirteen-source count above. Rows without linked community evidence are useful hardware context only.\n\n| System | CPU | GPU | RAM | Notes |\n|--------|-----|-----|-----|-------|\n| **Beelink GTR9 Pro** | Ryzen AI MAX+ 395 | Radeon 8060S (40 CU) | 128GB LPDDR5X-8000 | This guide's primary test system |\n| Corsair AI Workstation 300 | Ryzen AI MAX+ 395 | Radeon 8060S (40 CU) | 128GB LPDDR5X-8000 | Three community systems reproduced the Qwen3-Coder path |\n| Framework Desktop | Ryzen AI MAX+ 395 | Radeon 8060S (40 CU) | 128GB LPDDR5X-8000 | Used by kyuz0, lhl |\n| GMKtec EVO-X2 | Ryzen AI MAX+ 395 | Radeon 8060S (40 CU) | 96GB or 128GB LPDDR5X-8000 | Native 96GB community rows reproduce Qwen3.6 and add a stock Gemma 4 direct control; a separate corrected 96GB Kyanite package reaches 262K-class Qwen3.8 retrieval on a patched HIP route; a tuned Reddit report touched 100.0 t/s on Qwen3-Coder `Q4_K_S`; [pablo-ross guide](https://github.com/pablo-ross/strix-halo-gmktec-evo-x2) |\n| Minisforum MS-S1-Max | Ryzen AI MAX+ 395 | Radeon 8060S (40 CU) | 128GB LPDDR5X | Windows LM Studio community serving report imported; not a same-shape Linux comparison |\n| Nimo AI Mini PC | Ryzen AI MAX+ 395 | Radeon 8060S (40 CU) | 128GB LPDDR5X | Community issue #4 bundle imported; adds compact-chassis large-model serving, MTP, StepFun/Qwen 122B, Gemma 4 QAT/MTP assistant-head follow-up data, and thermal context |\n| HP ZBook Ultra G1a | Ryzen AI MAX+ 395 | Radeon 8060S (40 CU) | 128GB LPDDR5X-8000 | Workstation laptop |\n\n---\n\n### Hardware Comparison\n\nCapacity and bandwidth are hardware context, not matched inference benchmarks.\n\n| Hardware | Memory capacity context | Workload evidence in this guide |\n|----------|-------------------------|---------------------------------|\n| RTX 4090 / RTX 3090 | 24GB dedicated VRAM; host offload is a separate route | Older 100-122 / 100-112 t/s ranges lacked exact matched artifact/build/workload provenance and are not","github_created_at":"2026-03-11T22:33:43+00:00","created_at":"2026-07-15T11:00:34.406008+00:00","updated_at":"2026-09-20T05:05:29.250543+00:00","categories":[{"slug":"inference-serving","name":"Inference & Serving","url":"https://www.graphcanon.com/categories/inference-serving","markdown_url":"https://www.graphcanon.com/categories/inference-serving.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/inference-serving"},{"slug":"llm-frameworks","name":"LLM Frameworks","url":"https://www.graphcanon.com/categories/llm-frameworks","markdown_url":"https://www.graphcanon.com/categories/llm-frameworks.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/llm-frameworks"}],"tags":[{"slug":"amd","name":"amd"},{"slug":"llama-cpp","name":"llama-cpp"},{"slug":"local-llm","name":"local-llm"},{"slug":"ollama","name":"ollama"},{"slug":"radeon-8060s","name":"radeon-8060s"},{"slug":"raf-ryzen-ai-max-395","name":"raf-ryzen-ai-max-395"},{"slug":"rocm","name":"rocm"},{"slug":"vulkan","name":"vulkan"}],"trust":{"provenance":{"is_fork":false,"github_id":1179303509,"owner_type":"User","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-09-20T05:05:26.245Z","maintenance":{"label":"Very active","score":96,"methodology":"github_public_v1","releases_90d":1,"days_since_push":0,"last_release_at":"2026-08-25T15:48:39Z","stars_delta_30d":69,"open_issues_delta_30d":1},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-15T11:00:35.760Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-09-20T05:05:27.512Z"},"languages":{"value":["python"],"source":"github.language","observed_at":"2026-09-20T05:05:27.512Z"},"license_spdx":{"value":"MIT","source":"github.license","observed_at":"2026-09-20T05:05:27.512Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":{"notes":["This guide is specifically for AMD-based systems equipped with Radeon 8060S GPU."],"min_ram_gb":null,"requires_docker":false},"constraints":{"min_ram_gb":null,"requires_docker":false},"when_to_use":["When you need a setup guide specifically designed for AMD Ryzen AI MAX+ 395 processor and Radeon 8060S GPU, optimized for performance in LLM environments.","To leverage ROCm support and Vulkan for optimal inference speed on Radeon hardware with frameworks such as Ollama and llama.cpp."],"when_not_to_use":["If your setup involves non-AMD hardware, especially systems without Radeon GPUs that do not benefit from the guide's specialized instructions regarding Radeon hardware and ROCm.","Avoid this guide if you require setups for other CPU or GPU brands as it is tailored to Ryzen AI MAX+ 395 and Radeon 8060S configurations."],"source":"enrich:decision_facts","observed_at":"2026-07-17T13:40:13.307Z"},"constraint_facets":{"min_ram_gb":null,"requires_docker":false},"decision_summary":[{"label":"Requirements","value":"This guide is specifically for AMD-based systems equipped with Radeon 8060S GPU."},{"label":"Adopt for","value":"strix-halo-guide is an AMD-specific guide tailored for setting up LLM environments on Radeon hardware using local AI frameworks like Ollama and llama.cpp with Vulkan support."}]}}