{"data":{"slug":"raketenkater-ggrun","name":"ggrun","tagline":"Auto-tuned launcher for GGUF models on llama.cpp with OpenAI-compatible server","github_url":"https://github.com/raketenkater/ggrun","owner":"raketenkater","repo":"ggrun","owner_avatar_url":"https://avatars.githubusercontent.com/u/49783786?v=4","primary_language":"Go","stars":275,"forks":18,"topics":["cuda","gguf","golang","inference-server","llama-cpp","llamacpp","llm","local-llm","localllama","metal","moe","multi-gpu","ollama-alternative","openai-api","self-hosted","speculative-decoding","vulkan"],"archived":false,"github_pushed_at":"2026-09-19T09:42:32+00:00","maintenance_label":"Very active","stars_delta_30d":11,"url":"https://www.graphcanon.com/tools/raketenkater-ggrun","markdown_url":"https://www.graphcanon.com/tools/raketenkater-ggrun.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/raketenkater-ggrun","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=raketenkater-ggrun","description":"llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.","homepage_url":null,"license":"MIT","open_issues":4,"watchers":3,"ai_summary":"Offers auto-tuning capabilities and an OpenAI-compatible server for GGUF models leveraging llama.cpp, supports multi-GPU tensor-split and MoE expert placement.","readme_excerpt":"## Quick start\n\nLinux / macOS:\n\n```bash\ncurl -fsSL https://raw.githubusercontent.com/raketenkater/ggrun/main/setup.sh | bash\n```\n\nWindows (PowerShell):\n\n```powershell\niwr -useb https://raw.githubusercontent.com/raketenkater/ggrun/main/install.ps1 | iex\n```\n\nOpen a new terminal after adding ggrun to PATH. To launch immediately with the\nstandard install location, use `~/ggrun/ggrun` on Linux or\n`& \"$env:USERPROFILE\\ggrun\\ggrun.cmd\"` in PowerShell.\n\nThen run a local GGUF, download one from Hugging Face, or open the TUI:\n\n```bash\nggrun model.gguf\nggrun unsloth/Qwen3.6-27B-GGUF --download","github_created_at":"2026-03-11T11:15:58+00:00","created_at":"2026-07-15T11:00:21.26198+00:00","updated_at":"2026-09-20T05:04:53.534114+00:00","categories":[{"slug":"inference-serving","name":"Inference & Serving","url":"https://www.graphcanon.com/categories/inference-serving","markdown_url":"https://www.graphcanon.com/categories/inference-serving.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/inference-serving"}],"tags":[{"slug":"cuda","name":"cuda"},{"slug":"gguf","name":"gguf"},{"slug":"golang","name":"golang"},{"slug":"inference-server","name":"inference-server"},{"slug":"llama-cpp","name":"llama-cpp"},{"slug":"llm","name":"llm"},{"slug":"local-llm","name":"local-llm"},{"slug":"metal","name":"metal"}],"trust":{"provenance":{"is_fork":false,"github_id":1178786234,"owner_type":"User","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-09-20T05:04:51.734Z","maintenance":{"label":"Very active","score":96,"methodology":"github_public_v1","releases_90d":4,"days_since_push":0,"last_release_at":"2026-09-12T14:48:47Z","stars_delta_30d":11,"open_issues_delta_30d":3},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-15T11:00:22.527Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-09-20T05:04:52.751Z"},"languages":{"value":["go"],"source":"github.language","observed_at":"2026-09-20T05:04:52.751Z"},"license_spdx":{"value":"MIT","source":"github.license","observed_at":"2026-09-20T05:04:52.751Z"}},"decision_facts":{"hosting":null,"pricing":{"model":"freemium","summary":"Free to use under MIT license; no direct costs involved in usage."},"requirements":null,"constraints":{"pricing_model":"freemium"},"when_to_use":["When developing systems that require automatic hardware optimization and tuning for GGUF models on multiple GPUs","If you need a deployment method that supports OpenAI compatible APIs while offering crash recovery mechanisms"],"when_not_to_use":["For environments where single-GPU setups are preferred, as ggrun specializes in multi-GPU configurations and may offer limited advantage or additional complexity","When you do not require auto-tuning capabilities for hardware performance optimization since this feature is specific to ggrun"],"source":"enrich:decision_facts","observed_at":"2026-07-17T08:07:44.431Z"},"constraint_facets":{"pricing_model":"freemium"},"decision_summary":[{"label":"Pricing","value":"freemium - Free to use under MIT license; no direct costs involved in usage."},{"label":"Adopt for","value":"ggrun, an auto-tuned launcher for GGUF models using llama.cpp, offers OpenAI-compatible server support with multi-GPU tensor-split and MoE expert placement capabilities."},{"label":"License detail","value":"MIT License allows using ggrun freely in both open source and commercial projects, with conditions that the copyright notice and permission notice are preserved."}]}}