{"data":{"node":{"slug":"ggml-org-llama-cpp","name":"llama.cpp","tagline":"LLM inference in C/C++","github_url":"https://github.com/ggml-org/llama.cpp","owner":"ggml-org","repo":"llama.cpp","owner_avatar_url":"https://avatars.githubusercontent.com/u/134263123?v=4","primary_language":"C++","stars":122941,"forks":21406,"topics":["ggml"],"archived":false,"github_pushed_at":"2026-08-07T05:28:54+00:00","maintenance_label":"Very active","stars_delta_30d":3353,"url":"https://www.graphcanon.com/tools/ggml-org-llama-cpp","markdown_url":"https://www.graphcanon.com/tools/ggml-org-llama-cpp.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/ggml-org-llama-cpp","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=ggml-org-llama-cpp"},"categories":[{"slug":"inference-serving","name":"Inference & Serving","url":"https://www.graphcanon.com/categories/inference-serving","markdown_url":"https://www.graphcanon.com/categories/inference-serving.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/inference-serving"}],"tags":[{"slug":"c","name":"c++"},{"slug":"ggml","name":"ggml"}],"edges":[{"type":"related","direction":"out","explanation":null,"successor_context":null,"tool":{"slug":"unslothai-unsloth","name":"unsloth","tagline":"A web UI for training and running open models locally.","github_url":"https://github.com/unslothai/unsloth","owner":"unslothai","repo":"unsloth","owner_avatar_url":"https://avatars.githubusercontent.com/u/150920049?v=4","primary_language":"Python","stars":69621,"forks":6285,"topics":["agent","deepseek","fine-tuning","gemma","gemma3","gpt-oss","llama","llama3","llm","llms","mistral","openai","qwen","reinforcement-learning","self-hosted","text-to-speech","tts","ui","unsloth"],"archived":false,"github_pushed_at":"2026-08-06T06:01:56+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/unslothai-unsloth","markdown_url":"https://www.graphcanon.com/tools/unslothai-unsloth.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/unslothai-unsloth","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=unslothai-unsloth"}},{"type":"alternative","direction":"out","explanation":"Both airllm and llama.cpp offer lightweight GPU inference options for large language models, differing mainly in their implementation and optimization approaches.","successor_context":null,"tool":{"slug":"lyogavin-airllm","name":"airllm","tagline":"AirLLM 70B inference with single 4GB GPU","github_url":"https://github.com/lyogavin/airllm","owner":"lyogavin","repo":"airllm","owner_avatar_url":"https://avatars.githubusercontent.com/u/1113905?v=4","primary_language":"Jupyter Notebook","stars":24183,"forks":2722,"topics":["chinese-llm","chinese-nlp","finetune","generative-ai","instruct-gpt","instruction-set","llama","llm","lora","open-models","open-source","open-source-models","qlora"],"archived":false,"github_pushed_at":"2026-07-23T08:29:43+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/lyogavin-airllm","markdown_url":"https://www.graphcanon.com/tools/lyogavin-airllm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/lyogavin-airllm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=lyogavin-airllm"}},{"type":"integrates_with","direction":"out","explanation":"`wllama` provides WebAssembly bindings for `llama.cpp`, enabling on-browser LLM inference, indicating they integrate well.","successor_context":null,"tool":{"slug":"ngxson-wllama","name":"wllama","tagline":"WebAssembly binding for llama.cpp - Enabling on-browser LLM inference","github_url":"https://github.com/ngxson/wllama","owner":"ngxson","repo":"wllama","owner_avatar_url":"https://avatars.githubusercontent.com/u/7702203?v=4","primary_language":"TypeScript","stars":1159,"forks":117,"topics":["llama","llamacpp","llm","wasm","webassembly"],"archived":false,"github_pushed_at":"2026-06-17T17:32:59+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/ngxson-wllama","markdown_url":"https://www.graphcanon.com/tools/ngxson-wllama.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/ngxson-wllama","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=ngxson-wllama"}},{"type":"alternative","direction":"out","explanation":"Both `tiny-vllm` and `llama.cpp` are designed to provide high-performance LLM inference engines in C++, making them alternatives.","successor_context":null,"tool":{"slug":"jmaczan-tiny-vllm","name":"tiny-vllm","tagline":"Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM","github_url":"https://github.com/jmaczan/tiny-vllm","owner":"jmaczan","repo":"tiny-vllm","owner_avatar_url":"https://avatars.githubusercontent.com/u/18054202?v=4","primary_language":"C++","stars":947,"forks":68,"topics":["ai","attention","batching","course","cpp","cuda","hpc","inference","llm","llm-inference","pagedattention","tiny-vllm","vllm"],"archived":false,"github_pushed_at":"2026-07-02T18:32:16+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/jmaczan-tiny-vllm","markdown_url":"https://www.graphcanon.com/tools/jmaczan-tiny-vllm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/jmaczan-tiny-vllm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=jmaczan-tiny-vllm"}},{"type":"related","direction":"out","explanation":"`Awesome-LLM-Inference` is a curated list of inference resources and papers which may include mentions or resources related to `llama.cpp`, making them adjacent yet not directly connected.","successor_context":null,"tool":{"slug":"xlite-dev-awesome-llm-inference","name":"Awesome-LLM-Inference","tagline":"A curated list of LLM/VLM inference papers with codes","github_url":"https://github.com/xlite-dev/Awesome-LLM-Inference","owner":"xlite-dev","repo":"Awesome-LLM-Inference","owner_avatar_url":"https://avatars.githubusercontent.com/u/204302598?v=4","primary_language":"Python","stars":5415,"forks":428,"topics":["awesome-llm","deepseek","deepseek-r1","deepseek-v3","flash-attention","flash-attention-3","flash-mla","llm-inference","minimax-01","mla","paged-attention","qwen3","tensorrt-llm","vllm"],"archived":false,"github_pushed_at":"2026-06-23T03:48:43+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/xlite-dev-awesome-llm-inference","markdown_url":"https://www.graphcanon.com/tools/xlite-dev-awesome-llm-inference.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/xlite-dev-awesome-llm-inference","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=xlite-dev-awesome-llm-inference"}},{"type":"integrates_with","direction":"out","explanation":"`TurboLLM` can use `llama.cpp` as one of its engines for running LLM models, showing integration potential.","successor_context":null,"tool":{"slug":"mohitsoni48-turbollm","name":"TurboLLM","tagline":"Run any local LLM engine auto-tuned to your GPU with polished web UI and OpenAI/Anthropic-compatible API","github_url":"https://github.com/mohitsoni48/TurboLLM","owner":"mohitsoni48","repo":"TurboLLM","owner_avatar_url":"https://avatars.githubusercontent.com/u/63787789?v=4","primary_language":"TypeScript","stars":225,"forks":36,"topics":["ai","anthropic-api","claude-code","gguf","gpu","inference","llama-cpp","llama-server","llm","local-llm","offline","openai-api","self-hosted"],"archived":false,"github_pushed_at":"2026-08-11T13:28:32+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/mohitsoni48-turbollm","markdown_url":"https://www.graphcanon.com/tools/mohitsoni48-turbollm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/mohitsoni48-turbollm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=mohitsoni48-turbollm"}},{"type":"alternative","direction":"out","explanation":"`VLLM` and `llama.cpp` both offer high-throughput and efficient LLM inference engines, thus they are considered alternatives.","successor_context":null,"tool":{"slug":"vllm-project-vllm","name":"vllm","tagline":"A high-throughput and memory-efficient inference and serving engine for LLMs","github_url":"https://github.com/vllm-project/vllm","owner":"vllm-project","repo":"vllm","owner_avatar_url":"https://avatars.githubusercontent.com/u/136984999?v=4","primary_language":"Python","stars":87847,"forks":20135,"topics":["amd","blackwell","cuda","deepseek","deepseek-v3","gpt","gpt-oss","inference","kimi","llama","llm","llm-serving","model-serving","moe","openai","pytorch","qwen","qwen3","tpu","transformer"],"archived":false,"github_pushed_at":"2026-08-01T11:55:36+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/vllm-project-vllm","markdown_url":"https://www.graphcanon.com/tools/vllm-project-vllm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/vllm-project-vllm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=vllm-project-vllm"}},{"type":"integrates_with","direction":"out","explanation":"`ggrun` requires `llama.cpp` to function as the engine for launching GGUF models, indicating a dependency.","successor_context":null,"tool":{"slug":"raketenkater-ggrun","name":"ggrun","tagline":"Auto-tuned launcher for GGUF models on llama.cpp with OpenAI-compatible server","github_url":"https://github.com/raketenkater/ggrun","owner":"raketenkater","repo":"ggrun","owner_avatar_url":"https://avatars.githubusercontent.com/u/49783786?v=4","primary_language":"Go","stars":264,"forks":14,"topics":["cuda","gguf","golang","inference-server","llama-cpp","llamacpp","llm","local-llm","localllama","metal","moe","multi-gpu","ollama-alternative","openai-api","self-hosted","speculative-decoding","vulkan"],"archived":false,"github_pushed_at":"2026-08-11T21:16:38+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/raketenkater-ggrun","markdown_url":"https://www.graphcanon.com/tools/raketenkater-ggrun.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/raketenkater-ggrun","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=raketenkater-ggrun"}},{"type":"alternative","direction":"in","explanation":"Both LocalAI and llama.cpp provide means to run inference on large language models, with LocalAI offering a more versatile runtime.","successor_context":null,"tool":{"slug":"mudler-localai","name":"LocalAI","tagline":"Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.","github_url":"https://github.com/mudler/LocalAI","owner":"mudler","repo":"LocalAI","owner_avatar_url":"https://avatars.githubusercontent.com/u/2420543?v=4","primary_language":"Go","stars":48500,"forks":4362,"topics":["agents","ai","api","audio-generation","decentralized","distributed","image-generation","libp2p","llama","llm","mamba","mcp","musicgen","object-detection","rerank","stable-diffusion","text-generation","tts"],"archived":false,"github_pushed_at":"2026-08-16T05:07:25+00:00","maintenance_label":"Very active","stars_delta_30d":924,"url":"https://www.graphcanon.com/tools/mudler-localai","markdown_url":"https://www.graphcanon.com/tools/mudler-localai.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/mudler-localai","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=mudler-localai"}},{"type":"depends_on","direction":"in","explanation":"Jan runs LLMs offline, which likely uses llama.cpp under the hood for inference in C/C++.","successor_context":null,"tool":{"slug":"janhq-jan","name":"jan","tagline":"open source alternative to ChatGPT that runs offline locally","github_url":"https://github.com/janhq/jan","owner":"janhq","repo":"jan","owner_avatar_url":"https://avatars.githubusercontent.com/u/102363196?v=4","primary_language":"TypeScript","stars":44020,"forks":2971,"topics":["chatgpt","gpt","llamacpp","llm","localai","open-source","self-hosted","tauri"],"archived":false,"github_pushed_at":"2026-08-14T12:24:52+00:00","maintenance_label":"Very active","stars_delta_30d":426,"url":"https://www.graphcanon.com/tools/janhq-jan","markdown_url":"https://www.graphcanon.com/tools/janhq-jan.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/janhq-jan","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=janhq-jan"}},{"type":"alternative","direction":"in","explanation":"Both MLC-LLM and llama.cpp are focused on LLM inference but with different hardware support and optimizations.","successor_context":null,"tool":{"slug":"mlc-ai-mlc-llm","name":"mlc-llm","tagline":"Universal LLM Deployment Engine with ML Compilation","github_url":"https://github.com/mlc-ai/mlc-llm","owner":"mlc-ai","repo":"mlc-llm","owner_avatar_url":"https://avatars.githubusercontent.com/u/106173866?v=4","primary_language":"Python","stars":23063,"forks":2111,"topics":["language-model","llm","machine-learning-compilation","tvm"],"archived":false,"github_pushed_at":"2026-07-31T03:03:18+00:00","maintenance_label":"Active","stars_delta_30d":103,"url":"https://www.graphcanon.com/tools/mlc-ai-mlc-llm","markdown_url":"https://www.graphcanon.com/tools/mlc-ai-mlc-llm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/mlc-ai-mlc-llm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=mlc-ai-mlc-llm"}},{"type":"depends_on","direction":"in","explanation":"LEANN can potentially depend on llama.cpp for efficient LLM inference in C/C++, which aligns with its aim of providing fast and accurate RAG applications locally.","successor_context":null,"tool":{"slug":"startrail-org-leann","name":"LEANN","tagline":"RAG on Everything with LEANN","github_url":"https://github.com/StarTrail-org/LEANN","owner":"StarTrail-org","repo":"LEANN","owner_avatar_url":"https://avatars.githubusercontent.com/u/288858980?v=4","primary_language":"Python","stars":12785,"forks":1145,"topics":["ai","faiss","gpt-oss","langchain","llama-index","llm","localstorage","offline-first","ollama","privacy","python","rag","retrieval-augmented-generation","vector-database","vector-search","vectors"],"archived":false,"github_pushed_at":"2026-07-31T18:53:24+00:00","maintenance_label":"Active","stars_delta_30d":81,"url":"https://www.graphcanon.com/tools/startrail-org-leann","markdown_url":"https://www.graphcanon.com/tools/startrail-org-leann.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/startrail-org-leann","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=startrail-org-leann"}},{"type":"integrates_with","direction":"in","explanation":"Jan uses C/C++ inference for running LLM models and can potentially integrate with llama.cpp or other similar libraries.","successor_context":null,"tool":{"slug":"janhq-jan","name":"jan","tagline":"open source alternative to ChatGPT that runs offline locally","github_url":"https://github.com/janhq/jan","owner":"janhq","repo":"jan","owner_avatar_url":"https://avatars.githubusercontent.com/u/102363196?v=4","primary_language":"TypeScript","stars":44020,"forks":2971,"topics":["chatgpt","gpt","llamacpp","llm","localai","open-source","self-hosted","tauri"],"archived":false,"github_pushed_at":"2026-08-14T12:24:52+00:00","maintenance_label":"Very active","stars_delta_30d":426,"url":"https://www.graphcanon.com/tools/janhq-jan","markdown_url":"https://www.graphcanon.com/tools/janhq-jan.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/janhq-jan","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=janhq-jan"}},{"type":"alternative","direction":"in","explanation":"Both PowerInfer and llama.cpp are focused on LLM inference, particularly for local deployment. While they serve similar purposes, they do so through different implementations and potentially with varying performance characteristics.","successor_context":null,"tool":{"slug":"tiiny-ai-powerinfer","name":"PowerInfer","tagline":"High-speed Large Language Model Serving for Local Deployment","github_url":"https://github.com/Tiiny-AI/PowerInfer","owner":"Tiiny-AI","repo":"PowerInfer","owner_avatar_url":"https://avatars.githubusercontent.com/u/256922953?v=4","primary_language":"C++","stars":9718,"forks":591,"topics":["large-language-models","llama","llm","llm-inference","local-inference"],"archived":false,"github_pushed_at":"2026-05-11T06:48:06+00:00","maintenance_label":"Slowing","stars_delta_30d":76,"url":"https://www.graphcanon.com/tools/tiiny-ai-powerinfer","markdown_url":"https://www.graphcanon.com/tools/tiiny-ai-powerinfer.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/tiiny-ai-powerinfer","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=tiiny-ai-powerinfer"}},{"type":"alternative","direction":"in","explanation":"OptiMate and llama.cpp both focus on optimizing LLMs, but they use different programming languages (Python vs C/C++).","successor_context":null,"tool":{"slug":"nebuly-ai-optimate","name":"optimate","tagline":"A collection of libraries to optimize AI model performances","github_url":"https://github.com/nebuly-ai/optimate","owner":"nebuly-ai","repo":"optimate","owner_avatar_url":"https://avatars.githubusercontent.com/u/83510798?v=4","primary_language":"Python","stars":8329,"forks":617,"topics":["ai","analytics","artificial-intelligence","deeplearning","large-language-models","llm"],"archived":false,"github_pushed_at":"2024-07-22T02:07:03+00:00","maintenance_label":"Dormant","stars_delta_30d":-3,"url":"https://www.graphcanon.com/tools/nebuly-ai-optimate","markdown_url":"https://www.graphcanon.com/tools/nebuly-ai-optimate.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/nebuly-ai-optimate","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=nebuly-ai-optimate"}},{"type":"alternative","direction":"in","explanation":"Both Video-LLaMA and llama.cpp offer inference capabilities for Large Language Models, but Video-LLaMA is geared towards instruction-tuned video understanding.","successor_context":null,"tool":{"slug":"damo-nlp-sg-video-llama","name":"Video-LLaMA","tagline":"Instruction-tuned Audio-Visual Language Model for Video Understanding","github_url":"https://github.com/DAMO-NLP-SG/Video-LLaMA","owner":"DAMO-NLP-SG","repo":"Video-LLaMA","owner_avatar_url":"https://avatars.githubusercontent.com/u/130957594?v=4","primary_language":"Python","stars":3141,"forks":287,"topics":["blip2","cross-modal-pretraining","large-language-models","llama","minigpt4","multi-modal-chatgpt","video-language-pretraining","vision-language-pretraining"],"archived":false,"github_pushed_at":"2024-06-04T07:06:41+00:00","maintenance_label":"Dormant","stars_delta_30d":2,"url":"https://www.graphcanon.com/tools/damo-nlp-sg-video-llama","markdown_url":"https://www.graphcanon.com/tools/damo-nlp-sg-video-llama.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/damo-nlp-sg-video-llama","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=damo-nlp-sg-video-llama"}},{"type":"depends_on","direction":"in","explanation":"Paddler uses a built-in llama.cpp engine for inference, thus Paddler depends on ggml-org-llama-cpp.","successor_context":null,"tool":{"slug":"intentee-paddler","name":"paddler","tagline":"Open-source LLM/VLM load balancer and serving platform for self-hosting at scale","github_url":"https://github.com/intentee/paddler","owner":"intentee","repo":"paddler","owner_avatar_url":"https://avatars.githubusercontent.com/u/215040511?v=4","primary_language":"Rust","stars":1663,"forks":97,"topics":["ai","llamacpp","llm","llmops","load-balancer"],"archived":false,"github_pushed_at":"2026-07-19T19:36:21+00:00","maintenance_label":"Steady","stars_delta_30d":21,"url":"https://www.graphcanon.com/tools/intentee-paddler","markdown_url":"https://www.graphcanon.com/tools/intentee-paddler.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/intentee-paddler","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=intentee-paddler"}},{"type":"related","direction":"in","explanation":"Both work on LLM inference but ggml focuses on low-level C/C++ implementation whereas petals provides a higher level distributed execution framework.","successor_context":null,"tool":{"slug":"bigscience-workshop-petals","name":"petals","tagline":"Run LLMs at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading","github_url":"https://github.com/bigscience-workshop/petals","owner":"bigscience-workshop","repo":"petals","owner_avatar_url":"https://avatars.githubusercontent.com/u/82455566?v=4","primary_language":"Python","stars":10496,"forks":642,"topics":["bloom","chatbot","deep-learning","distributed-systems","falcon","gpt","guanaco","language-models","large-language-models","llama","machine-learning","mixtral","neural-networks","nlp","pipeline-parallelism","pretrained-models","pytorch","tensor-parallelism","transformer","volunteer-computing"],"archived":false,"github_pushed_at":"2024-09-07T11:54:28+00:00","maintenance_label":"Dormant","stars_delta_30d":212,"url":"https://www.graphcanon.com/tools/bigscience-workshop-petals","markdown_url":"https://www.graphcanon.com/tools/bigscience-workshop-petals.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/bigscience-workshop-petals","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=bigscience-workshop-petals"}},{"type":"related","direction":"in","explanation":"Both Ollama and llama.cpp are related to LLM inference, but while Ollama seems to provide a broader ecosystem (models + infrastructure), llama.cpp is specific to C/C++ inference.","successor_context":null,"tool":{"slug":"ollama-ollama","name":"ollama","tagline":"Get up and running with various large language models using Ollama.","github_url":"https://github.com/ollama/ollama","owner":"ollama","repo":"ollama","owner_avatar_url":"https://avatars.githubusercontent.com/u/151674099?v=4","primary_language":"Go","stars":177524,"forks":17229,"topics":["deepseek","gemma","gemma3","glm","go","golang","gpt-oss","llama","llama3","llm","llms","minimax","mistral","ollama","qwen"],"archived":false,"github_pushed_at":"2026-07-31T23:59:29+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/ollama-ollama","markdown_url":"https://www.graphcanon.com/tools/ollama-ollama.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/ollama-ollama","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=ollama-ollama"}}],"neighbours":[{"slug":"vllm-project-vllm","name":"vllm","tagline":"A high-throughput and memory-efficient inference and serving engine for LLMs","github_url":"https://github.com/vllm-project/vllm","owner":"vllm-project","repo":"vllm","owner_avatar_url":"https://avatars.githubusercontent.com/u/136984999?v=4","primary_language":"Python","stars":87847,"forks":20135,"topics":["amd","blackwell","cuda","deepseek","deepseek-v3","gpt","gpt-oss","inference","kimi","llama","llm","llm-serving","model-serving","moe","openai","pytorch","qwen","qwen3","tpu","transformer"],"archived":false,"github_pushed_at":"2026-08-01T11:55:36+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/vllm-project-vllm","markdown_url":"https://www.graphcanon.com/tools/vllm-project-vllm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/vllm-project-vllm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=vllm-project-vllm","shared_categories":["inference-serving"]},{"slug":"nomic-ai-gpt4all","name":"gpt4all","tagline":"Run Local LLMs on Any Device","github_url":"https://github.com/nomic-ai/gpt4all","owner":"nomic-ai","repo":"gpt4all","owner_avatar_url":"https://avatars.githubusercontent.com/u/102670180?v=4","primary_language":"C++","stars":77396,"forks":8304,"topics":["ai-chat","llm-inference"],"archived":false,"github_pushed_at":"2025-05-27T20:05:19+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/nomic-ai-gpt4all","markdown_url":"https://www.graphcanon.com/tools/nomic-ai-gpt4all.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/nomic-ai-gpt4all","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=nomic-ai-gpt4all","shared_categories":["inference-serving"]},{"slug":"mozilla-ai-llamafile","name":"llamafile","tagline":"Distribute and run LLMs with a single file.","github_url":"https://github.com/mozilla-ai/llamafile","owner":"mozilla-ai","repo":"llamafile","owner_avatar_url":"https://avatars.githubusercontent.com/u/129804596?v=4","primary_language":"C++","stars":25470,"forks":1530,"topics":["cross-platform","gguf","llama-cpp","local-ai","local-inference","local-llm","open-source-ai","single-file-executable","speech-to-text"],"archived":false,"github_pushed_at":"2026-07-27T14:21:26+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/mozilla-ai-llamafile","markdown_url":"https://www.graphcanon.com/tools/mozilla-ai-llamafile.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/mozilla-ai-llamafile","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=mozilla-ai-llamafile","shared_categories":["inference-serving"]},{"slug":"lyogavin-airllm","name":"airllm","tagline":"AirLLM 70B inference with single 4GB GPU","github_url":"https://github.com/lyogavin/airllm","owner":"lyogavin","repo":"airllm","owner_avatar_url":"https://avatars.githubusercontent.com/u/1113905?v=4","primary_language":"Jupyter Notebook","stars":24183,"forks":2722,"topics":["chinese-llm","chinese-nlp","finetune","generative-ai","instruct-gpt","instruction-set","llama","llm","lora","open-models","open-source","open-source-models","qlora"],"archived":false,"github_pushed_at":"2026-07-23T08:29:43+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/lyogavin-airllm","markdown_url":"https://www.graphcanon.com/tools/lyogavin-airllm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/lyogavin-airllm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=lyogavin-airllm","shared_categories":["inference-serving"]},{"slug":"jundot-omlx","name":"omlx","tagline":"LLM inference server with continuous batching and SSD caching for Apple Silicon","github_url":"https://github.com/jundot/omlx","owner":"jundot","repo":"omlx","owner_avatar_url":"https://avatars.githubusercontent.com/u/64250138?v=4","primary_language":"Python","stars":18679,"forks":1617,"topics":["apple-silicon","inference-server","llm","macos","mlx","openai-api"],"archived":false,"github_pushed_at":"2026-08-14T09:12:48+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/jundot-omlx","markdown_url":"https://www.graphcanon.com/tools/jundot-omlx.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/jundot-omlx","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=jundot-omlx","shared_categories":["inference-serving"]},{"slug":"lightning-ai-litgpt","name":"litgpt","tagline":"High-performance LLMs with recipes for pretraining, finetuning and deployment","github_url":"https://github.com/Lightning-AI/litgpt","owner":"Lightning-AI","repo":"litgpt","owner_avatar_url":"https://avatars.githubusercontent.com/u/58386951?v=4","primary_language":"Python","stars":13605,"forks":1483,"topics":["ai","artificial-intelligence","deep-learning","large-language-models","llm","llm-inference","llms"],"archived":false,"github_pushed_at":"2026-07-20T10:24:12+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/lightning-ai-litgpt","markdown_url":"https://www.graphcanon.com/tools/lightning-ai-litgpt.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/lightning-ai-litgpt","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=lightning-ai-litgpt","shared_categories":["inference-serving"]},{"slug":"flashinfer-ai-flashinfer","name":"flashinfer","tagline":"FlashInfer is a kernel library for serving large language models","github_url":"https://github.com/flashinfer-ai/flashinfer","owner":"flashinfer-ai","repo":"flashinfer","owner_avatar_url":"https://avatars.githubusercontent.com/u/145061914?v=4","primary_language":"Python","stars":6024,"forks":1196,"topics":["attention","cuda","distributed-inference","gpu","jit","large-large-models","llm-inference","moe","nvidia","pytorch"],"archived":false,"github_pushed_at":"2026-07-25T05:01:59+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/flashinfer-ai-flashinfer","markdown_url":"https://www.graphcanon.com/tools/flashinfer-ai-flashinfer.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/flashinfer-ai-flashinfer","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=flashinfer-ai-flashinfer","shared_categories":["inference-serving"]},{"slug":"turboderp-exllama","name":"exllama","tagline":"Memory-efficient rewrite of HF transformers for Llama with quantized weights","github_url":"https://github.com/turboderp/exllama","owner":"turboderp","repo":"exllama","owner_avatar_url":"https://avatars.githubusercontent.com/u/11859846?v=4","primary_language":"Python","stars":2937,"forks":220,"topics":[],"archived":false,"github_pushed_at":"2023-09-30T19:06:04+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/turboderp-exllama","markdown_url":"https://www.graphcanon.com/tools/turboderp-exllama.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/turboderp-exllama","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=turboderp-exllama","shared_categories":["inference-serving"]},{"slug":"liltom-eth-llama2-webui","name":"llama2-webui","tagline":"Run Llama 2 locally with gradio UI on GPU or CPU","github_url":"https://github.com/liltom-eth/llama2-webui","owner":"liltom-eth","repo":"llama2-webui","owner_avatar_url":"https://avatars.githubusercontent.com/u/11456256?v=4","primary_language":"Jupyter Notebook","stars":1937,"forks":201,"topics":["llama-2","llama2","llm","llm-inference"],"archived":false,"github_pushed_at":"2024-03-22T09:50:24+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/liltom-eth-llama2-webui","markdown_url":"https://www.graphcanon.com/tools/liltom-eth-llama2-webui.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/liltom-eth-llama2-webui","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=liltom-eth-llama2-webui","shared_categories":["inference-serving"]},{"slug":"xllm-ai-xllm","name":"xllm","tagline":"A high-performance inference engine for LLM, VLM, DiT and REC models","github_url":"https://github.com/xLLM-AI/xllm","owner":"xLLM-AI","repo":"xllm","owner_avatar_url":"https://avatars.githubusercontent.com/u/205719415?v=4","primary_language":"C++","stars":1493,"forks":269,"topics":["deepseek","glm","inference","inference-engine","large-language-models","llm-inference","qwen"],"archived":false,"github_pushed_at":"2026-07-24T10:38:15+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/xllm-ai-xllm","markdown_url":"https://www.graphcanon.com/tools/xllm-ai-xllm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/xllm-ai-xllm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=xllm-ai-xllm","shared_categories":["inference-serving"]},{"slug":"ngxson-wllama","name":"wllama","tagline":"WebAssembly binding for llama.cpp - Enabling on-browser LLM inference","github_url":"https://github.com/ngxson/wllama","owner":"ngxson","repo":"wllama","owner_avatar_url":"https://avatars.githubusercontent.com/u/7702203?v=4","primary_language":"TypeScript","stars":1159,"forks":117,"topics":["llama","llamacpp","llm","wasm","webassembly"],"archived":false,"github_pushed_at":"2026-06-17T17:32:59+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/ngxson-wllama","markdown_url":"https://www.graphcanon.com/tools/ngxson-wllama.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/ngxson-wllama","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=ngxson-wllama","shared_categories":["inference-serving"]},{"slug":"jmaczan-tiny-vllm","name":"tiny-vllm","tagline":"Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM","github_url":"https://github.com/jmaczan/tiny-vllm","owner":"jmaczan","repo":"tiny-vllm","owner_avatar_url":"https://avatars.githubusercontent.com/u/18054202?v=4","primary_language":"C++","stars":947,"forks":68,"topics":["ai","attention","batching","course","cpp","cuda","hpc","inference","llm","llm-inference","pagedattention","tiny-vllm","vllm"],"archived":false,"github_pushed_at":"2026-07-02T18:32:16+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/jmaczan-tiny-vllm","markdown_url":"https://www.graphcanon.com/tools/jmaczan-tiny-vllm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/jmaczan-tiny-vllm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=jmaczan-tiny-vllm","shared_categories":["inference-serving"]}]}}