{"data":{"node":{"slug":"intentee-paddler","name":"paddler","tagline":"Open-source LLM/VLM load balancer and serving platform for self-hosting at scale","github_url":"https://github.com/intentee/paddler","owner":"intentee","repo":"paddler","owner_avatar_url":"https://avatars.githubusercontent.com/u/215040511?v=4","primary_language":"Rust","stars":1663,"forks":97,"topics":["ai","llamacpp","llm","llmops","load-balancer"],"archived":false,"github_pushed_at":"2026-07-19T19:36:21+00:00","maintenance_label":"Steady","stars_delta_30d":21,"url":"https://www.graphcanon.com/tools/intentee-paddler","markdown_url":"https://www.graphcanon.com/tools/intentee-paddler.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/intentee-paddler","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=intentee-paddler"},"categories":[{"slug":"inference-serving","name":"Inference & Serving","url":"https://www.graphcanon.com/categories/inference-serving","markdown_url":"https://www.graphcanon.com/categories/inference-serving.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/inference-serving"}],"tags":[{"slug":"ai","name":"ai"},{"slug":"cpu","name":"cpu"},{"slug":"gpu","name":"gpu"},{"slug":"llamacpp","name":"llamacpp"},{"slug":"llm","name":"llm"},{"slug":"llmops","name":"llmops"},{"slug":"load-balancer","name":"load-balancer"}],"edges":[{"type":"alternative","direction":"out","explanation":"Paddler and ollama both serve as platforms to run various LLMs but Paddler focuses more on load balancing and simple deployment, making them alternatives in this context.","successor_context":null,"tool":{"slug":"ollama-ollama","name":"ollama","tagline":"Get up and running with various large language models using Ollama.","github_url":"https://github.com/ollama/ollama","owner":"ollama","repo":"ollama","owner_avatar_url":"https://avatars.githubusercontent.com/u/151674099?v=4","primary_language":"Go","stars":177524,"forks":17229,"topics":["deepseek","gemma","gemma3","glm","go","golang","gpt-oss","llama","llama3","llm","llms","minimax","mistral","ollama","qwen"],"archived":false,"github_pushed_at":"2026-07-31T23:59:29+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/ollama-ollama","markdown_url":"https://www.graphcanon.com/tools/ollama-ollama.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/ollama-ollama","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=ollama-ollama"}},{"type":"integrates_with","direction":"out","explanation":"Paddler, as an LLM serving platform, aligns well with the tools and practices listed in Awesome-LLMOps. These resources contain guidelines, software, and best practices that can be integrated into Paddler setup or deployment processes.","successor_context":null,"tool":{"slug":"tensorchord-awesome-llmops","name":"Awesome-LLMOps","tagline":"An awesome & curated list of best LLMOps tools for developers","github_url":"https://github.com/tensorchord/Awesome-LLMOps","owner":"tensorchord","repo":"Awesome-LLMOps","owner_avatar_url":"https://avatars.githubusercontent.com/u/100543303?v=4","primary_language":"Shell","stars":5915,"forks":993,"topics":["ai-development-tools","awesome-list","llmops","mlops"],"archived":false,"github_pushed_at":"2026-05-21T09:12:50+00:00","maintenance_label":"Slowing","stars_delta_30d":28,"url":"https://www.graphcanon.com/tools/tensorchord-awesome-llmops","markdown_url":"https://www.graphcanon.com/tools/tensorchord-awesome-llmops.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/tensorchord-awesome-llmops","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=tensorchord-awesome-llmops"}},{"type":"depends_on","direction":"out","explanation":null,"successor_context":null,"tool":{"slug":"ggml-org-ggml","name":"ggml","tagline":"Tensor library for machine learning","github_url":"https://github.com/ggml-org/ggml","owner":"ggml-org","repo":"ggml","owner_avatar_url":"https://avatars.githubusercontent.com/u/134263123?v=4","primary_language":"C++","stars":15185,"forks":1780,"topics":["automatic-differentiation","large-language-models","machine-learning","tensor-algebra"],"archived":false,"github_pushed_at":"2026-08-14T15:14:09+00:00","maintenance_label":"Very active","stars_delta_30d":183,"url":"https://www.graphcanon.com/tools/ggml-org-ggml","markdown_url":"https://www.graphcanon.com/tools/ggml-org-ggml.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/ggml-org-ggml","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=ggml-org-ggml"}},{"type":"integrates_with","direction":"out","explanation":"Paddler's built-in metrics and observability features can integrate well with evidently, an ML and LLM observability framework. This integration allows for advanced monitoring and debugging of deployed models.","successor_context":null,"tool":{"slug":"evidentlyai-evidently","name":"evidently","tagline":"An open-source ML and LLM observability framework.","github_url":"https://github.com/evidentlyai/evidently","owner":"evidentlyai","repo":"evidently","owner_avatar_url":"https://avatars.githubusercontent.com/u/75031056?v=4","primary_language":"Jupyter Notebook","stars":7790,"forks":895,"topics":["data-drift","data-quality","data-science","data-validation","generative-ai","hacktoberfest","html-report","jupyter-notebook","llm","llmops","machine-learning","mlops","model-monitoring","pandas-dataframe"],"archived":false,"github_pushed_at":"2026-08-05T16:29:57+00:00","maintenance_label":"Very active","stars_delta_30d":117,"url":"https://www.graphcanon.com/tools/evidentlyai-evidently","markdown_url":"https://www.graphcanon.com/tools/evidentlyai-evidently.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/evidentlyai-evidently","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=evidentlyai-evidently"}},{"type":"alternative","direction":"out","explanation":"Both sglang and Paddler offer serving frameworks for large language models with considerations for multimodal models; however, their approaches to deployment and management may differ.","successor_context":null,"tool":{"slug":"sgl-project-sglang","name":"sglang","tagline":"High-performance serving framework for large language and multimodal models","github_url":"https://github.com/sgl-project/sglang","owner":"sgl-project","repo":"sglang","owner_avatar_url":"https://avatars.githubusercontent.com/u/147780389?v=4","primary_language":"Python","stars":31454,"forks":7720,"topics":["attention","blackwell","cuda","deepseek","diffusion","glm","gpt-oss","inference","llama","llm","minimax","moe","qwen","qwen-image","reinforcement-learning","transformer","vlm","wan"],"archived":false,"github_pushed_at":"2026-08-07T06:00:20+00:00","maintenance_label":"Very active","stars_delta_30d":1409,"url":"https://www.graphcanon.com/tools/sgl-project-sglang","markdown_url":"https://www.graphcanon.com/tools/sgl-project-sglang.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/sgl-project-sglang","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=sgl-project-sglang"}},{"type":"alternative","direction":"out","explanation":"Paddler and vllm both serve as LLM/VLM serving platforms focused on ease of use, performance, and scaling. They solve similar problems in the space but may differ in specific features or underlying architecture.","successor_context":null,"tool":{"slug":"vllm-project-vllm","name":"vllm","tagline":"A high-throughput and memory-efficient inference and serving engine for LLMs","github_url":"https://github.com/vllm-project/vllm","owner":"vllm-project","repo":"vllm","owner_avatar_url":"https://avatars.githubusercontent.com/u/136984999?v=4","primary_language":"Python","stars":87847,"forks":20135,"topics":["amd","blackwell","cuda","deepseek","deepseek-v3","gpt","gpt-oss","inference","kimi","llama","llm","llm-serving","model-serving","moe","openai","pytorch","qwen","qwen3","tpu","transformer"],"archived":false,"github_pushed_at":"2026-08-01T11:55:36+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/vllm-project-vllm","markdown_url":"https://www.graphcanon.com/tools/vllm-project-vllm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/vllm-project-vllm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=vllm-project-vllm"}},{"type":"integrates_with","direction":"out","explanation":"Paddler, an open-source LLM load balancer and serving platform for deploying and scaling Language and Vision Models on private infrastructure, has an 'integrates with' relationship to vLLM, a fast, high-throughput, memory-efficient inference engine. This integration allows Paddler to leverage vLLM's performance optimizations for efficient model servicing.","successor_context":null,"tool":{"slug":"vllm-project-vllm","name":"vllm","tagline":"A high-throughput and memory-efficient inference and serving engine for LLMs","github_url":"https://github.com/vllm-project/vllm","owner":"vllm-project","repo":"vllm","owner_avatar_url":"https://avatars.githubusercontent.com/u/136984999?v=4","primary_language":"Python","stars":87847,"forks":20135,"topics":["amd","blackwell","cuda","deepseek","deepseek-v3","gpt","gpt-oss","inference","kimi","llama","llm","llm-serving","model-serving","moe","openai","pytorch","qwen","qwen3","tpu","transformer"],"archived":false,"github_pushed_at":"2026-08-01T11:55:36+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/vllm-project-vllm","markdown_url":"https://www.graphcanon.com/tools/vllm-project-vllm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/vllm-project-vllm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=vllm-project-vllm"}},{"type":"alternative","direction":"out","explanation":"Both Paddler and Agenta offer platforms for deploying, managing, and scaling LLMs on self-hosted infrastructure. They provide similar functionalities with unique approaches to solving the same challenge.","successor_context":null,"tool":{"slug":"agenta-ai-agenta","name":"agenta","tagline":"The open-source LLMOps platform for prompt management, evaluation, and observability.","github_url":"https://github.com/Agenta-AI/agenta","owner":"Agenta-AI","repo":"agenta","owner_avatar_url":"https://avatars.githubusercontent.com/u/127993667?v=4","primary_language":"TypeScript","stars":4445,"forks":609,"topics":["agent-builder","agent-observability","agent-orchestration","agent-workspace","agentic-ai","ai-agent","ai-agents","ai-automation","ai-skills-manager","ai-workflow-builder","harness","mcp","open-source","self-hosted","workflow-automation"],"archived":false,"github_pushed_at":"2026-08-07T10:41:36+00:00","maintenance_label":"Very active","stars_delta_30d":170,"url":"https://www.graphcanon.com/tools/agenta-ai-agenta","markdown_url":"https://www.graphcanon.com/tools/agenta-ai-agenta.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/agenta-ai-agenta","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=agenta-ai-agenta"}},{"type":"depends_on","direction":"out","explanation":"Paddler uses a built-in llama.cpp engine for inference, thus Paddler depends on ggml-org-llama-cpp.","successor_context":null,"tool":{"slug":"ggml-org-llama-cpp","name":"llama.cpp","tagline":"LLM inference in C/C++","github_url":"https://github.com/ggml-org/llama.cpp","owner":"ggml-org","repo":"llama.cpp","owner_avatar_url":"https://avatars.githubusercontent.com/u/134263123?v=4","primary_language":"C++","stars":122941,"forks":21406,"topics":["ggml"],"archived":false,"github_pushed_at":"2026-08-07T05:28:54+00:00","maintenance_label":"Very active","stars_delta_30d":3353,"url":"https://www.graphcanon.com/tools/ggml-org-llama-cpp","markdown_url":"https://www.graphcanon.com/tools/ggml-org-llama-cpp.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/ggml-org-llama-cpp","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=ggml-org-llama-cpp"}},{"type":"integrates_with","direction":"out","explanation":"Paddler's focus on monitoring and observability metrics aligns well with OpenLit, a platform aimed at AI engineering observability. This integration can provide comprehensive monitoring and analysis of LLM serving performance.","successor_context":null,"tool":{"slug":"openlit-openlit","name":"openlit","tagline":"A comprehensive open-source platform for AI Engineering with LLM Observability, Monitoring, and Management","github_url":"https://github.com/openlit/openlit","owner":"openlit","repo":"openlit","owner_avatar_url":"https://avatars.githubusercontent.com/u/149867240?v=4","primary_language":"TypeScript","stars":2664,"forks":342,"topics":["ai-observability","amd-gpu","clickhouse","distributed-tracing","genai","gpu-monitoring","grafana","langchain","llmops","llms","metrics","monitoring-tool","nvidia-smi","observability","open-source","openai","opentelemetry","otlp","python","tracing"],"archived":false,"github_pushed_at":"2026-07-31T18:39:37+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/openlit-openlit","markdown_url":"https://www.graphcanon.com/tools/openlit-openlit.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/openlit-openlit","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=openlit-openlit"}},{"type":"integrates_with","direction":"out","explanation":"Paddler's focus on monitoring and observation can benefit from integrating with OpenLLMetry, an observability tool designed for LLM applications. This integration would enhance Paddler’s capability to diagnose issues and monitor application performance.","successor_context":null,"tool":{"slug":"traceloop-openllmetry","name":"openllmetry","tagline":"Open-source observability for GenAI and LLM applications based on OpenTelemetry.","github_url":"https://github.com/traceloop/openllmetry","owner":"traceloop","repo":"openllmetry","owner_avatar_url":"https://avatars.githubusercontent.com/u/125419530?v=4","primary_language":"Python","stars":7377,"forks":1047,"topics":["artifical-intelligence","datascience","generative-ai","good-first-issue","good-first-issues","help-wanted","llm","llmops","metrics","ml","model-monitoring","monitoring","observability","open-source","open-telemetry","opentelemetry","opentelemetry-python","python"],"archived":false,"github_pushed_at":"2026-08-10T08:49:01+00:00","maintenance_label":"Very active","stars_delta_30d":75,"url":"https://www.graphcanon.com/tools/traceloop-openllmetry","markdown_url":"https://www.graphcanon.com/tools/traceloop-openllmetry.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/traceloop-openllmetry","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=traceloop-openllmetry"}},{"type":"related","direction":"in","explanation":"PaddleOCR and paddler are both related to the PaddlePaddle ecosystem but serve different purposes. While PaddleOCR focuses on OCR and document parsing, paddler is a platform for serving LLMs. They do not directly integrate or depend on each other.","successor_context":null,"tool":{"slug":"paddlepaddle-paddleocr","name":"PaddleOCR","tagline":"A powerful, lightweight OCR toolkit to convert images and PDFs into structured data","github_url":"https://github.com/PaddlePaddle/PaddleOCR","owner":"PaddlePaddle","repo":"PaddleOCR","owner_avatar_url":"https://avatars.githubusercontent.com/u/23534030?v=4","primary_language":"Python","stars":87808,"forks":11187,"topics":["ai4science","chineseocr","document-parsing","document-translation","kie","ocr","paddleocr-vl","pdf-extractor-rag","pdf-parser","pdf2markdown","pp-ocr","pp-structure","rag"],"archived":false,"github_pushed_at":"2026-07-22T11:59:34+00:00","maintenance_label":"Active","stars_delta_30d":2062,"url":"https://www.graphcanon.com/tools/paddlepaddle-paddleocr","markdown_url":"https://www.graphcanon.com/tools/paddlepaddle-paddleocr.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/paddlepaddle-paddleocr","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=paddlepaddle-paddleocr"}},{"type":"depends_on","direction":"in","explanation":"PaddleOCR could potentially depend on paddler for serving its OCR models at scale and load balancing, although this is speculative without more specific integration details.","successor_context":null,"tool":{"slug":"paddlepaddle-paddleocr","name":"PaddleOCR","tagline":"A powerful, lightweight OCR toolkit to convert images and PDFs into structured data","github_url":"https://github.com/PaddlePaddle/PaddleOCR","owner":"PaddlePaddle","repo":"PaddleOCR","owner_avatar_url":"https://avatars.githubusercontent.com/u/23534030?v=4","primary_language":"Python","stars":87808,"forks":11187,"topics":["ai4science","chineseocr","document-parsing","document-translation","kie","ocr","paddleocr-vl","pdf-extractor-rag","pdf-parser","pdf2markdown","pp-ocr","pp-structure","rag"],"archived":false,"github_pushed_at":"2026-07-22T11:59:34+00:00","maintenance_label":"Active","stars_delta_30d":2062,"url":"https://www.graphcanon.com/tools/paddlepaddle-paddleocr","markdown_url":"https://www.graphcanon.com/tools/paddlepaddle-paddleocr.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/paddlepaddle-paddleocr","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=paddlepaddle-paddleocr"}}],"neighbours":[{"slug":"alexsjones-llmfit","name":"llmfit","tagline":"Hundreds of models & providers. One command to find what runs on your hardware.","github_url":"https://github.com/AlexsJones/llmfit","owner":"AlexsJones","repo":"llmfit","owner_avatar_url":"https://avatars.githubusercontent.com/u/1235925?v=4","primary_language":"Rust","stars":31867,"forks":1978,"topics":["gguf","llm","localai","mlx","skill","unsloth"],"archived":false,"github_pushed_at":"2026-08-14T07:36:41+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/alexsjones-llmfit","markdown_url":"https://www.graphcanon.com/tools/alexsjones-llmfit.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/alexsjones-llmfit","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=alexsjones-llmfit","shared_categories":[]},{"slug":"eugeneyan-open-llms","name":"open-llms","tagline":"A list of open LLMs available for commercial use.","github_url":"https://github.com/eugeneyan/open-llms","owner":"eugeneyan","repo":"open-llms","owner_avatar_url":"https://avatars.githubusercontent.com/u/6831355?v=4","primary_language":null,"stars":12849,"forks":985,"topics":["commercial","large-language-models","llm","llms"],"archived":false,"github_pushed_at":"2025-02-13T06:37:12+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/eugeneyan-open-llms","markdown_url":"https://www.graphcanon.com/tools/eugeneyan-open-llms.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/eugeneyan-open-llms","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=eugeneyan-open-llms","shared_categories":[]},{"slug":"steven2358-awesome-generative-ai","name":"awesome-generative-ai","tagline":"A curated list of modern Generative Artificial Intelligence projects and services","github_url":"https://github.com/steven2358/awesome-generative-ai","owner":"steven2358","repo":"awesome-generative-ai","owner_avatar_url":"https://avatars.githubusercontent.com/u/164072?v=4","primary_language":null,"stars":12501,"forks":1990,"topics":["ai","artificial-intelligence","awesome","awesome-list","generative-ai","generative-art","large-language-models","llm"],"archived":false,"github_pushed_at":"2026-08-03T10:58:05+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/steven2358-awesome-generative-ai","markdown_url":"https://www.graphcanon.com/tools/steven2358-awesome-generative-ai.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/steven2358-awesome-generative-ai","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=steven2358-awesome-generative-ai","shared_categories":["inference-serving"]},{"slug":"andyyyy64-whichllm","name":"whichllm","tagline":"Command-line tool to find and benchmark local LLM performance","github_url":"https://github.com/Andyyyy64/whichllm","owner":"Andyyyy64","repo":"whichllm","owner_avatar_url":"https://avatars.githubusercontent.com/u/105579829?v=4","primary_language":"Python","stars":6225,"forks":330,"topics":["ai","apple-silicon","benchmarks","cli","command-line-tool","gguf","gpu","huggingface","inference","llm","local-llm","ollama","python","vram"],"archived":false,"github_pushed_at":"2026-08-05T07:15:32+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/andyyyy64-whichllm","markdown_url":"https://www.graphcanon.com/tools/andyyyy64-whichllm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/andyyyy64-whichllm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=andyyyy64-whichllm","shared_categories":["inference-serving"]},{"slug":"b4rtaz-distributed-llama","name":"distributed-llama","tagline":"Distributed LLM inference using home devices cluster","github_url":"https://github.com/b4rtaz/distributed-llama","owner":"b4rtaz","repo":"distributed-llama","owner_avatar_url":"https://avatars.githubusercontent.com/u/12797776?v=4","primary_language":"C++","stars":3012,"forks":242,"topics":["distributed-computing","distributed-llm","llama2","llama3","llm","llm-inference","llms","neural-network","open-llm"],"archived":false,"github_pushed_at":"2026-07-05T16:47:20+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/b4rtaz-distributed-llama","markdown_url":"https://www.graphcanon.com/tools/b4rtaz-distributed-llama.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/b4rtaz-distributed-llama","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=b4rtaz-distributed-llama","shared_categories":["inference-serving"]},{"slug":"rafska-awesome-local-llm","name":"awesome-local-llm","tagline":"Resources for running LLMs locally","github_url":"https://github.com/rafska/awesome-local-llm","owner":"rafska","repo":"awesome-local-llm","owner_avatar_url":"https://avatars.githubusercontent.com/u/17859377?v=4","primary_language":null,"stars":2518,"forks":316,"topics":["ai","awesome","awesome-list","llm","local","local-ai","local-llm","resources","self-hosted","selfhosted"],"archived":false,"github_pushed_at":"2026-08-04T23:08:15+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/rafska-awesome-local-llm","markdown_url":"https://www.graphcanon.com/tools/rafska-awesome-local-llm.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/rafska-awesome-local-llm","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=rafska-awesome-local-llm","shared_categories":["inference-serving"]},{"slug":"superlinked-sie","name":"sie","tagline":"Open-source inference server and production cluster for all the models your agent needs.","github_url":"https://github.com/superlinked/sie","owner":"superlinked","repo":"sie","owner_avatar_url":"https://avatars.githubusercontent.com/u/94243920?v=4","primary_language":"Python","stars":2297,"forks":215,"topics":["bge","colbert","data-pipeline","deep-learning","embeddings","inference","inference-server","information-retrieval","llm","ml","mlops","natural-language-processing","nlp","python","reranking","retrieval","retrieval-augmented-generation","semantic-search","splade","vector-search"],"archived":false,"github_pushed_at":"2026-07-22T08:53:55+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/superlinked-sie","markdown_url":"https://www.graphcanon.com/tools/superlinked-sie.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/superlinked-sie","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=superlinked-sie","shared_categories":["inference-serving"]},{"slug":"waybarrios-vllm-mlx","name":"vllm-mlx","tagline":"Server for LLMs and vision-language models compatible with Apple Silicon","github_url":"https://github.com/waybarrios/vllm-mlx","owner":"waybarrios","repo":"vllm-mlx","owner_avatar_url":"https://avatars.githubusercontent.com/u/6794828?v=4","primary_language":"Python","stars":1472,"forks":205,"topics":["anthropic","apple-silicon","audio-processing","claude-code","computer-vision","image-understanding","inference","llm","machine-learning","macos","mllm","mlx","multimodal-ai","speech-to-text","stt","text-to-speech","tts","video-understanding","vision-language-model","vllm"],"archived":false,"github_pushed_at":"2026-06-28T20:18:31+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/waybarrios-vllm-mlx","markdown_url":"https://www.graphcanon.com/tools/waybarrios-vllm-mlx.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/waybarrios-vllm-mlx","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=waybarrios-vllm-mlx","shared_categories":["inference-serving"]},{"slug":"ebhy-budgetml","name":"budgetml","tagline":"Deploys ML inference service economically","github_url":"https://github.com/ebhy/budgetml","owner":"ebhy","repo":"budgetml","owner_avatar_url":"https://avatars.githubusercontent.com/u/76654256?v=4","primary_language":"Python","stars":1343,"forks":65,"topics":["api","data-science","deployment","fastapi","inference","machine-learning","mlops"],"archived":false,"github_pushed_at":"2024-02-12T17:29:24+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/ebhy-budgetml","markdown_url":"https://www.graphcanon.com/tools/ebhy-budgetml.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/ebhy-budgetml","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=ebhy-budgetml","shared_categories":["inference-serving"]},{"slug":"kubeai-project-kubeai","name":"kubeai","tagline":"AI Inference Operator for Kubernetes","github_url":"https://github.com/kubeai-project/kubeai","owner":"kubeai-project","repo":"kubeai","owner_avatar_url":"https://avatars.githubusercontent.com/u/232319222?v=4","primary_language":"Go","stars":1237,"forks":131,"topics":["ai","autoscaler","faster-whisper","inference-operator","k8s","kubernetes","llm","ollama","ollama-operator","openai-api","vllm","vllm-operator","whisper"],"archived":false,"github_pushed_at":"2026-07-31T01:04:47+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/kubeai-project-kubeai","markdown_url":"https://www.graphcanon.com/tools/kubeai-project-kubeai.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/kubeai-project-kubeai","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=kubeai-project-kubeai","shared_categories":["inference-serving"]},{"slug":"openinfer-project-openinfer","name":"openinfer","tagline":"Pure Rust CUDA LLM inference engine serving multiple models including Qwen3 and Kimi-K2","github_url":"https://github.com/openinfer-project/openinfer","owner":"openinfer-project","repo":"openinfer","owner_avatar_url":"https://avatars.githubusercontent.com/u/292134277?v=4","primary_language":"Rust","stars":585,"forks":89,"topics":["cuda","cuda-kernels","deepseek","gpu","inference","inference-engine","kimi","kimi-k2","kv-cache","llm","llm-inference","llm-serving","model-serving","moe","openai-api","paged-attention","qwen","qwen3","rust","vllm"],"archived":false,"github_pushed_at":"2026-07-25T14:08:34+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/openinfer-project-openinfer","markdown_url":"https://www.graphcanon.com/tools/openinfer-project-openinfer.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/openinfer-project-openinfer","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=openinfer-project-openinfer","shared_categories":["inference-serving"]},{"slug":"kaito-project-aikit","name":"aikit","tagline":"Fine-tune, build, and deploy open-source LLMs easily!","github_url":"https://github.com/kaito-project/aikit","owner":"kaito-project","repo":"aikit","owner_avatar_url":"https://avatars.githubusercontent.com/u/186863079?v=4","primary_language":"Go","stars":534,"forks":57,"topics":["ai","buildkit","chatgpt","docker","fine-tuning","finetuning","gemma","gpt","inference","kubernetes","large-language-models","llama","llm","localllama","mistral","mixtral","nvidia","open-llm","open-source-llm","openai"],"archived":false,"github_pushed_at":"2026-07-20T03:11:50+00:00","maintenance_label":"Steady","url":"https://www.graphcanon.com/tools/kaito-project-aikit","markdown_url":"https://www.graphcanon.com/tools/kaito-project-aikit.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/kaito-project-aikit","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=kaito-project-aikit","shared_categories":["inference-serving"]}]}}