---
title: "Inference & Serving"
type: "category"
slug: "inference-serving"
canonical_url: "https://www.graphcanon.com/categories/inference-serving"
tool_count: 406
---

# Inference & Serving

*GraphCanon updated Aug 21, 2026*

Model inference, serving, and local runtimes — deploying and running models efficiently (Ollama, vLLM, llama.cpp, TGI).

406 tools in this category (showing the top 60 by stars).

## Featured comparisons

- [Yi vs Qwen](/compare/01-ai-yi-vs-qwenlm-qwen.md)
- [Yi vs LLMs-from-scratch](/compare/01-ai-yi-vs-rasbt-llms-from-scratch.md)
- [1Panel vs open-webui](/compare/1panel-dev-1panel-vs-open-webui-open-webui.md)
- [CV vs self-llm](/compare/accumulatemore-cv-vs-datawhalechina-self-llm.md)
- [gateway vs litellm](/compare/adaline-gateway-vs-berriai-litellm.md)
- [gateway vs 9router](/compare/adaline-gateway-vs-decolua-9router.md)
- [gateway vs litgpt](/compare/adaline-gateway-vs-lightning-ai-litgpt.md)
- [gateway vs ollama](/compare/adaline-gateway-vs-ollama-ollama.md)

## Stacks

- [The local / self-hosted LLM stack](/stacks/local-llm.md)

## Tools

- [transformers](/tools/huggingface-transformers.md) - Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models (★ 164,121) [Very active]
- [open-webui](/tools/open-webui-open-webui.md) - User-friendly AI Interface (Supports Ollama, OpenAI API, ...) (★ 148,875) [Very active]
- [ollama](/tools/ollama-ollama.md) - Get up and running with various large language models using Ollama. (★ 177,524) [Very active]
- [vllm](/tools/vllm-project-vllm.md) - A high-throughput and memory-efficient inference and serving engine for LLMs (★ 87,847) [Very active]
- [claude-mem](/tools/thedotmack-claude-mem.md) - Persistent Context Across Sessions for Every Agent (★ 91,018) [Very active]
- [unsloth](/tools/unslothai-unsloth.md) - A web UI for training and running open models locally. (★ 69,621) [Very active]
- [anything-llm](/tools/mintplex-labs-anything-llm.md) - Self-hosted agent experience with deployment scripts for multiple environments (★ 64,716) [Very active]
- [JeecgBoot](/tools/jeecgboot-jeecgboot.md) - AI低代码平台，实现快速生成前后端系统及模块 (★ 47,405) [Very active]
- [litellm](/tools/berriai-litellm.md) - Python SDK and Proxy Server for calling multiple LLM APIs (★ 55,221) [Very active]
- [OmniRoute](/tools/diegosouzapw-omniroute.md) - Free AI gateway with multi-provider support and token savings (★ 51,233) [Very active]
- [ray](/tools/ray-project-ray.md) - Ray is an AI compute engine with a core distributed runtime and AI Libraries for accelerating ML workloads. (★ 43,526) [Very active]
- [LibreChat](/tools/danny-avila-librechat.md) - Enhanced ChatGPT Clone with extensive features and integrations for self-hosting (★ 41,282) [Very active]
- [sglang](/tools/sgl-project-sglang.md) - High-performance serving framework for large language and multimodal models (★ 31,454) [Very active]
- [whisper.cpp](/tools/ggml-org-whisper-cpp.md) - Port of OpenAI's Whisper model in C/C++ for speech-to-text inference (★ 52,501) [Very active]
- [mlflow](/tools/mlflow-mlflow.md) - AI engineering platform for debugging, evaluating, monitoring, and optimizing AI applications (★ 27,591) [Very active]
- [netron](/tools/lutzroeder-netron.md) - Visualizer for neural network, deep learning and machine learning models (★ 33,302) [Very active]
- [ruflo](/tools/ruvnet-ruflo.md) - The leading agent meta-harness for intelligent multi-player swarms and autonomous workflows (★ 68,322) [Very active]
- [9router](/tools/decolua-9router.md) - Unlimited FREE AI coding with auto-fallback and token savings (★ 25,841) [Very active]
- [ncnn](/tools/tencent-ncnn.md) - High-performance neural network inference framework optimized for mobile platforms (★ 23,644) [Very active]
- [screenpipe](/tools/screenpipe-screenpipe.md) - AI that records and analyzes everything you do, say, hear locally (★ 20,534) [Very active]
- [ai-guide](/tools/liyupi-ai-guide.md) - 免费开放的AI知识共享平台 (★ 18,766) [Active]
- [jan](/tools/janhq-jan.md) - open source alternative to ChatGPT that runs offline locally (★ 44,020) [Very active]
- [ml-engineering](/tools/stas00-ml-engineering.md) - Machine Learning Engineering Open Book (★ 18,632) [Very active]
- [airllm](/tools/lyogavin-airllm.md) - AirLLM 70B inference with single 4GB GPU (★ 24,183) [Very active]
- [generative-ai](/tools/googlecloudplatform-generative-ai.md) - Sample code and notebooks for Generative AI on Google Cloud, with Gemini Enterprise Agent Platform (★ 17,594) [Very active]
- [openvino](/tools/openvinotoolkit-openvino.md) - OpenVINO is an open source toolkit for optimizing and deploying AI inference. (★ 10,566) [Very active]
- [DeepSpeed](/tools/deepspeedai-deepspeed.md) - Deep learning optimization library for efficient distributed training and inference (★ 42,870) [Very active]
- [omlx](/tools/jundot-omlx.md) - LLM inference server with continuous batching and SSD caching for Apple Silicon (★ 18,679) [Very active]
- [pytorch](/tools/pytorch-pytorch.md) - Tensors and Dynamic neural networks in Python with strong GPU acceleration (★ 102,144) [Very active]
- [OpenLLM](/tools/bentoml-openllm.md) - Run any open-source LLMs as OpenAI compatible API endpoint in the cloud. (★ 12,454) [Very active]
- [inference](/tools/xorbitsai-inference.md) - Unified production-ready inference API for various models (★ 9,470) [Very active]
- [awesome-LLM-resources](/tools/wangrongsheng-awesome-llm-resources.md) - Summary of the world's best LLM resources. (★ 8,845) [Very active]
- [oumi](/tools/oumi-ai-oumi.md) - Easily fine-tune, evaluate and deploy open source LLMs/VLMs (★ 9,359) [Very active]
- [dynamo](/tools/ai-dynamo-dynamo.md) - A Datacenter Scale Distributed Inference Serving Framework (★ 7,575) [Very active]
- [runanywhere-sdks](/tools/runanywhereai-runanywhere-sdks.md) - Production ready toolkit to run AI locally (★ 10,300) [Very active]
- [CosyVoice](/tools/funaudiollm-cosyvoice.md) - Multi-lingual large voice generation model with full-stack abilities for inference, training and deployment. (★ 22,373) [Steady]
- [train-llm-from-scratch](/tools/fareedkhan-dev-train-llm-from-scratch.md) - A straightforward method for training your LLM from raw text to aligned model generation (★ 9,141) [Very active]
- [awesome-free-llm-apis](/tools/mnfst-awesome-free-llm-apis.md) - List of Permanent Free LLM API (★ 6,532) [Active]
- [plano](/tools/katanemo-plano.md) - An AI-native proxy and data plane for agentic apps (★ 7,004) [Very active]
- [BentoML](/tools/bentoml-bentoml.md) - The easiest way to serve AI apps and models (★ 8,793) [Active]
- [serving](/tools/tensorflow-serving.md) - A flexible, high-performance serving system for machine learning models (★ 6,359) [Very active]
- [llama.cpp](/tools/ggml-org-llama-cpp.md) - LLM inference in C/C++ (★ 122,941) [Very active]
- [MNN](/tools/alibaba-mnn.md) - Blazing-fast, lightweight inference engine for high-performance on-device LLMs and Edge AI (★ 15,830) [Very active]
- [lmdeploy](/tools/internlm-lmdeploy.md) - Toolkit for compressing, deploying, and serving LLMs (★ 7,995) [Very active]
- [metaflow](/tools/netflix-metaflow.md) - Build, Manage and Deploy AI/ML Systems (★ 10,228) [Very active]
- [llamafile](/tools/mozilla-ai-llamafile.md) - Distribute and run LLMs with a single file. (★ 25,470) [Very active]
- [llm-universe](/tools/datawhalechina-llm-universe.md) - 面向小白开发者的LLM应用开发教程 (★ 13,803) [Active]
- [shell_gpt](/tools/ther1d-shell-gpt.md) - A command-line productivity tool powered by AI large language models (★ 12,235) [Steady]
- [llm-action](/tools/liguodongiot-llm-action.md) - Aims to share large model technology principles and practical experience (large model engineering, application implementation) (★ 24,898) [Active]
- [bifrost](/tools/maximhq-bifrost.md) - Fast Enterprise AI Gateway with Adaptive Load Balancer and Guardrails (★ 7,449) [Very active]
- [TensorRT-LLM](/tools/nvidia-tensorrt-llm.md) - Python API for defining and optimizing Large Language Models (LLMs) on NVIDIA GPUs (★ 14,317) [Very active]
- [mcp](/tools/awslabs-mcp.md) - Open source MCP Servers for AWS (★ 9,503) [Very active]
- [litgpt](/tools/lightning-ai-litgpt.md) - High-performance LLMs with recipes for pretraining, finetuning and deployment (★ 13,605) [Active]
- [self-llm](/tools/datawhalechina-self-llm.md) - A guide for fine-tuning and deploying open-source large language models tailored for a Chinese audience on Linux. (★ 31,722) [Active]
- [llm](/tools/simonw-llm.md) - Access large language models from the command-line (★ 12,324) [Very active]
- [pytorch-lightning](/tools/lightning-ai-pytorch-lightning.md) - Pretrain, finetune ANY AI model of ANY size on 1 or 10,000+ GPUs with zero code changes. (★ 31,267) [Very active]
- [bitsandbytes](/tools/bitsandbytes-foundation-bitsandbytes.md) - Large language model quantization toolkit for PyTorch. (★ 8,385) [Very active]
- [go-stock](/tools/arvinlovegood-go-stock.md) - AI-empowered stock analysis tool for multiple markets (★ 7,236) [Very active]
- [private-gpt](/tools/zylon-ai-private-gpt.md) - Complete API layer for private AI applications on local models (★ 57,415) [Very active]
- [llm-engineer-toolkit](/tools/kalyanks-nlp-llm-engineer-toolkit.md) - A curated list of over 120 LLM libraries categorized. (★ 10,767) [Very active]

## Common questions

### What are the best inference & serving tools?

GraphCanon ranks Inference & Serving tools by GitHub adoption and freshness. transformers is the current leader (164,121 stars). See the full list on this page - sorted by stars, with [maintenance labels](/glossary/trust-and-signals/maintenance-label) and graph relationships.

### How does GraphCanon rank Inference & Serving tools?

We sort by GitHub stars and push recency on category pages, not paid placement. Alternatives and compare pages use [typed graph edges](/glossary/knowledge-graph/typed-edge) (alternative, successor, integrates_with) plus shared categories - constraint-first, not marketing votes.

### How many tools are in Inference & Serving?

406 published tools are tagged with Inference & Serving in the GraphCanon knowledge graph.

### What are popular Inference & Serving comparisons?

Head-to-head compare pages in this category include Yi vs Qwen, Yi vs LLMs-from-scratch, 1Panel vs open-webui. Each comparison uses live GitHub stats and optional [trust signals](/glossary/trust-and-signals/trust-signal) - see the comparisons block on this page.

### Which stacks use Inference & Serving?

Curated workflow pages that include Inference & Serving: [The local / self-hosted LLM stack](/stacks/local-llm). Each stack step includes when-not-to-use guidance.

### Is there a machine-readable Inference & Serving list?

Yes. Append `.md` to this URL or fetch [`/md/categories/inference-serving`](/md/categories/inference-serving) for a markdown twin. The JSON API exposes the same corpus at [`/api/graphcanon/categories/inference-serving`](/api/graphcanon/categories/inference-serving).

---

**Machine-readable endpoints**

- JSON: [`/api/graphcanon/categories/inference-serving`](/api/graphcanon/categories/inference-serving)
- LLM index: [/llms.txt](/llms.txt)
- Full corpus: [/llms-full.txt](/llms-full.txt)

_GraphCanon - The knowledge graph for AI development. https://www.graphcanon.com/_
