{"data":{"slug":"mani-kantap-llm-inference-solutions","name":"llm-inference-solutions","tagline":"A collection of all available inference solutions for the LLMs","github_url":"https://github.com/mani-kantap/llm-inference-solutions","owner":"mani-kantap","repo":"llm-inference-solutions","owner_avatar_url":"https://avatars.githubusercontent.com/u/31481823?v=4","primary_language":null,"stars":95,"forks":7,"topics":["llm-inference","llm-serving","llmops"],"archived":false,"github_pushed_at":"2025-03-01T13:49:13+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/mani-kantap-llm-inference-solutions","markdown_url":"https://www.graphcanon.com/tools/mani-kantap-llm-inference-solutions.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/mani-kantap-llm-inference-solutions","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=mani-kantap-llm-inference-solutions","description":"A collection of all available inference solutions for the LLMs","homepage_url":null,"license":"MIT","open_issues":1,"watchers":3,"ai_summary":"This repository lists various tools and frameworks used for efficient inference and serving of large language models (LLMs). It provides an overview including supported hardware, key features, and licenses.","readme_excerpt":"# llm-inference-solutions\nA collection of all available inference solutions for the LLMs\n\n| Name | Organization | Description | Supported Hardware | Key Features | License |\n|------|--------------|-------------|--------------------|--------------|---------|\n| [vLLM](https://github.com/vllm-project/vllm) | UC Berkeley | High-throughput and memory-efficient inference and serving engine for LLMs. | CPU, GPU | PagedAttention for optimized memory management, high-throughput serving. | Apache 2.0 |\n| [Text-Generation-Inference](https://github.com/huggingface/text-generation-inference) | Hugging Face 🤗 | Efficient and scalable text generation inference for LLMs. | CPU, GPU | Multi-model serving, dynamic batching, optimized for transformers. | Apache 2.0 |\n| [llm-engine](https://github.com/scaleapi/llm-engine) | Scale AI | Scale LLM Engine public repository for efficient inference. | CPU, GPU | Scalable deployment, monitoring tools, integration with Scale AI services. | Apache 2.0 |\n| [DeepSpeed](https://github.com/microsoft/DeepSpeed) | Microsoft | Deep learning optimization library for easy, efficient, and effective distributed training and inference. | CPU, GPU | ZeRO redundancy optimizer, mixed-precision training, model parallelism. | MIT |\n| [OpenLLM](https://github.com/bentoml/OpenLLM) | BentoML | Operating LLMs in production with ease. | CPU, GPU | Model serving, deployment orchestration, integration with BentoML. | Apache 2.0 |\n| [LMDeploy](https://github.com/InternLM/lmdeploy) | InternLM Team | Toolkit for compressing, deploying, and serving LLMs. | CPU, GPU | Model compression, deployment automation, serving optimization. | Apache 2.0 |\n| [FlexFlow](https://github.com/flexflow/FlexFlow) | CMU, Stanford, UCSD | A distributed deep learning framework. | CPU, GPU, TPU | Automatic parallelization, support for complex models, scalability. | Apache 2.0 |\n| [CTranslate2](https://github.com/OpenNMT/CTranslate2) | OpenNMT | Fast inference engine for Transformer models. | CPU, GPU | Int8 quantization, multi-threaded execution, optimized for translation models. | MIT |\n| [FastChat](https://github.com/lm-sys/FastChat) | lm-sys | Open platform for training, serving, and evaluating large language models; release repo for Vicuna and Chatbot Arena. | CPU, GPU | Chatbot framework, multi-turn conversations, evaluation tools. | Apache 2.0 |\n| [Triton Inference Server](https://github.com/triton-inference-server/server) | NVIDIA | Optimized cloud and edge inferencing solution. | CPU, GPU | Model ensemble, dynamic batching, support for multiple frameworks. | BSD-3-Clause |\n| [Lepton.AI](https://github.com/leptonai/leptonai) | lepton.ai | Pythonic framework to simplify AI service building. | CPU, GPU | Service orchestration, API generation, scalability. | MIT |\n| [ScaleLLM](https://github.com/vectorch-ai/ScaleLLM) | Vectorch | High-performance inference system for LLMs, designed for production environments. | CPU, GPU | Low-latency serving, high throughput, production-ready. | Apache 2.0 |\n| [Lorax](https://predibase.com/blog/lorax-the-open-source-framework-for-serving-100s-of-fine-tuned-llms-in) | Predibase | Serve hundreds of fine-tuned LLMs in production for the cost of one. | CPU, GPU | Model multiplexing, cost-efficient serving, scalability. | Apache 2.0 |\n| [TensorRT-LLM](https://github.com/NVIDIA/TensorRT-LLM) | NVIDIA | Provides users with an easy-to-use Python API to define LLMs and build TensorRT engines. | GPU | TensorRT optimization, high-performance inference, integration with NVIDIA GPUs. | Apache 2.0 |\n| [mistral.rs](https://github.com/EricLBuehler/mistral.rs) | mistral.rs | Blazingly fast LLM inference. | CPU, GPU | Rust-based implementation, performance optimization, lightweight. | MIT |\n| [NanoFlow](https://github.com/efeslab/Nanoflow) | NanoFlow | Throughput-oriented high-performance serving framework for LLMs. | CPU, GPU | High throughput, low latency, optimized for large-scale deployments. | Apache 2.0 |\n| [LMCache](https://gi","github_created_at":"2023-07-23T20:39:23+00:00","created_at":"2026-07-11T10:36:41.864128+00:00","updated_at":"2026-08-07T06:01:03.130852+00:00","categories":[{"slug":"inference-serving","name":"Inference & Serving","url":"https://www.graphcanon.com/categories/inference-serving","markdown_url":"https://www.graphcanon.com/categories/inference-serving.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/inference-serving"}],"tags":[{"slug":"llm-inference","name":"llm-inference"},{"slug":"llm-serving","name":"llm-serving"},{"slug":"llmops","name":"llmops"}],"trust":{"provenance":{"is_fork":false,"github_id":669911538,"owner_type":"User","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-07T06:01:02.390Z","maintenance":{"label":"Dormant","score":18,"methodology":"github_public_v1","releases_90d":0,"days_since_push":523,"last_release_at":null},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-11T10:36:43.126Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-07T06:01:02.859Z"},"license_spdx":{"value":"MIT","source":"github.license","observed_at":"2026-08-07T06:01:02.859Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":null,"constraints":null,"when_to_use":["Need a comprehensive catalog to compare multiple inference solutions for LLMs like vLLM's memory management or Triton Inference Server's framework diversity","Require insights into different licensing options like MIT or Apache 2.0 across tools"],"when_not_to_use":["Looking for direct technical implementation details instead of a curated list, as it primarily serves as an overview repository","In need of real-time updates since the repository's content may not be continuously updated to reflect new developments in inference solutions"],"source":"enrich:decision_facts","observed_at":"2026-07-12T11:57:17.214Z"},"constraint_facets":null,"decision_summary":[{"label":"Adopt for","value":"Curated listings of tools for efficient inference and deployment of LLMs with details on hardware support, features, and licenses."}]}}