Home/Inference & Serving/llm-inference-solutions
llm-inference-solutions logo

llm-inference-solutions

mani-kantap/llm-inference-solutions

A collection of all available inference solutions for the LLMs

GraphCanon updated 2w · GitHub synced 2w · 29 views this month

95 stars7 forksLast push 1y MIT

Decision brief

Curated listings of tools for efficient inference and deployment of LLMs with details on hardware support, features, and licenses.

Good fit when

  • Need a comprehensive catalog to compare multiple inference solutions for LLMs like vLLM's memory management or Triton Inference Server's framework diversity
  • Require insights into different licensing options like MIT or Apache 2.0 across tools

Avoid when

  • Looking for direct technical implementation details instead of a curated list, as it primarily serves as an overview repository
  • In need of real-time updates since the repository's content may not be continuously updated to reflect new developments in inference solutions

Observed Jul 12, 2026 · Source: enrich:decision_facts

Verify the decision

Maintenance and security

Full trust report
Maintenance
Dormant (523d since push)
As of 2w
Provenance
Not a fork · Personal account
As of 2w
Security (OSV)
No lockfile
As of 1mo

Public GitHub metadata and optional OSV scans. Signals, not a guarantee. Trust methodology.

Install

git clone https://github.com/mani-kantap/llm-inference-solutions

Similar tools

Same-category neighbours. No typed graph edges are catalogued for this tool yet.

Evidence and technical details

Sourced facts, taxonomy, compatibility claims, README excerpt, and machine-readable endpoints.

Overview

This repository lists various tools and frameworks used for efficient inference and serving of large language models (LLMs). It provides an overview including supported hardware, key features, and licenses.

Capability facts

No sourced capability facts yet. Facts appear after ingest scans repo manifests (Dockerfile, package.json, MCP configs).

Categories

Compatibility

Sourced claims from the README excerpt - not unsourced marketing copy.

Python runtimePython

Source: README excerpt (regex_v1, Aug 7, 2026)

//github.com/NVIDIA/TensorRT-LLM) | NVIDIA | Provides users with an easy-to-use Python API to define LLMs and build TensorRT engines. | GPU | TensorRT optimization, hi
Source link

Tags

README

llm-inference-solutions

A collection of all available inference solutions for the LLMs

NameOrganizationDescriptionSupported HardwareKey FeaturesLicense
vLLMUC BerkeleyHigh-throughput and memory-efficient inference and serving engine for LLMs.CPU, GPUPagedAttention for optimized memory management, high-throughput serving.Apache 2.0
Text-Generation-InferenceHugging Face 🤗Efficient and scalable text generation inference for LLMs.CPU, GPUMulti-model serving, dynamic batching, optimized for transformers.Apache 2.0
llm-engineScale AIScale LLM Engine public repository for efficient inference.CPU, GPUScalable deployment, monitoring tools, integration with Scale AI services.Apache 2.0
DeepSpeedMicrosoftDeep learning optimization library for easy, efficient, and effective distributed training and inference.CPU, GPUZeRO redundancy optimizer, mixed-precision training, model parallelism.MIT
OpenLLMBentoMLOperating LLMs in production with ease.CPU, GPUModel serving, deployment orchestration, integration with BentoML.Apache 2.0
LMDeployInternLM TeamToolkit for compressing, deploying, and serving LLMs.CPU, GPUModel compression, deployment automation, serving optimization.Apache 2.0
FlexFlowCMU, Stanford, UCSDA distributed deep learning framework.CPU, GPU, TPUAutomatic parallelization, support for complex models, scalability.Apache 2.0
CTranslate2OpenNMTFast inference engine for Transformer models.CPU, GPUInt8 quantization, multi-threaded execution, optimized for translation models.MIT
FastChatlm-sysOpen platform for training, serving, and evaluating large language models; release repo for Vicuna and Chatbot Arena.CPU, GPUChatbot framework, multi-turn conversations, evaluation tools.Apache 2.0
Triton Inference ServerNVIDIAOptimized cloud and edge inferencing solution.CPU, GPUModel ensemble, dynamic batching, support for multiple frameworks.BSD-3-Clause
Lepton.AIlepton.aiPythonic framework to simplify AI service building.CPU, GPUService orchestration, API generation, scalability.MIT
ScaleLLMVectorchHigh-performance inference system for LLMs, designed for production environments.CPU, GPULow-latency serving, high throughput, production-ready.Apache 2.0
LoraxPredibaseServe hundreds of fine-tuned LLMs in production for the cost of one.CPU, GPUModel multiplexing, cost-efficient serving, scalability.Apache 2.0
TensorRT-LLMNVIDIAProvides users with an easy-to-use Python API to define LLMs and build TensorRT engines.GPUTensorRT optimization, high-performance inference, integration with NVIDIA GPUs.Apache 2.0
mistral.rsmistral.rsBlazingly fast LLM inference.CPU, GPURust-based implementation, performance optimization, lightweight.MIT
NanoFlowNanoFlowThroughput-oriented high-performance serving framework for LLMs.CPU, GPUHigh throughput, low latency, optimized for large-scale deployments.Apache 2.0
[LMCache](https://gi

For agents

This page has a .md twin and JSON over the API.

Was this helpful?

Anonymous feedback helps us improve pages and translations.