GraphCanon updated today · GitHub synced today · 26 views this month
Decision brief
Sarathi Serve targets efficient low-latency and high-throughput inference for LLMs using Python.
Good fit when
- Optimize Python-based projects needing quick responses from large language models.
- Demand high throughput alongside minimal inference delay.
Avoid when
- Necessitate a non-Python environment for deployment and operation.
- Prefer a tool that incorporates more than just low-latency, high-throughput focus such as multi-language support or specialized optimizations.
Observed Jul 17, 2026 · Source: enrich:decision_facts
Verify the decision
Maintenance and security
Full trust report- Maintenance
- Slowing (229d since push)
- As of today
- Provenance
- Not a fork · Organization account
- As of today
- Security (OSV)
- No lockfile
- As of 1mo
Public GitHub metadata and optional OSV scans. Signals, not a guarantee. Trust methodology.
Backing
Company context for Microsoft. Display-only - separate from trust and ranking.
- Company
- Microsoft·GitHub org profile·1mo
- Employees
- 221,000·Wikidata (P1128 employees)·1mo
- Commercial model
- Pure OSS·GitHub org profile (public repos)·1mo
Install
pip install sarathi-serve PyPISimilar tools
Same-category neighbours. No typed graph edges are catalogued for this tool yet.
Evidence and technical details
Sourced facts, taxonomy, compatibility claims, README excerpt, and machine-readable endpoints.
Overview
Sarathi Serve is a Python-based project designed to provide efficient inference services with emphasis on low latency and high throughput for large language models.
Capability facts
- Languages
- python
Source: github.language+pyproject.toml · Aug 25, 2026
Categories
Graph entities
Compatibility
Sourced claims from the README excerpt - not unsourced marketing copy.
Tags
README
Install Sarathi-Serve
pip install -e .
For agents
This page has a .md twin and JSON over the API.