GraphCanon updated 3w · GitHub synced 3w
Decision brief
NanoLLM optimizes local inference for LLMs via HuggingFace-compatible APIs, supporting quantization and multimodal applications like vision, speech, RAG, and vector databases.
Good fit when
- When building edge-ai solutions requiring optimized local inference
- For multimodal AI projects needing integrations with speech and image models
Avoid when
- In scenarios where a fully cloud-based solution is preferred over local inference
- If the project does not benefit from multimodal or RAG capabilities
Observed Jul 17, 2026 · Source: enrich:decision_facts
Verify the decision
Maintenance and security
Full trust report- Maintenance
- Dormant (645d since push)
- As of 3w
- Provenance
- Not a fork · Personal account
- As of 3w
- Security (OSV)
- No lockfile
- As of 1mo
Public GitHub metadata and optional OSV scans. Signals, not a guarantee. Trust methodology.
Install
pip install NanoLLM PyPIHow it fits your stack(1)
Typed graph edges - alternatives, integrations, successors, and dependencies. Ranked by relationship type, not raw GitHub stars.
Relationship graph
Optional deeper exploration of typed edges and category neighbours.
Similar tools
Same-category neighbours not already linked as typed edges.
Evidence and technical details
Sourced facts, taxonomy, compatibility claims, README excerpt, and machine-readable endpoints.
Overview
NanoLLM supports optimized local inference with HuggingFace-compatible interfaces for tasks such as quantization and multimodal applications including vision, language, speech, vector databases, and RAG.
Capability facts
- Languages
- python
Source: github.language · Jul 26, 2026
Categories
Tags
README
NanoLLM
Optimized local inference for LLMs with HuggingFace-like APIs for quantization, vision/language models, multimodal agents, speech, vector DB, and RAG.
[!NOTE]
Seedusty-nv.github.io/NanoLLMfor docs and Jetson AI Lab for tutorials.
Latest Release: 24.7 (dustynv/nano_llm:24.7-r36.2.0)
For agents
This page has a .md twin and JSON over the API.
