xllm logo

xllm

xLLM-AI/xllm

A high-performance inference engine for LLM, VLM, DiT and REC models

GraphCanon updated 1mo · GitHub synced 1mo

1.5k stars269 forksLast push 1mo C++ Apache-2.0

Decision brief

A high-performance inference engine for LLM, VLM, DiT, and REC models by the OpenAtom Foundation.

Good fit when

  • When developing applications that require optimized performance on various AI accelerators
  • For projects involving deepseek, glm, Qwen large language models

Avoid when

  • If your project strictly requires Python-based inference engines for backend support
  • In cases preferring proprietary licenses over the Apache-2.0 open-source framework used here

Observed Jul 14, 2026 · Source: enrich:decision_facts

Verify the decision

Maintenance and security

Full trust report
Maintenance
Very active (0d since push)
As of 1mo
Provenance
Not a fork · Organization account
As of 1mo
Security (OSV)
No lockfile
As of 1mo

Public GitHub metadata and optional OSV scans. Signals, not a guarantee. Trust methodology.

Install

git clone https://github.com/xLLM-AI/xllm

Similar tools

Same-category neighbours. No typed graph edges are catalogued for this tool yet.

Evidence and technical details

Sourced facts, taxonomy, compatibility claims, README excerpt, and machine-readable endpoints.

Overview

Hosted by the OpenAtom Foundation, this C++-based repository offers optimized performances across diverse AI accelerators.

Capability facts

CLI
CLI entrypoint

Source: pyproject.toml:[project.scripts] · Jul 25, 2026

Languages
c++, python

Source: github.language+pyproject.toml · Jul 25, 2026

Categories

Tags

README

Hardware Support

HardwareAbbreviationExampleRemark
Ascend NPUNPUA2, A3HDK Driver 25.2.0 +
Cambricon MLUMLUMLU
Moore Threads GPUMUSAS5000
Hygon DCUDCUBW1000
MetaX MACAMACAMXC500
Iluvatar CoreX GPUILUBI150

Getting Started

For agents

This page has a .md twin and JSON over the API.

Was this helpful?

Anonymous feedback helps us improve pages and translations.