mistral.rs logo

mistral.rs

EricLBuehler/mistral.rs

Fast flexible LLM inference

GraphCanon updated 2w · GitHub synced 2w

7.6k stars671 forksLast push 3w Rust MIT

Decision brief

Mistral.rs is ideal for developers requiring fast and flexible LLM inference with support across multiple platforms. It provides prebuilt binaries and a simple installation process.

Good fit when

  • Mistral.rs should be used when seeking Rust-based implementation that supports quick and flexible deployment of large language models, particularly on Linux, macOS, or Windows systems
  • When you need to simplify setup without requiring the Rust compiler or CUDA toolkit for initial use

Avoid when

  • Avoid Mistral.rs if your project is strictly dependent on another programming language framework as it is implemented in Rust
  • If needing tight control over model-specific optimizations not provided by default prebuild paths, then consider alternatives with extensive fine-tuning options out-of-the-box

Observed Jul 14, 2026 · Source: enrich:decision_facts

Verify the decision

Maintenance and security

Full trust report
Maintenance
Active (8d since push)
As of 2w
Provenance
Not a fork · Personal account
As of 2w
Security (OSV)
No lockfile
As of 1mo

Public GitHub metadata and optional OSV scans. Signals, not a guarantee. Trust methodology.

Install

cargo add mistral.rs
crates.io

Similar tools

Same-category neighbours. No typed graph edges are catalogued for this tool yet.

Evidence and technical details

Sourced facts, taxonomy, compatibility claims, README excerpt, and machine-readable endpoints.

Overview

A Rust implementation for fast and flexible large language model (LLM) inference, supporting prebuilt binaries for multiple platforms including Linux, macOS, and Windows.

Capability facts

Deploy
Self-host

Source: dockerfile:Dockerfile · Aug 7, 2026

Docker
Dockerfile present

Source: dockerfile:Dockerfile · Aug 7, 2026

Languages
rust

Source: github.language · Aug 7, 2026

Categories

Tags

README

Install

Linux/macOS:

curl --proto '=https' --tlsv1.2 -sSf https://raw.githubusercontent.com/EricLBuehler/mistral.rs/master/install.sh | sh

Windows (PowerShell):

irm https://raw.githubusercontent.com/EricLBuehler/mistral.rs/master/install.ps1 | iex

Downloads a self-contained prebuilt binary for your platform (Metal on Apple Silicon; per-GPU CUDA or CPU on Linux; CPU on Windows), falling back to a source build if none matches. Standard acceleration needs no Rust or CUDA toolkit. Optional cuTile acceleration requires NVIDIA's separately installed tileiras tool.

Manual installation, accelerator details & other platforms


Recommend settings for your hardware and emit a config file

mistralrs tune -m Qwen/Qwen3-4B --emit-config config.toml


Docker

Prebuilt CPU and CUDA images are published to GHCR. Pull commands, tags, and Kubernetes notes are in the Docker guide.

For agents

This page has a .md twin and JSON over the API.

Was this helpful?

Anonymous feedback helps us improve pages and translations.