llama.cpp logo

llama.cpp

ggml-org/llama.cpp

LLM inference in C/C++

GraphCanon updated 2w · GitHub synced 2w · 25 views this month

123k stars21k forksLast push 2w C++ MIT

Decision brief

llama.cpp is a C++ framework for LLM inference, offering versatile installation options including package managers, Docker, and binary downloads.

Good fit when

  • - You need high-performance inference capabilities in a lightweight environment where C++ performance benefits are critical.
  • - Your deployment requires direct model quantization which can be optimized with the provided tools and framework.

Avoid when

  • - If you prefer a language other than C++, as this tool lacks support for Python or JavaScript bindings that provide higher-level abstractions.
  • - When your project demands extensive runtime customization and flexibility that is more easily achieved in languages like Python with libraries such as PyTorch.
Hosting:
unknown - llama.cpp supports various installation methods including package managers (like brew), Docker containers for isolation, pre-built binaries for ease of deployment, and source builds for flexibility.
Requirements:
Installation can be done via multiple channels including package managers, Docker, and direct downloads.

Observed Jul 11, 2026 · Source: enrich:decision_facts

Verify the decision

Maintenance and security

Full trust report
Maintenance
Very active (0d since push)
As of 2w
Provenance
Not a fork · Organization account
As of 2w
Security (OSV)
No criticals
As of 1mo

Public GitHub metadata and optional OSV scans. Signals, not a guarantee. Trust methodology.

Backing

Company context for ggml. Display-only - separate from trust and ranking.

Company
ggml·GitHub org profile·1mo
Commercial model
Pure OSS·GitHub org profile (public repos)·1mo

Install

git clone https://github.com/ggml-org/llama.cpp

How it fits your stack(18)

Typed graph edges - alternatives, integrations, successors, and dependencies. Ranked by relationship type, not raw GitHub stars.

Alternative

Relationship graph

Optional deeper exploration of typed edges and category neighbours.

Similar tools

Same-category neighbours not already linked as typed edges.

Evidence and technical details

Sourced facts, taxonomy, compatibility claims, README excerpt, and machine-readable endpoints.

Overview

llama.cpp provides a framework for LLM inference using C++. It supports installation via package managers, Docker, pre-built binaries, and source builds.

Capability facts

CLI
CLI entrypoint

Source: pyproject.toml:[project.scripts] · Aug 7, 2026

Languages
c++, python

Source: github.language+pyproject.toml · Aug 7, 2026

Categories

Graph entities

Tags

README

Quick start

A few options to get llama.cpp installed on your machine:

  • Visit https://llama.app and follow the instructions
  • Run with Docker - see our Docker documentation
  • Download pre-built binaries from the releases page
  • Build from source by cloning this repository - check out our build guide

Once installed:

For agents

This page has a .md twin and JSON over the API.

Was this helpful?

Anonymous feedback helps us improve pages and translations.