Home/Inference & Serving/Awesome-LLM-Compression
Awesome-LLM-Compression logo

Awesome-LLM-Compression

HuangOwen/Awesome-LLM-Compression

Awesome LLM compression research papers and tools to accelerate LLM training and inference.

GraphCanon updated 1w · GitHub synced 1w

1.9k stars129 forksLast push 1mo MIT

Decision brief

Awesome LLM-Compression curates a comprehensive collection of research papers and tools aimed at compressing large language models, focusing on enhancing computational efficiency during both training and serving phases.

Good fit when

  • When you need to explore the latest advancements in LLM compression techniques and their impact on both training and inference.
  • If your project requires a detailed survey of model compression for large language models covering various aspects like quantization, pruning, distillation, and efficient prompting.

Avoid when

  • Avoid relying solely on Awesome LLM-Compression if you require a hands-on toolset rather than theoretical frameworks and research papers, as it focuses more on consolidating the survey information.
  • If your immediate need is for proprietary or commercial tools that offer out-of-the-box functionality, since this resource mainly links to academic research and open-source projects.
Requirements:
The repository provides curated listings but does not develop its own software; hence specific language requirements are not applicable.

Observed Jul 11, 2026 · Source: enrich:decision_facts

Verify the decision

Maintenance and security

Full trust report
Maintenance
Steady (37d since push)
As of 1w
Provenance
Not a fork · Personal account
As of 1w
Security (OSV)
No lockfile
As of 1mo

Public GitHub metadata and optional OSV scans. Signals, not a guarantee. Trust methodology.

Install

git clone https://github.com/HuangOwen/Awesome-LLM-Compression

How it fits your stack(1)

Typed graph edges - alternatives, integrations, successors, and dependencies. Ranked by relationship type, not raw GitHub stars.

Relationship graph

Optional deeper exploration of typed edges and category neighbours.

Similar tools

Same-category neighbours not already linked as typed edges.

Evidence and technical details

Sourced facts, taxonomy, compatibility claims, README excerpt, and machine-readable endpoints.

Overview

Compilation of research papers and tools focused on compressing large language models for improved computational efficiency during both training and serving phases.

Capability facts

No sourced capability facts yet. Facts appear after ingest scans repo manifests (Dockerfile, package.json, MCP configs).

Categories

Tags

README

Awesome LLM Compression

Badge image Badge image Badge image

Awesome LLM compression research papers and tools to accelerate LLM training and inference.

Contents

  • 📑 Papers
    • Survey
    • Quantization
    • Pruning and Sparsity
    • Distillation
    • Efficient Prompting
    • KV Cache Compression
    • Other
  • 🔧 Tools
  • 🙌 Contributing
  • 🌟 Star History

Papers

Survey

  • Compressed but Compromised? A Study of Jailbreaking in Compressed LLMs
    NeurIPS Lock-LLM Workshop 2025 [Paper] [[Blog]] (https://namburisrinath.medium.com/compressed-but-compromised-a-study-of-jailbreaking-in-compressed-llms-02a6e40aaf17)

  • A Survey on Model Compression for Large Language Models
    TACL [Paper]

  • The Cost of Compression: Investigating the Impact of Compression on Parametric Knowledge in Language Models
    EMNLP 2023 [Paper] [Code]

  • The Efficiency Spectrum of Large Language Models: An Algorithmic Survey
    Arxiv 2023 [Paper]

  • Efficient Large Language Models: A Survey
    TMLR [Paper] [GitHub Page]

  • Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
    ICML 2024 Tutorial [Paper] [Tutorial]

  • Understanding LLMs: A Comprehensive Overview from Training to Inference
    Arxiv 2024 [Paper]

  • Faster and Lighter LLMs: A Survey on Current Challenges and Way Forward
    IJCAI 2024 (Survey Track) [Paper] [GitHub Page]

  • A Survey of Resource-efficient LLM and Multimodal Foundation Models
    Arxiv 2024 [Paper]

  • A Survey on Hardware Accelerators for Large Language Models
    Arxiv 2024 [Paper]

  • A Comprehensive Survey of Compression Algorithms for Language Models
    Arxiv 2024 [Paper]

  • A Survey on Transformer Compression
    Arxiv 2024 [Paper]

  • Model Compression and Efficient Inference for Large Language Models: A Survey
    Arxiv 2024 [Paper]

  • LLM Inference Unveiled: Survey and Roofline Model Insights
    Arxiv 2024 [Paper]

  • A Survey on Knowledge Distillation of Large Language Models
    Arxiv 2024 [Paper] [GitHub Page]

  • Efficient Prompting Methods for Large Language Models: A Survey
    Arxiv 2024 [Paper]

  • Survey on Knowledge Distillation for Large Language Models: Methods, Evaluation, and Application
    Arxiv 2024 [Paper]

  • On-Device Language Models: A Comprehensive Review
    Arxiv 2024 [Paper] [GitHub Page] [Download On-device LLMs]

  • A Survey of Low-bit Large Language Models: Basics, Systems, and Algorithms
    Arxiv 2024 [Paper]

  • Contextual Com

For agents

This page has a .md twin and JSON over the API.

Was this helpful?

Anonymous feedback helps us improve pages and translations.