Home/Compare/flash-linear-attention vs Liger-Kernel

Comparison

flash-linear-attention vs Liger-Kernel

Verdict

Pick flash-linear-attention if flash-linear-attention accelerates linear attention mechanisms in large language models, using CUDA for optimal performance; pick Liger-Kernel if optimized Triton kernels for accelerating LLM training, especially on ROCm PyTorch installations.

Markdown twin · flash-linear-attention alternatives · Liger-Kernel alternatives

GraphCanon updated 6d

flash-linear-attention logo

flash-linear-attention

fla-org/flash-linear-attention

5.6kpushed Aug 17, 2026
vs
Liger-Kernel logo

Liger-Kernel

linkedin/Liger-Kernel

6.6kpushed Aug 7, 2026

Trust & integrity

Signalflash-linear-attentionLiger-Kernel
Maintenance
Very active (0d since push)
As of 6d · github_public_v1
Very active (0d since push)
As of 2w · github_public_v1
Provenance
Not a fork · Organization account
As of 6d · github_public_v1
Not a fork · Organization account
As of 2w · github_public_v1
OSV dependency advisories
No lockfile (source not queried)
As of 1mo · osv@v1
No lockfile (source not queried)
As of 1mo · osv@v1
deps.dev advisories
Not queried
deps.dev@v1
Not queried
deps.dev@v1
OpenSSF Scorecard
Not queried
openssf-scorecard@v1
Not queried
openssf-scorecard@v1

Tagline

flash-linear-attention
🚀 Efficient implementations for emerging model architectures
Liger-Kernel
Efficient Triton Kernels for LLM Training

Stars

flash-linear-attention
5.6k
Liger-Kernel
6.6k

Forks

flash-linear-attention
661
Liger-Kernel
573

Open issues

flash-linear-attention
98
Liger-Kernel
190

Language

flash-linear-attention
Python
Liger-Kernel
Python

Adopt for

flash-linear-attention
Flash-linear-attention accelerates linear attention mechanisms in large language models, using CUDA for optimal performance.
Liger-Kernel
Optimized Triton kernels for accelerating LLM training, especially on ROCm PyTorch installations.

Persona

flash-linear-attention
-
Liger-Kernel
-

Runtime

flash-linear-attention
-
Liger-Kernel
-

License

flash-linear-attention
MIT
Liger-Kernel
BSD-2-Clause

Last pushed

flash-linear-attention
Aug 17, 2026
Liger-Kernel
Aug 7, 2026

Categories

flash-linear-attention
Model Training
Liger-Kernel
Model Training

Trust and health

Open issues (now)

flash-linear-attention
98
Liger-Kernel
190

Stars delta

flash-linear-attention
+208 (30d)
Liger-Kernel
Unknown

Open issues delta

flash-linear-attention
+21 (30d)
Liger-Kernel
Unknown

Full report

flash-linear-attention
Trust report
Liger-Kernel
Trust report

Shared compatibility

  • Python · flash-linear-attention: Python runtime · Liger-Kernel: Python runtime

Choose flash-linear-attention if…

  • License: flash-linear-attention is MIT, Liger-Kernel is BSD-2-Clause.
  • Tags unique to flash-linear-attention: large language models, machine-learning-systems, natural-language-processing, sequence-modeling.
  • High-performance requirements with Nvidia GPUs where CUDA can offer significant speed-ups

When NOT to use flash-linear-attention

  • Limited GPU hardware or no support for backend flavors like CUDA, ROCM, XPU, NPU, or CPU
  • Do not require linear attention mechanism in modeling large language models or sequence data

Choose Liger-Kernel if…

  • License: Liger-Kernel is BSD-2-Clause, flash-linear-attention is MIT.
  • Tags unique to Liger-Kernel: finetuning, gemma2, llama, mistral.
  • When enhancing training speed of large language models with ROCm-compatible hardware.

When NOT to use Liger-Kernel

  • Avoid if only CUDA environments are supported, as Liger-Kernel emphasizes ROCm compatibility.
  • Skip for simple setup requirements; prefer more streamlined tools without extensive customization options.

Explore

Sources

Every stat on this page traces to a dated GitHub sync, license file, enrichment field, or trust scan.

GitHub stars on cards: flash-linear-attention 5.6k · Liger-Kernel 6.6k (synced Aug 17, 2026).

Common questions

What is the difference between flash-linear-attention and Liger-Kernel?
flash-linear-attention: 🚀 Efficient implementations for emerging model architectures. Liger-Kernel: Efficient Triton Kernels for LLM Training. See the comparison table for live GitHub stats and shared categories.
When should I choose flash-linear-attention over Liger-Kernel?
Choose flash-linear-attention over Liger-Kernel when License: flash-linear-attention is MIT, Liger-Kernel is BSD-2-Clause; Tags unique to flash-linear-attention: large language models, machine-learning-systems, natural-language-processing, sequence-modeling; High-performance requirements with Nvidia GPUs where CUDA can offer significant speed-ups.
When should I choose Liger-Kernel over flash-linear-attention?
Choose Liger-Kernel over flash-linear-attention when License: Liger-Kernel is BSD-2-Clause, flash-linear-attention is MIT; Tags unique to Liger-Kernel: finetuning, gemma2, llama, mistral; When enhancing training speed of large language models with ROCm-compatible hardware.
When should I avoid flash-linear-attention?
Limited GPU hardware or no support for backend flavors like CUDA, ROCM, XPU, NPU, or CPU Do not require linear attention mechanism in modeling large language models or sequence data
When should I avoid Liger-Kernel?
Avoid if only CUDA environments are supported, as Liger-Kernel emphasizes ROCm compatibility. Skip for simple setup requirements; prefer more streamlined tools without extensive customization options.
Is flash-linear-attention or Liger-Kernel more popular on GitHub?
Liger-Kernel has more GitHub stars (6,555 vs 5,568). Stars measure visibility, not whether either tool fits your constraints.
Are flash-linear-attention and Liger-Kernel open source?
Yes - both are open-source projects on GitHub (flash-linear-attention: MIT, Liger-Kernel: BSD-2-Clause).
Where can I find alternatives to flash-linear-attention or Liger-Kernel?
GraphCanon lists graph-backed alternatives at flash-linear-attention alternatives and Liger-Kernel alternatives (flash-linear-attention markdown twin, Liger-Kernel markdown twin), ranked by typed relationship edges rather than popularity votes.
Is there a machine-readable version of this comparison?
Yes. The markdown twin at this comparison mirrors this page for agents and LLM crawlers, with the same stats table and FAQ answers.
Which is better maintained, flash-linear-attention or Liger-Kernel?
flash-linear-attention: Very active. Liger-Kernel: Very active. Compare maintenance labels, days since push, and release cadence in the trust section below - stars alone do not measure maintenance.
Where are the full trust reports for flash-linear-attention and Liger-Kernel?
GraphCanon publishes per-repo trust reports with dated maintenance, provenance, and scan summaries: flash-linear-attention trust report; Liger-Kernel trust report.

Was this helpful?

Anonymous feedback helps us improve pages and translations.