Home/Compare/awesome-evals vs lmms-eval

Comparison

awesome-evals vs lmms-eval

Verdict

Pick awesome-evals if curated resources for AI agent evaluation with BenchFlow backing its maintenance; pick lmms-eval if lmms-eval is a one-stop solution for benchmarking multimodal large language models across various tasks including text, image, video, and audio.

Markdown twin · awesome-evals alternatives · lmms-eval alternatives

GraphCanon updated 3d

awesome-evals logo

awesome-evals

benchflow-ai/awesome-evals

761pushed Jul 1, 2026
vs
lmms-eval logo

lmms-eval

EvolvingLMMs-Lab/lmms-eval

4.4kpushed Aug 6, 2026

Trust & integrity

Signalawesome-evalslmms-eval
Maintenance
Active (26d since push)
As of 3w · github_public_v1
Active (11d since push)
As of 3d · github_public_v1
Provenance
Not a fork · Organization account
As of 3w · github_public_v1
Not a fork · Organization account
As of 3d · github_public_v1
OSV dependency advisories
No lockfile (source not queried)
As of 1mo · osv@v1
No lockfile (source not queried)
As of 1mo · osv@v1
deps.dev advisories
Not queried
deps.dev@v1
Not queried
deps.dev@v1
OpenSSF Scorecard
Not queried
openssf-scorecard@v1
Not queried
openssf-scorecard@v1

Tagline

awesome-evals
A curated library of resources for building and evaluating AI agents
lmms-eval
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks

Stars

awesome-evals
761
lmms-eval
4.4k

Forks

awesome-evals
71
lmms-eval
639

Open issues

awesome-evals
21
lmms-eval
49

Language

awesome-evals
-
lmms-eval
Python

Adopt for

awesome-evals
Curated resources for AI agent evaluation with BenchFlow backing its maintenance
lmms-eval
lmms-eval is a one-stop solution for benchmarking multimodal large language models across various tasks including text, image, video, and audio.

Persona

awesome-evals
-
lmms-eval
-

Runtime

awesome-evals
-
lmms-eval
-

License

awesome-evals
Other
lmms-eval
Other

Last pushed

awesome-evals
Jul 1, 2026
lmms-eval
Aug 6, 2026

Categories

awesome-evals
AI Agents, Evaluation & Observability
lmms-eval
Evaluation & Observability

Trust and health

Days since push

awesome-evals
26d
lmms-eval
11d

Open issues (now)

awesome-evals
21
lmms-eval
49

Stars delta

awesome-evals
Unknown
lmms-eval
+52 (30d)

Open issues delta

awesome-evals
Unknown
lmms-eval
+9 (30d)

Full report

awesome-evals
Trust report
lmms-eval
Trust report

Choose awesome-evals if…

  • Tags unique to awesome-evals: agent-evaluation, ai-agents, awesome-list, benchmarks.
  • Also covers AI Agents.
  • Need diverse resources encompassing papers, blogs, talks, tools, and benchmarks specifically curated for AI agent evaluation

When NOT to use awesome-evals

  • Require real-time interactive support or direct tool integrations not covered by a static resource list
  • Seeking proprietary tools from specific vendors rather than open resources and community content

Choose lmms-eval if…

  • Tags unique to lmms-eval: agi, audio-evaluation, benchmark, evaluation.
  • You need to evaluate LLaVA series models on different datasets with precise control over reproducibility details like torch/cuda versions.
  • More GitHub stars (4.4k vs 761) - visibility, not fit.

When NOT to use lmms-eval

  • Looking for a tool that supports less than Python 3.12, as uv setup mandates this version.
  • Requiring support beyond text, image, video, and audio modalities which lmms-eval specifically covers.
  • Your project doesn't benefit from extensive results tracking in Google Sheets or relies solely on alternative reproducibility mechanisms without external dependencies.

Explore

Sources

Every stat on this page traces to a dated GitHub sync, license file, enrichment field, or trust scan.

GitHub stars on cards: awesome-evals 761 · lmms-eval 4.4k (synced Jul 28, 2026).

Common questions

What is the difference between awesome-evals and lmms-eval?
awesome-evals: A curated library of resources for building and evaluating AI agents. lmms-eval: One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks. See the comparison table for live GitHub stats and shared categories.
When should I choose awesome-evals over lmms-eval?
Choose awesome-evals over lmms-eval when Tags unique to awesome-evals: agent-evaluation, ai-agents, awesome-list, benchmarks; Also covers AI Agents; Need diverse resources encompassing papers, blogs, talks, tools, and benchmarks specifically curated for AI agent evaluation.
When should I choose lmms-eval over awesome-evals?
Choose lmms-eval over awesome-evals when Tags unique to lmms-eval: agi, audio-evaluation, benchmark, evaluation; You need to evaluate LLaVA series models on different datasets with precise control over reproducibility details like torch/cuda versions; More GitHub stars (4.4k vs 761) - visibility, not fit.
When should I avoid awesome-evals?
Require real-time interactive support or direct tool integrations not covered by a static resource list Seeking proprietary tools from specific vendors rather than open resources and community content
When should I avoid lmms-eval?
Looking for a tool that supports less than Python 3.12, as uv setup mandates this version. Requiring support beyond text, image, video, and audio modalities which lmms-eval specifically covers. Your project doesn't benefit from extensive results tracking in Google Sheets or relies solely on alternative reproducibility mechanisms without external dependencies.
Is awesome-evals or lmms-eval more popular on GitHub?
lmms-eval has more GitHub stars (4,368 vs 761). Stars measure visibility, not whether either tool fits your constraints.
Are awesome-evals and lmms-eval open source?
Yes - both are open-source projects on GitHub (awesome-evals: Other, lmms-eval: Other).
Where can I find alternatives to awesome-evals or lmms-eval?
GraphCanon lists graph-backed alternatives at awesome-evals alternatives and lmms-eval alternatives (awesome-evals markdown twin, lmms-eval markdown twin), ranked by typed relationship edges rather than popularity votes.
Is there a machine-readable version of this comparison?
Yes. The markdown twin at this comparison mirrors this page for agents and LLM crawlers, with the same stats table and FAQ answers.
Which is better maintained, awesome-evals or lmms-eval?
awesome-evals: Active. lmms-eval: Active. Compare maintenance labels, days since push, and release cadence in the trust section below - stars alone do not measure maintenance.
Where are the full trust reports for awesome-evals and lmms-eval?
GraphCanon publishes per-repo trust reports with dated maintenance, provenance, and scan summaries: awesome-evals trust report; lmms-eval trust report.

Was this helpful?

Anonymous feedback helps us improve pages and translations.