Home/Compare/lm-evaluation-harness vs jiwer

Comparison

lm-evaluation-harness vs jiwer

Verdict

Pick lm-evaluation-harness if lm-evaluation-harness is a Python framework for evaluating language models in various parallelism modes using different checkpoint formats, compatible with the Megatron-LM backend; pick jiwer if a Python library for assessing speech-to-text systems with word error rate measures.

Markdown twin · lm-evaluation-harness alternatives · jiwer alternatives

GraphCanon updated 2w

lm-evaluation-harness logo

lm-evaluation-harness

EleutherAI/lm-evaluation-harness

14kpushed Jul 13, 2026
vs
jiwer logo

jiwer

jitsi/jiwer

917pushed Apr 16, 2026

Trust & integrity

Signallm-evaluation-harnessjiwer
Maintenance
Active (24d since push)
As of 2w · github_public_v1
Slowing (105d since push)
As of 3w · github_public_v1
Provenance
Not a fork · Organization account
As of 2w · github_public_v1
Not a fork · Organization account
As of 3w · github_public_v1
OSV dependency advisories
No lockfile (source not queried)
As of 1mo · osv@v1
No lockfile (source not queried)
As of 1mo · osv@v1
deps.dev advisories
Not queried
deps.dev@v1
Not queried
deps.dev@v1
OpenSSF Scorecard
Not queried
openssf-scorecard@v1
Not queried
openssf-scorecard@v1

Tagline

lm-evaluation-harness
A framework for few-shot evaluation of language models.
jiwer
Evaluate speech-to-text systems with word error rate (WER) measures

Stars

lm-evaluation-harness
14k
jiwer
917

Forks

lm-evaluation-harness
3.5k
jiwer
107

Open issues

lm-evaluation-harness
938
jiwer
19

Language

lm-evaluation-harness
Python
jiwer
Python

Adopt for

lm-evaluation-harness
lm-evaluation-harness is a Python framework for evaluating language models in various parallelism modes using different checkpoint formats, compatible with the Megatron-LM backend.
jiwer
A Python library for assessing speech-to-text systems with word error rate measures.

Persona

lm-evaluation-harness
-
jiwer
-

Runtime

lm-evaluation-harness
-
jiwer
-

License

lm-evaluation-harness
MIT
jiwer
Apache-2.0

Last pushed

lm-evaluation-harness
Jul 13, 2026
jiwer
Apr 16, 2026

Categories

lm-evaluation-harness
Evaluation & Observability
jiwer
Evaluation & Observability

Trust and health

Maintenance

lm-evaluation-harness
Active (82%)
jiwer
Slowing (36%)

Days since push

lm-evaluation-harness
24d
jiwer
105d

Open issues (now)

lm-evaluation-harness
938
jiwer
19

Full report

lm-evaluation-harness
Trust report

Shared compatibility

  • Python · lm-evaluation-harness: Python runtime · jiwer: Python runtime

Choose lm-evaluation-harness if…

  • License: lm-evaluation-harness is MIT, jiwer is Apache-2.0.
  • Tags unique to lm-evaluation-harness: data-parallelism, evaluation-framework, expert-parallelism, language-model.
  • - When you need to evaluate large language models across multiple GPUs in data or tensor parallel configurations.

When NOT to use lm-evaluation-harness

  • - If your evaluation setup requires pipeline parallelism not currently supported by this framework.

Choose jiwer if…

  • License: jiwer is Apache-2.0, lm-evaluation-harness is MIT.
  • Tags unique to jiwer: automatic-speech-recognition, evaluation-metrics, python3, speech-to-text.
  • For accurate speech recognition evaluation needing direct WER calculations

When NOT to use jiwer

  • Avoid if your project requires only basic, non-word level performance indicators
  • Not suitable for evaluating text-to-speech systems, as it focuses exclusively on speech-to-text output assessment

Explore

Sources

Every stat on this page traces to a dated GitHub sync, license file, enrichment field, or trust scan.

GitHub stars on cards: lm-evaluation-harness 14k · jiwer 917 (synced Aug 7, 2026).

Common questions

What is the difference between lm-evaluation-harness and jiwer?
lm-evaluation-harness: A framework for few-shot evaluation of language models.. jiwer: Evaluate speech-to-text systems with word error rate (WER) measures. See the comparison table for live GitHub stats and shared categories.
When should I choose lm-evaluation-harness over jiwer?
Choose lm-evaluation-harness over jiwer when License: lm-evaluation-harness is MIT, jiwer is Apache-2.0; Tags unique to lm-evaluation-harness: data-parallelism, evaluation-framework, expert-parallelism, language-model; - When you need to evaluate large language models across multiple GPUs in data or tensor parallel configurations.
When should I choose jiwer over lm-evaluation-harness?
Choose jiwer over lm-evaluation-harness when License: jiwer is Apache-2.0, lm-evaluation-harness is MIT; Tags unique to jiwer: automatic-speech-recognition, evaluation-metrics, python3, speech-to-text; For accurate speech recognition evaluation needing direct WER calculations.
When should I avoid lm-evaluation-harness?
- If your evaluation setup requires pipeline parallelism not currently supported by this framework.
When should I avoid jiwer?
Avoid if your project requires only basic, non-word level performance indicators Not suitable for evaluating text-to-speech systems, as it focuses exclusively on speech-to-text output assessment
Is lm-evaluation-harness or jiwer more popular on GitHub?
lm-evaluation-harness has more GitHub stars (13,560 vs 917). Stars measure visibility, not whether either tool fits your constraints.
Are lm-evaluation-harness and jiwer open source?
Yes - both are open-source projects on GitHub (lm-evaluation-harness: MIT, jiwer: Apache-2.0).
Where can I find alternatives to lm-evaluation-harness or jiwer?
GraphCanon lists graph-backed alternatives at lm-evaluation-harness alternatives and jiwer alternatives (lm-evaluation-harness markdown twin, jiwer markdown twin), ranked by typed relationship edges rather than popularity votes.
Is there a machine-readable version of this comparison?
Yes. The markdown twin at this comparison mirrors this page for agents and LLM crawlers, with the same stats table and FAQ answers.
Which is better maintained, lm-evaluation-harness or jiwer?
lm-evaluation-harness: Active. jiwer: Slowing. Compare maintenance labels, days since push, and release cadence in the trust section below - stars alone do not measure maintenance.
Where are the full trust reports for lm-evaluation-harness and jiwer?
GraphCanon publishes per-repo trust reports with dated maintenance, provenance, and scan summaries: lm-evaluation-harness trust report; jiwer trust report.

Was this helpful?

Anonymous feedback helps us improve pages and translations.