Home/Compare/olmo-eval vs hallucination-index

Comparison

olmo-eval vs hallucination-index

Verdict

Pick olmo-eval if olmo-eval is an evaluation framework for large language models, using uv for reproducible builds. It focuses on modular task implementations and integrates with various datasets via defined tasks; pick hallucination-index if hallucination-Index helps users identify LLMs with the lowest propensity for factual errors across varying context lengths and source types.

Markdown twin · olmo-eval alternatives · hallucination-index alternatives

GraphCanon updated 2w

olmo-eval logo

olmo-eval

allenai/olmo-eval

65pushed Aug 6, 2026
vs
hallucination-index logo

hallucination-index

rungalileo/hallucination-index

116pushed Jul 28, 2025

Trust & integrity

Signalolmo-evalhallucination-index
Maintenance
Very active (0d since push)
As of 2w · github_public_v1
Dormant (365d since push)
As of 3w · github_public_v1
Provenance
Not a fork · Organization account
As of 2w · github_public_v1
Not a fork · Organization account
As of 3w · github_public_v1
OSV dependency advisories
No lockfile (source not queried)
As of 1mo · osv@v1
No lockfile (source not queried)
As of 1mo · osv@v1
deps.dev advisories
Not queried
deps.dev@v1
Not queried
deps.dev@v1
OpenSSF Scorecard
Not queried
openssf-scorecard@v1
Not queried
openssf-scorecard@v1

Tagline

olmo-eval
Olmo Evaluation Framework for LLM Tasks
hallucination-index
Initiative to evaluate and rank popular LLMs based on hallucination propensity

Stars

olmo-eval
65
hallucination-index
116

Forks

olmo-eval
14
hallucination-index
8

Open issues

olmo-eval
38
hallucination-index
1

Language

olmo-eval
Python
hallucination-index
-

Adopt for

olmo-eval
Olmo-eval is an evaluation framework for large language models, using uv for reproducible builds. It focuses on modular task implementations and integrates with various datasets via defined tasks.
hallucination-index
Hallucination-Index helps users identify LLMs with the lowest propensity for factual errors across varying context lengths and source types.

Persona

olmo-eval
-
hallucination-index
-

Runtime

olmo-eval
-
hallucination-index
-

License

olmo-eval
Apache-2.0
hallucination-index
-

Last pushed

olmo-eval
Aug 6, 2026
hallucination-index
Jul 28, 2025

Categories

olmo-eval
Evaluation & Observability
hallucination-index
Evaluation & Observability

Trust and health

Maintenance

olmo-eval
Very active (96%)
hallucination-index
Dormant (18%)

Days since push

olmo-eval
0d
hallucination-index
365d

Open issues (now)

olmo-eval
38
hallucination-index
1

Full report

olmo-eval
Trust report
hallucination-index
Trust report

Choose olmo-eval if…

  • Tags unique to olmo-eval: datasets, evaluation, llm, python.
  • olmo-eval ships Docker support for self-hosted deployment.
  • When you need a flexible evaluation setup that works with a variety of LLMs and datasets.

When NOT to use olmo-eval

  • When you require a simpler setup that doesn't need the reproducibility constraints of uv builds.
  • If your project already has an established evaluation toolchain and does not benefit from introducing a new framework for manageability reasons.

Choose hallucination-index if…

  • Tags unique to hallucination-index: hallucinations, large language models, llm-evaluation, openai.
  • Use when you need to ensure accuracy in short-context tasks, as it tests models like Chain-of-Note prompting techniques specifically for such scenarios.
  • More GitHub stars (116 vs 65) - visibility, not fit.

When NOT to use hallucination-index

  • Avoid using Hallucination-Index when your application requires real-time evaluation of hallucinations, as it focuses on predefined tests rather than live model performance.
  • Do not rely solely on this index if your primary concern is the latest updates to LLM models; its data might not reflect recent improvements in models or the introduction of new ones.

Explore

Sources

Every stat on this page traces to a dated GitHub sync, license file, enrichment field, or trust scan.

GitHub stars on cards: olmo-eval 65 · hallucination-index 116 (synced Aug 7, 2026).

Common questions

What is the difference between olmo-eval and hallucination-index?
olmo-eval: Olmo Evaluation Framework for LLM Tasks. hallucination-index: Initiative to evaluate and rank popular LLMs based on hallucination propensity. See the comparison table for live GitHub stats and shared categories.
When should I choose olmo-eval over hallucination-index?
Choose olmo-eval over hallucination-index when Tags unique to olmo-eval: datasets, evaluation, llm, python; olmo-eval ships Docker support for self-hosted deployment; When you need a flexible evaluation setup that works with a variety of LLMs and datasets.
When should I choose hallucination-index over olmo-eval?
Choose hallucination-index over olmo-eval when Tags unique to hallucination-index: hallucinations, large language models, llm-evaluation, openai; Use when you need to ensure accuracy in short-context tasks, as it tests models like Chain-of-Note prompting techniques specifically for such scenarios; More GitHub stars (116 vs 65) - visibility, not fit.
When should I avoid olmo-eval?
When you require a simpler setup that doesn't need the reproducibility constraints of uv builds. If your project already has an established evaluation toolchain and does not benefit from introducing a new framework for manageability reasons.
When should I avoid hallucination-index?
Avoid using Hallucination-Index when your application requires real-time evaluation of hallucinations, as it focuses on predefined tests rather than live model performance. Do not rely solely on this index if your primary concern is the latest updates to LLM models; its data might not reflect recent improvements in models or the introduction of new ones.
Is olmo-eval or hallucination-index more popular on GitHub?
hallucination-index has more GitHub stars (116 vs 65). Stars measure visibility, not whether either tool fits your constraints.
Are olmo-eval and hallucination-index open source?
Yes - both are open-source projects on GitHub.
Where can I find alternatives to olmo-eval or hallucination-index?
GraphCanon lists graph-backed alternatives at olmo-eval alternatives and hallucination-index alternatives (olmo-eval markdown twin, hallucination-index markdown twin), ranked by typed relationship edges rather than popularity votes.
Is there a machine-readable version of this comparison?
Yes. The markdown twin at this comparison mirrors this page for agents and LLM crawlers, with the same stats table and FAQ answers.
Which is better maintained, olmo-eval or hallucination-index?
olmo-eval: Very active. hallucination-index: Dormant. Compare maintenance labels, days since push, and release cadence in the trust section below - stars alone do not measure maintenance.
Where are the full trust reports for olmo-eval and hallucination-index?
GraphCanon publishes per-repo trust reports with dated maintenance, provenance, and scan summaries: olmo-eval trust report; hallucination-index trust report.

Was this helpful?

Anonymous feedback helps us improve pages and translations.