Comparison
olmo-eval vs hallucination-index
Verdict
Pick olmo-eval if olmo-eval is an evaluation framework for large language models, using uv for reproducible builds. It focuses on modular task implementations and integrates with various datasets via defined tasks; pick hallucination-index if hallucination-Index helps users identify LLMs with the lowest propensity for factual errors across varying context lengths and source types.
Markdown twin · olmo-eval alternatives · hallucination-index alternatives
GraphCanon updated 2w
Trust & integrity
| Signal | olmo-eval | hallucination-index |
|---|---|---|
| Maintenance | Very active (0d since push) As of 2w · github_public_v1 | Dormant (365d since push) As of 3w · github_public_v1 |
| Provenance | Not a fork · Organization account As of 2w · github_public_v1 | Not a fork · Organization account As of 3w · github_public_v1 |
| OSV dependency advisories | No lockfile (source not queried) As of 1mo · osv@v1 | No lockfile (source not queried) As of 1mo · osv@v1 |
| deps.dev advisories | Not queried deps.dev@v1 | Not queried deps.dev@v1 |
| OpenSSF Scorecard | Not queried openssf-scorecard@v1 | Not queried openssf-scorecard@v1 |
Tagline
- olmo-eval
- Olmo Evaluation Framework for LLM Tasks
- hallucination-index
- Initiative to evaluate and rank popular LLMs based on hallucination propensity
Stars
- olmo-eval
- 65
- hallucination-index
- 116
Forks
- olmo-eval
- 14
- hallucination-index
- 8
Open issues
- olmo-eval
- 38
- hallucination-index
- 1
Language
- olmo-eval
- Python
- hallucination-index
- -
Adopt for
- olmo-eval
- Olmo-eval is an evaluation framework for large language models, using uv for reproducible builds. It focuses on modular task implementations and integrates with various datasets via defined tasks.
- hallucination-index
- Hallucination-Index helps users identify LLMs with the lowest propensity for factual errors across varying context lengths and source types.
Persona
- olmo-eval
- -
- hallucination-index
- -
Runtime
- olmo-eval
- -
- hallucination-index
- -
License
- olmo-eval
- Apache-2.0
- hallucination-index
- -
Last pushed
- olmo-eval
- Aug 6, 2026
- hallucination-index
- Jul 28, 2025
Categories
- olmo-eval
- Evaluation & Observability
- hallucination-index
- Evaluation & Observability
Trust and health
Maintenance
- olmo-eval
- Very active (96%)
- hallucination-index
- Dormant (18%)
Days since push
- olmo-eval
- 0d
- hallucination-index
- 365d
Open issues (now)
- olmo-eval
- 38
- hallucination-index
- 1
Full report
- olmo-eval
- Trust report
- hallucination-index
- Trust report
Choose olmo-eval if…
- Tags unique to olmo-eval: datasets, evaluation, llm, python.
- olmo-eval ships Docker support for self-hosted deployment.
- When you need a flexible evaluation setup that works with a variety of LLMs and datasets.
When NOT to use olmo-eval
- When you require a simpler setup that doesn't need the reproducibility constraints of uv builds.
- If your project already has an established evaluation toolchain and does not benefit from introducing a new framework for manageability reasons.
Choose hallucination-index if…
- Tags unique to hallucination-index: hallucinations, large language models, llm-evaluation, openai.
- Use when you need to ensure accuracy in short-context tasks, as it tests models like Chain-of-Note prompting techniques specifically for such scenarios.
- More GitHub stars (116 vs 65) - visibility, not fit.
When NOT to use hallucination-index
- Avoid using Hallucination-Index when your application requires real-time evaluation of hallucinations, as it focuses on predefined tests rather than live model performance.
- Do not rely solely on this index if your primary concern is the latest updates to LLM models; its data might not reflect recent improvements in models or the introduction of new ones.
Explore
Sources
Every stat on this page traces to a dated GitHub sync, license file, enrichment field, or trust scan.
- GitHub stars (allenai/olmo-eval) · observed Aug 7, 2026
- GitHub forks (allenai/olmo-eval) · observed Aug 7, 2026
- Last push (allenai/olmo-eval) · observed Aug 6, 2026
- License file (Apache-2.0) · observed Aug 7, 2026
- Decision facts (enrichment) · observed Jul 12, 2026
- Trust scan (lockfile / OSV) · observed Jul 11, 2026
- GitHub stars (rungalileo/hallucination-index) · observed Jul 29, 2026
- GitHub forks (rungalileo/hallucination-index) · observed Jul 29, 2026
- Last push (rungalileo/hallucination-index) · observed Jul 28, 2025
- License file (unknown) · observed Jul 29, 2026
- Decision facts (enrichment) · observed Jul 17, 2026
- Trust scan (lockfile / OSV) · observed Jul 11, 2026
GitHub stars on cards: olmo-eval 65 · hallucination-index 116 (synced Aug 7, 2026).
Common questions
- What is the difference between olmo-eval and hallucination-index?
- olmo-eval: Olmo Evaluation Framework for LLM Tasks. hallucination-index: Initiative to evaluate and rank popular LLMs based on hallucination propensity. See the comparison table for live GitHub stats and shared categories.
- When should I choose olmo-eval over hallucination-index?
- Choose olmo-eval over hallucination-index when Tags unique to olmo-eval: datasets, evaluation, llm, python; olmo-eval ships Docker support for self-hosted deployment; When you need a flexible evaluation setup that works with a variety of LLMs and datasets.
- When should I choose hallucination-index over olmo-eval?
- Choose hallucination-index over olmo-eval when Tags unique to hallucination-index: hallucinations, large language models, llm-evaluation, openai; Use when you need to ensure accuracy in short-context tasks, as it tests models like Chain-of-Note prompting techniques specifically for such scenarios; More GitHub stars (116 vs 65) - visibility, not fit.
- When should I avoid olmo-eval?
- When you require a simpler setup that doesn't need the reproducibility constraints of uv builds. If your project already has an established evaluation toolchain and does not benefit from introducing a new framework for manageability reasons.
- When should I avoid hallucination-index?
- Avoid using Hallucination-Index when your application requires real-time evaluation of hallucinations, as it focuses on predefined tests rather than live model performance. Do not rely solely on this index if your primary concern is the latest updates to LLM models; its data might not reflect recent improvements in models or the introduction of new ones.
- Is olmo-eval or hallucination-index more popular on GitHub?
- hallucination-index has more GitHub stars (116 vs 65). Stars measure visibility, not whether either tool fits your constraints.
- Are olmo-eval and hallucination-index open source?
- Yes - both are open-source projects on GitHub.
- Where can I find alternatives to olmo-eval or hallucination-index?
- GraphCanon lists graph-backed alternatives at olmo-eval alternatives and hallucination-index alternatives (olmo-eval markdown twin, hallucination-index markdown twin), ranked by typed relationship edges rather than popularity votes.
- Is there a machine-readable version of this comparison?
- Yes. The markdown twin at this comparison mirrors this page for agents and LLM crawlers, with the same stats table and FAQ answers.
- Which is better maintained, olmo-eval or hallucination-index?
- olmo-eval: Very active. hallucination-index: Dormant. Compare maintenance labels, days since push, and release cadence in the trust section below - stars alone do not measure maintenance.
- Where are the full trust reports for olmo-eval and hallucination-index?
- GraphCanon publishes per-repo trust reports with dated maintenance, provenance, and scan summaries: olmo-eval trust report; hallucination-index trust report.