Awesome-LLM-Eval logo

Awesome-LLM-Eval

onejune2018/Awesome-LLM-Eval

Curated list for evaluation of large language models

GraphCanon updated 3w · GitHub synced 3w

654 stars82 forksLast push 9mo MIT

Decision brief

Awesome-LLM-Eval provides a comprehensive curated list of resources for evaluating large language models including tools, datasets, and benchmarks.

Good fit when

  • When you specifically need access to an extensive compilation of evaluation-related resources tailored towards large language model assessment.
  • If your focus is on the latest advancements in LLM evaluations as documented by a curated list that covers wide-ranging aspects from benchmarks to demonstrations.

Avoid when

  • You require real-time testing capabilities or interactive features; Awesome-LLM-Eval is a static resource list and not an interactive platform.
  • If integration with specific third-party platforms or direct API access is necessary, since the repository predominantly serves as a reference point rather than an operational tool.
Pricing:
freemium - The core resources listed in Awesome-LLM-Eval are freely accessible under MIT license, however, certain datasets or tools might have individual licensing terms.
Requirements:
The resources listed may vary in their own requirements, including software dependencies and hardware specifications.

Observed Jul 17, 2026 · Source: enrich:decision_facts

Verify the decision

Maintenance and security

Full trust report
Maintenance
Slowing (246d since push)
As of 3w
Provenance
Not a fork · Personal account
As of 3w
Security (OSV)
No lockfile
As of 1mo

Public GitHub metadata and optional OSV scans. Signals, not a guarantee. Trust methodology.

Install

git clone https://github.com/onejune2018/Awesome-LLM-Eval

Similar tools

Same-category neighbours. No typed graph edges are catalogued for this tool yet.

Evidence and technical details

Sourced facts, taxonomy, compatibility claims, README excerpt, and machine-readable endpoints.

Overview

Awesome-LLM-Eval: a curated list featuring tools, benchmarks, datasets, leaderboard, and documents for evaluating large language models.

Capability facts

No sourced capability facts yet. Facts appear after ingest scans repo manifests (Dockerfile, package.json, MCP configs).

Categories

Tags

README

Leaderboards for popular Provider (performance and cost, 2024-05-14)

Provider (link to pricing)OpenAIOpenAIAnthropicGoogleReplicateDeepSeekMistralAnthropicGoogleMistralCohereAnthropicMistralReplicateMistralOpenAIGroqOpenAIMistralAnthropicGroqAnthropicAnthropicMicrosoftMicrosoftMistral
Model nameGPT-4oGPT-4 TurboClaude 3 OpusGemini 1.5 ProLlama 3 70BDeepSeek-V2Mixtral 8x22BClaude 3 SonnetGemini 1.5 FlashMistral LargeCommand R+Claude 3 HaikuMistral SmallLlama 3 8BMixtral 8x7BGPT-3.5 TurboLlama 3 70B (Groq)GPT-4Mistral M

For agents

This page has a .md twin and JSON over the API.

Was this helpful?

Anonymous feedback helps us improve pages and translations.