GraphCanon updated 3w · GitHub synced 3w
Decision brief
Awesome-LLM-Eval provides a comprehensive curated list of resources for evaluating large language models including tools, datasets, and benchmarks.
Good fit when
- When you specifically need access to an extensive compilation of evaluation-related resources tailored towards large language model assessment.
- If your focus is on the latest advancements in LLM evaluations as documented by a curated list that covers wide-ranging aspects from benchmarks to demonstrations.
Avoid when
- You require real-time testing capabilities or interactive features; Awesome-LLM-Eval is a static resource list and not an interactive platform.
- If integration with specific third-party platforms or direct API access is necessary, since the repository predominantly serves as a reference point rather than an operational tool.
- Pricing:
- freemium - The core resources listed in Awesome-LLM-Eval are freely accessible under MIT license, however, certain datasets or tools might have individual licensing terms.
- Requirements:
- The resources listed may vary in their own requirements, including software dependencies and hardware specifications.
Observed Jul 17, 2026 · Source: enrich:decision_facts
Verify the decision
Maintenance and security
Full trust report- Maintenance
- Slowing (246d since push)
- As of 3w
- Provenance
- Not a fork · Personal account
- As of 3w
- Security (OSV)
- No lockfile
- As of 1mo
Public GitHub metadata and optional OSV scans. Signals, not a guarantee. Trust methodology.
Install
git clone https://github.com/onejune2018/Awesome-LLM-EvalSimilar tools
Same-category neighbours. No typed graph edges are catalogued for this tool yet.
Evidence and technical details
Sourced facts, taxonomy, compatibility claims, README excerpt, and machine-readable endpoints.
Overview
Awesome-LLM-Eval: a curated list featuring tools, benchmarks, datasets, leaderboard, and documents for evaluating large language models.
Capability facts
No sourced capability facts yet. Facts appear after ingest scans repo manifests (Dockerfile, package.json, MCP configs).
Categories
Tags
README
Leaderboards for popular Provider (performance and cost, 2024-05-14)
| Provider (link to pricing) | OpenAI | OpenAI | Anthropic | Replicate | DeepSeek | Mistral | Anthropic | Mistral | Cohere | Anthropic | Mistral | Replicate | Mistral | OpenAI | Groq | OpenAI | Mistral | Anthropic | Groq | Anthropic | Anthropic | Microsoft | Microsoft | Mistral | |||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Model name | GPT-4o | GPT-4 Turbo | Claude 3 Opus | Gemini 1.5 Pro | Llama 3 70B | DeepSeek-V2 | Mixtral 8x22B | Claude 3 Sonnet | Gemini 1.5 Flash | Mistral Large | Command R+ | Claude 3 Haiku | Mistral Small | Llama 3 8B | Mixtral 8x7B | GPT-3.5 Turbo | Llama 3 70B (Groq) | GPT-4 | Mistral M |
For agents
This page has a .md twin and JSON over the API.