awesome-evals
A curated library of resources for building and evaluating AI agents
GraphCanon updated 3w · GitHub synced 3w
Decision brief
Curated resources for AI agent evaluation with BenchFlow backing its maintenance
Good fit when
- Need diverse resources encompassing papers, blogs, talks, tools, and benchmarks specifically curated for AI agent evaluation
- Looking to understand how AI agents are evaluated in both LLMs and RL environments
Avoid when
- Require real-time interactive support or direct tool integrations not covered by a static resource list
- Seeking proprietary tools from specific vendors rather than open resources and community content
Observed Jul 17, 2026 · Source: enrich:decision_facts
Verify the decision
Maintenance and security
Full trust report- Maintenance
- Active (26d since push)
- As of 3w
- Provenance
- Not a fork · Organization account
- As of 3w
- Security (OSV)
- No lockfile
- As of 1mo
Public GitHub metadata and optional OSV scans. Signals, not a guarantee. Trust methodology.
Install
git clone https://github.com/benchflow-ai/awesome-evalsSimilar tools
Same-category neighbours. No typed graph edges are catalogued for this tool yet.
Evidence and technical details
Sourced facts, taxonomy, compatibility claims, README excerpt, and machine-readable endpoints.
Overview
Maintained by BenchFlow, this resource offers papers, blogs, talks, tools, and benchmarks focused on agent evaluation, featuring components related to AI agents, LLMs, and RL environments.
Capability facts
No sourced capability facts yet. Facts appear after ingest scans repo manifests (Dockerfile, package.json, MCP configs).
Categories
Tags
README
5 · Evaluation infrastructure (the eval stack: datasets, scorers, online/offline, tracing, CI)
(All repos URL-verified via GitHub API, Jun 2026. 🆕 = released/expanded 2025–2026. ⚠️ = caveat/discontinued.)
License
To the extent possible under law, BenchFlow and contributors have waived all copyright and related rights to this work (CC0 1.0). The linked resources remain under their respective licenses.
For agents
This page has a .md twin and JSON over the API.