llm-leaderboard
Comprehensive LLM benchmark scores and provider prices
GraphCanon updated 4w · GitHub synced 4w
Decision brief
llm-leaderboard provides deprecated benchmark data for large language models alongside service provider pricing information.
Good fit when
- When you need to compare historical performance and service costs of different LLMs within the constraints of outdated data.
Avoid when
- If timely or updated benchmarking data is a requirement, as llm-leaderboard's repository has been deprecated.
- For real-time evaluations, as this tool does not provide current or recent performance metrics and pricing details.
Observed Jul 17, 2026 · Source: enrich:decision_facts
Verify the decision
Maintenance and security
Full trust report- Maintenance
- Slowing (277d since push)
- As of 4w
- Provenance
- Not a fork · Personal account
- As of 4w
- Security (OSV)
- No lockfile
- As of 1mo
Public GitHub metadata and optional OSV scans. Signals, not a guarantee. Trust methodology.
Install
npm install llm-leaderboard npmSimilar tools
Same-category neighbours. No typed graph edges are catalogued for this tool yet.
Evidence and technical details
Sourced facts, taxonomy, compatibility claims, README excerpt, and machine-readable endpoints.
Overview
A deprecated repository containing benchmark data for large language models along with service provider pricing information.
Capability facts
- MCP server
- No MCP server detected
Source: repo_scan · Jul 28, 2026
- Languages
- javascript
Source: github.language+package.json · Jul 28, 2026
Categories
Tags
README
DEPRECATED - Updates and contributions
This repository is now depracated and won't be getting any new updates. For contributions and corrections of the data seen in LLM Stats please create a post with the tag "Issue" in the official community section of the website.
For model and/or benchmark specific corrections, please visit create an Issue under the "Discussion" tab of the model/benchmark, as seen in the example below.
LLM-Stats.com
A community-driven repository of LLM data and benchmarks. Compare and explore language models through our interactive dashboard at llm-stats.com.
Found an issue or have a feature request?
Open an issue here. Thank you!
Data
🔍 What's Inside
Our repository contains detailed information on hundreds of LLMs:
- Model parameters, context window sizes, licensing details, capabilities, and more
- Provider pricing and configurations
- Performance metrics (throughput, latency)
- Standardized benchmark results
- Organization and license information
📁 Data Structure
All data is organized in the data/ directory:
data/models/- Model metadata and configurationsdata/providers/- Provider informationdata/provider_models/- Provider-specific model pricing and featuresdata/benchmarks/- Benchmark definitionsdata/model_benchmarks/- Model benchmark scoresdata/organizations/- Organization informationdata/licenses/- License definitions
🤝 How to Contribute
We welcome community contributions to keep our data accurate and up-to-date:
-
Update Model Data
- Browse the
data/directory structure - Submit a PR following our contribution guidelines
- Check
schemas/for JSON Schema validation
- Browse the
📈 Data Quality
Accuracy is our priority. To ensure reliable information:
- All benchmark data requires verifiable source links
- Community review process for all changes
- Multiple source citations encouraged
- Regular validation of submitted data
There's no guarantee that the data is 100% accurate, but we do our best to ensure it's as accurate as possible.
🌟 Community
- Join our Discord for discussions
Leaderboard
| Name | Release Date | Input Context | Output Context | GPQA | MMLU | MMLU-Pro | MATH | HumanEval | MMMU | LiveCodeBench |
|---|---|---|---|---|---|---|---|---|---|---|
| GPT-5 | 2025-08-07 | N/A | N/A | 0.857 | 0.925 | N/A | 0.847 | 0.934 | 0.842 | N/A |
| o1 | 2024-12-17 | N/A | N/A | 0.780 | 0.918 | N/A | 0.964 | 0.881 | 0.776 | N/A |
| GPT-4.5 | 2025-02-27 | N/A | N/A | 0.695 | 0.908 | N/A | N/A | 0.880 | 0.752 | N/A |
| o1-preview | 2024-09-12 | N/A | N/A | 0.733 | 0.908 | N/A | 0.855 | N/A | N/A | N/A |
| Claude 3.5 Sonnet | 2024-10-22 | N/A | N/A | 0.672 | 0.904 | 0.776 | 0.783 | 0.937 | 0.683 | N/A |
| Claude 3.5 Sonnet | 2024-06-21 | N/A | N/A | 0.594 | 0.904 | 0.761 | 0.711 | 0.920 | N/A | N/A |
| Kimi K2 0905 |
For agents
This page has a .md twin and JSON over the API.