llm-leaderboard logo

llm-leaderboard

JonathanChavezTamales/llm-leaderboard

Comprehensive LLM benchmark scores and provider prices

GraphCanon updated 4w · GitHub synced 4w

359 stars40 forksLast push 10mo JavaScript Other

Decision brief

llm-leaderboard provides deprecated benchmark data for large language models alongside service provider pricing information.

Good fit when

  • When you need to compare historical performance and service costs of different LLMs within the constraints of outdated data.

Avoid when

  • If timely or updated benchmarking data is a requirement, as llm-leaderboard's repository has been deprecated.
  • For real-time evaluations, as this tool does not provide current or recent performance metrics and pricing details.

Observed Jul 17, 2026 · Source: enrich:decision_facts

Verify the decision

Maintenance and security

Full trust report
Maintenance
Slowing (277d since push)
As of 4w
Provenance
Not a fork · Personal account
As of 4w
Security (OSV)
No lockfile
As of 1mo

Public GitHub metadata and optional OSV scans. Signals, not a guarantee. Trust methodology.

Install

npm install llm-leaderboard
npm

Similar tools

Same-category neighbours. No typed graph edges are catalogued for this tool yet.

Evidence and technical details

Sourced facts, taxonomy, compatibility claims, README excerpt, and machine-readable endpoints.

Overview

A deprecated repository containing benchmark data for large language models along with service provider pricing information.

Capability facts

MCP server
No MCP server detected

Source: repo_scan · Jul 28, 2026

Languages
javascript

Source: github.language+package.json · Jul 28, 2026

Categories

Tags

README

DEPRECATED - Updates and contributions

This repository is now depracated and won't be getting any new updates. For contributions and corrections of the data seen in LLM Stats please create a post with the tag "Issue" in the official community section of the website.

For model and/or benchmark specific corrections, please visit create an Issue under the "Discussion" tab of the model/benchmark, as seen in the example below.

Screenshot 2025-10-24 at 1 43 52 p m
image

LLM-Stats.com

A community-driven repository of LLM data and benchmarks. Compare and explore language models through our interactive dashboard at llm-stats.com.

Found an issue or have a feature request?

Open an issue here. Thank you!

Data

🔍 What's Inside

Our repository contains detailed information on hundreds of LLMs:

  • Model parameters, context window sizes, licensing details, capabilities, and more
  • Provider pricing and configurations
  • Performance metrics (throughput, latency)
  • Standardized benchmark results
  • Organization and license information

📁 Data Structure

All data is organized in the data/ directory:

  • data/models/ - Model metadata and configurations
  • data/providers/ - Provider information
  • data/provider_models/ - Provider-specific model pricing and features
  • data/benchmarks/ - Benchmark definitions
  • data/model_benchmarks/ - Model benchmark scores
  • data/organizations/ - Organization information
  • data/licenses/ - License definitions

🤝 How to Contribute

We welcome community contributions to keep our data accurate and up-to-date:

  1. Update Model Data

    • Browse the data/ directory structure
    • Submit a PR following our contribution guidelines
    • Check schemas/ for JSON Schema validation

📈 Data Quality

Accuracy is our priority. To ensure reliable information:

  • All benchmark data requires verifiable source links
  • Community review process for all changes
  • Multiple source citations encouraged
  • Regular validation of submitted data

There's no guarantee that the data is 100% accurate, but we do our best to ensure it's as accurate as possible.

🌟 Community

Leaderboard

NameRelease DateInput ContextOutput ContextGPQAMMLUMMLU-ProMATHHumanEvalMMMULiveCodeBench
GPT-52025-08-07N/AN/A0.8570.925N/A0.8470.9340.842N/A
o12024-12-17N/AN/A0.7800.918N/A0.9640.8810.776N/A
GPT-4.52025-02-27N/AN/A0.6950.908N/AN/A0.8800.752N/A
o1-preview2024-09-12N/AN/A0.7330.908N/A0.855N/AN/AN/A
Claude 3.5 Sonnet2024-10-22N/AN/A0.6720.9040.7760.7830.9370.683N/A
Claude 3.5 Sonnet2024-06-21N/AN/A0.5940.9040.7610.7110.920N/AN/A
Kimi K2 0905

For agents

This page has a .md twin and JSON over the API.

Was this helpful?

Anonymous feedback helps us improve pages and translations.