{"data":{"slug":"jonathanchaveztamales-llm-leaderboard","name":"llm-leaderboard","tagline":"Comprehensive LLM benchmark scores and provider prices","github_url":"https://github.com/JonathanChavezTamales/llm-leaderboard","owner":"JonathanChavezTamales","repo":"llm-leaderboard","owner_avatar_url":"https://avatars.githubusercontent.com/u/22694942?v=4","primary_language":"JavaScript","stars":359,"forks":40,"topics":["llm","llm-agents","llm-evaluation","llmops","llms-benchmarking"],"archived":false,"github_pushed_at":"2025-10-24T17:47:59+00:00","maintenance_label":"Slowing","url":"https://www.graphcanon.com/tools/jonathanchaveztamales-llm-leaderboard","markdown_url":"https://www.graphcanon.com/tools/jonathanchaveztamales-llm-leaderboard.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/jonathanchaveztamales-llm-leaderboard","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=jonathanchaveztamales-llm-leaderboard","description":"A comprehensive set of LLM benchmark scores and provider prices. (deprecated, read more in README)","homepage_url":"https://llm-stats.com","license":"Other","open_issues":14,"watchers":12,"ai_summary":"A deprecated repository containing benchmark data for large language models along with service provider pricing information.","readme_excerpt":"# DEPRECATED - Updates and contributions\n\nThis repository is now depracated and won't be getting any new updates. For contributions and corrections of the data seen in [LLM Stats](https://llm-stats.com/) please create a post with the tag \"Issue\" in the [official community section](https://llm-stats.com/posts) of the website.\n\nFor model and/or benchmark specific corrections, please visit create an Issue under the \"Discussion\" tab of the model/benchmark, as seen in the example below.\n\n<img width=\"1156\" height=\"575\" alt=\"Screenshot 2025-10-24 at 1 43 52 p m\" src=\"https://github.com/user-attachments/assets/b78f2cf3-f3ff-4a51-bba4-d8643865d16b\" />\n\n---\n\n<img width=\"1208\" alt=\"image\" src=\"https://github.com/user-attachments/assets/835f1e1b-73e6-405a-b7ad-096d5f5f567a\" />\n\n# LLM-Stats.com\n\n\n\n\n\n\nA community-driven repository of LLM data and benchmarks. Compare and explore language models through our interactive dashboard at [llm-stats.com](https://llm-stats.com).\n\n## Found an issue or have a feature request?\n\n[Open an issue here](https://github.com/JonathanChavezTamales/llm-leaderboard/issues). Thank you!\n\n# Data\n\n## 🔍 What's Inside\n\nOur repository contains detailed information on hundreds of LLMs:\n\n- Model parameters, context window sizes, licensing details, capabilities, and more\n- Provider pricing and configurations\n- Performance metrics (throughput, latency)\n- Standardized benchmark results\n- Organization and license information\n\n## 📁 Data Structure\n\nAll data is organized in the `data/` directory:\n\n- `data/models/` - Model metadata and configurations\n- `data/providers/` - Provider information\n- `data/provider_models/` - Provider-specific model pricing and features\n- `data/benchmarks/` - Benchmark definitions\n- `data/model_benchmarks/` - Model benchmark scores\n- `data/organizations/` - Organization information\n- `data/licenses/` - License definitions\n\n## 🤝 How to Contribute\n\nWe welcome community contributions to keep our data accurate and up-to-date:\n\n1. **Update Model Data**\n\n   - Browse the [`data/`](data/) directory structure\n   - Submit a PR following our [contribution guidelines](CONTRIBUTING.md)\n   - Check [`schemas/`](schemas/) for JSON Schema validation\n\n## 📈 Data Quality\n\nAccuracy is our priority. To ensure reliable information:\n\n- All benchmark data requires verifiable source links\n- Community review process for all changes\n- Multiple source citations encouraged\n- Regular validation of submitted data\n\nThere's no guarantee that the data is 100% accurate, but we do our best to ensure it's as accurate as possible.\n\n## 🌟 Community\n\n- Join our [Discord](https://discord.gg/RxGUBvE42d) for discussions\n\n## Leaderboard\n\n| Name                                     | Release Date | Input Context | Output Context | GPQA  | MMLU  | MMLU-Pro | MATH  | HumanEval | MMMU  | LiveCodeBench |\n| ---------------------------------------- | ------------ | ------------- | -------------- | ----- | ----- | -------- | ----- | --------- | ----- | ------------- |\n| GPT-5                                    | 2025-08-07   | N/A           | N/A            | 0.857 | 0.925 | N/A      | 0.847 | 0.934     | 0.842 | N/A           |\n| o1                                       | 2024-12-17   | N/A           | N/A            | 0.780 | 0.918 | N/A      | 0.964 | 0.881     | 0.776 | N/A           |\n| GPT-4.5                                  | 2025-02-27   | N/A           | N/A            | 0.695 | 0.908 | N/A      | N/A   | 0.880     | 0.752 | N/A           |\n| o1-preview                               | 2024-09-12   | N/A           | N/A            | 0.733 | 0.908 | N/A      | 0.855 | N/A       | N/A   | N/A           |\n| Claude 3.5 Sonnet                        | 2024-10-22   | N/A           | N/A            | 0.672 | 0.904 | 0.776    | 0.783 | 0.937     | 0.683 | N/A           |\n| Claude 3.5 Sonnet                        | 2024-06-21   | N/A           | N/A            | 0.594 | 0.904 | 0.761    | 0.711 | 0.920     | N/A   | N/A           |\n| Kimi K2 0905","github_created_at":"2024-09-07T22:33:55+00:00","created_at":"2026-07-11T12:00:13.041546+00:00","updated_at":"2026-07-28T18:00:26.172027+00:00","categories":[{"slug":"evaluation-observability","name":"Evaluation & Observability","url":"https://www.graphcanon.com/categories/evaluation-observability","markdown_url":"https://www.graphcanon.com/categories/evaluation-observability.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/evaluation-observability"},{"slug":"llm-frameworks","name":"LLM Frameworks","url":"https://www.graphcanon.com/categories/llm-frameworks","markdown_url":"https://www.graphcanon.com/categories/llm-frameworks.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/llm-frameworks"}],"tags":[{"slug":"llm","name":"llm"},{"slug":"llm-agents","name":"llm-agents"},{"slug":"llm-evaluation","name":"llm-evaluation"},{"slug":"llmops","name":"llmops"},{"slug":"llms-benchmarking","name":"llms-benchmarking"}],"trust":{"provenance":{"is_fork":false,"github_id":853917746,"owner_type":"User","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-07-28T18:00:24.816Z","maintenance":{"label":"Slowing","score":36,"methodology":"github_public_v1","releases_90d":0,"days_since_push":277,"last_release_at":null},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-11T12:00:14.445Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"mcp":{"source":"repo_scan","observed_at":"2026-07-28T18:00:25.360Z","server_manifest":false},"scan":{"source":"repo_scan","observed_at":"2026-07-28T18:00:25.360Z"},"languages":{"value":["javascript"],"source":"github.language+package.json","observed_at":"2026-07-28T18:00:25.360Z"},"license_spdx":{"value":"Other","source":"github.license","observed_at":"2026-07-28T18:00:25.360Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":null,"constraints":null,"when_to_use":["When you need to compare historical performance and service costs of different LLMs within the constraints of outdated data."],"when_not_to_use":["If timely or updated benchmarking data is a requirement, as llm-leaderboard's repository has been deprecated.","For real-time evaluations, as this tool does not provide current or recent performance metrics and pricing details."],"source":"enrich:decision_facts","observed_at":"2026-07-17T00:21:59.410Z"},"constraint_facets":null,"decision_summary":[{"label":"Adopt for","value":"llm-leaderboard provides deprecated benchmark data for large language models alongside service provider pricing information."}]}}