{"data":{"slug":"livecodebench-livecodebench","name":"LiveCodeBench","tagline":"Holistic and contamination-free evaluation of large language models for code","github_url":"https://github.com/LiveCodeBench/LiveCodeBench","owner":"LiveCodeBench","repo":"LiveCodeBench","owner_avatar_url":"https://avatars.githubusercontent.com/u/161278213?v=4","primary_language":"Python","stars":925,"forks":195,"topics":["code-execution","code-generation","code-llms","code-repair","gpt-4","test-generation"],"archived":false,"github_pushed_at":"2025-07-16T00:58:38+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/livecodebench-livecodebench","markdown_url":"https://www.graphcanon.com/tools/livecodebench-livecodebench.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/livecodebench-livecodebench","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=livecodebench-livecodebench","description":"Official repository for the paper \"LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code\"","homepage_url":"https://livecodebench.github.io/","license":"MIT","open_issues":38,"watchers":7,"ai_summary":"LiveCodeBench evaluates the capabilities of large language models in generating, repairing, and executing code with a focus on maintaining high standards to prevent model contamination.","readme_excerpt":"## Installation\nYou can clone the repository using the following command:\n\n```bash\ngit clone https://github.com/LiveCodeBench/LiveCodeBench.git\ncd LiveCodeBench\n```\n\nWe recommend using [uv](https://github.com/astral-sh/uv)\nfor managing dependencies, which can be installed a [number of ways](https://github.com/astral-sh/uv?tab=readme-ov-file#installation).\n\nVerify that `uv` is installed on your system by running:\n\n```bash\nuv --version\n```\n\nOnce `uv` has been installed, use it to create a virtual environment for\nLiveCodeBench and install its dependencies with the following commands:\n\n```bash\nuv venv --python 3.11\nsource .venv/bin/activate\n\nuv pip install -e .\n```","github_created_at":"2024-03-12T23:34:37+00:00","created_at":"2026-07-11T23:45:46.907649+00:00","updated_at":"2026-08-05T18:01:02.29995+00:00","categories":[{"slug":"evaluation-observability","name":"Evaluation & Observability","url":"https://www.graphcanon.com/categories/evaluation-observability","markdown_url":"https://www.graphcanon.com/categories/evaluation-observability.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/evaluation-observability"}],"tags":[{"slug":"code-generation","name":"code generation"},{"slug":"code-execution","name":"code-execution"},{"slug":"code-repair","name":"code-repair"},{"slug":"gpt-4","name":"gpt-4"},{"slug":"python","name":"python"},{"slug":"test-generation","name":"test-generation"}],"trust":{"provenance":{"is_fork":false,"github_id":771234452,"owner_type":"Organization","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-05T18:01:01.438Z","maintenance":{"label":"Dormant","score":18,"methodology":"github_public_v1","releases_90d":0,"days_since_push":385,"last_release_at":null},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-11T23:45:53.933Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-05T18:01:01.969Z"},"languages":{"value":["python"],"source":"github.language+pyproject.toml","observed_at":"2026-08-05T18:01:01.969Z"},"license_spdx":{"value":"MIT","source":"github.license","observed_at":"2026-08-05T18:01:01.969Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":null,"constraints":null,"when_to_use":["When you need a holistic method to assess the effectiveness of LLMs in code tasks without risking contamination by earlier outputs or data leakage.","If you are working with Python-based projects and require precise and unbiased evaluations leveraging advanced features like GPT-4."],"when_not_to_use":["For broad, non-code-specific model assessments where a more generalized evaluation tool would suffice.","If your project is not compatible with Python 3.11 or if you do not want to use the uv dependency manager recommended by LiveCodeBench."],"source":"enrich:decision_facts","observed_at":"2026-07-17T06:13:46.768Z"},"constraint_facets":null,"decision_summary":[{"label":"Adopt for","value":"LiveCodeBench offers an in-depth approach to evaluating large language models specifically for code tasks such as generation and repair."}]}}