{"data":{"slug":"swe-bench-swe-bench","name":"SWE-bench","tagline":"Benchmark for assessing language models' capability to resolve real-world Github issues","github_url":"https://github.com/SWE-bench/SWE-bench","owner":"SWE-bench","repo":"SWE-bench","owner_avatar_url":"https://avatars.githubusercontent.com/u/139597579?v=4","primary_language":"Python","stars":5576,"forks":930,"topics":["benchmark","language-model","software-engineering"],"archived":false,"github_pushed_at":"2026-07-27T05:34:27+00:00","maintenance_label":"Active","url":"https://www.graphcanon.com/tools/swe-bench-swe-bench","markdown_url":"https://www.graphcanon.com/tools/swe-bench-swe-bench.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/swe-bench-swe-bench","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=swe-bench-swe-bench","description":"SWE-bench: Can Language Models Resolve Real-world Github Issues?","homepage_url":"https://www.swebench.com","license":"MIT","open_issues":131,"watchers":38,"ai_summary":"A benchmark designed to evaluate how effectively language models can address practical software engineering challenges found in GitHub issue reports","readme_excerpt":"## ✍️ Citation & license\nMIT license. Check `LICENSE.md`.\n\nIf you find our work helpful, please use the following citations.\n\nFor SWE-bench (Verified):\n```bibtex\n@inproceedings{\n    jimenez2024swebench,\n    title={{SWE}-bench: Can Language Models Resolve Real-world Github Issues?},\n    author={Carlos E Jimenez and John Yang and Alexander Wettig and Shunyu Yao and Kexin Pei and Ofir Press and Karthik R Narasimhan},\n    booktitle={The Twelfth International Conference on Learning Representations},\n    year={2024},\n    url={https://openreview.net/forum?id=VTF8yNQM66}\n}\n```\n\nFor SWE-bench Multimodal\n```bibtex\n@inproceedings{\n    yang2024swebenchmultimodal,\n    title={{SWE}-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?},\n    author={John Yang and Carlos E. Jimenez and Alex L. Zhang and Kilian Lieret and Joyce Yang and Xindi Wu and Ori Press and Niklas Muennighoff and Gabriel Synnaeve and Karthik R. Narasimhan and Diyi Yang and Sida I. Wang and Ofir Press},\n    booktitle={The Thirteenth International Conference on Learning Representations},\n    year={2025},\n    url={https://openreview.net/forum?id=riTiq3i21b}\n}\n```\n\nFor SWE-bench Multilingual\n```bibtex\n@misc{yang2025swesmith,\n    title={SWE-smith: Scaling Data for Software Engineering Agents},\n    author={John Yang and Kilian Lieret and Carlos E. Jimenez and Alexander Wettig and Kabir Khandpur and Yanzhe Zhang and Binyuan Hui and Ofir Press and Ludwig Schmidt and Diyi Yang},\n    year={2025},\n    eprint={2504.21798},\n    archivePrefix={arXiv},\n    primaryClass={cs.SE},\n    url={https://arxiv.org/abs/2504.21798},\n}\n```","github_created_at":"2023-10-04T01:22:46+00:00","created_at":"2026-07-11T23:46:00.597809+00:00","updated_at":"2026-08-05T18:01:07.255688+00:00","categories":[{"slug":"evaluation-observability","name":"Evaluation & Observability","url":"https://www.graphcanon.com/categories/evaluation-observability","markdown_url":"https://www.graphcanon.com/categories/evaluation-observability.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/evaluation-observability"}],"tags":[{"slug":"benchmark","name":"benchmark"},{"slug":"language-model","name":"language-model"},{"slug":"software-engineering","name":"software-engineering"}],"trust":{"provenance":{"is_fork":false,"github_id":700117650,"owner_type":"Organization","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-05T18:01:06.483Z","maintenance":{"label":"Active","score":82,"methodology":"github_public_v1","releases_90d":0,"days_since_push":9,"last_release_at":null},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-11T23:46:02.735Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-05T18:01:06.948Z"},"languages":{"value":["python"],"source":"github.language+pyproject.toml","observed_at":"2026-08-05T18:01:06.948Z"},"license_spdx":{"value":"MIT","source":"github.license","observed_at":"2026-08-05T18:01:06.948Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":null,"constraints":null,"when_to_use":["When you need to evaluate the effectiveness of your language model in resolving practical software engineering challenges found in open-source repositories like GitHub.","If your research or testing involves understanding how well AI systems can generalize to visual aspects of software domains, as seen in the SWE-bench Multimodal version."],"when_not_to_use":["Do not use SWE-bench if your language model's primary application is outside the context of real-world GitHub issue resolution.","Avoid using this tool if you are not interested in testing AI systems' capabilities across visual software domains; it's more specialized for that specific area, unlike general-purpose benchmarks."],"source":"enrich:decision_facts","observed_at":"2026-07-17T06:16:13.266Z"},"constraint_facets":null,"decision_summary":[{"label":"Adopt for","value":"SWE-bench serves as a benchmark for assessing how well language models can tackle real-world software engineering issues from GitHub."},{"label":"License detail","value":"The tool operates under the MIT license, detailed in LICENSE.md."}]}}