{"data":{"slug":"itbench-hub-itbench","name":"ITBench","tagline":"An open source benchmarking framework for IT automation","github_url":"https://github.com/itbench-hub/ITBench","owner":"itbench-hub","repo":"ITBench","owner_avatar_url":"https://avatars.githubusercontent.com/u/219890183?v=4","primary_language":"Jinja","stars":506,"forks":48,"topics":["ai","automation","hacktoberfest","it-automation","itops"],"archived":false,"github_pushed_at":"2026-09-19T18:38:22+00:00","maintenance_label":"Very active","stars_delta_30d":14,"url":"https://www.graphcanon.com/tools/itbench-hub-itbench","markdown_url":"https://www.graphcanon.com/tools/itbench-hub-itbench.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/itbench-hub-itbench","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=itbench-hub-itbench","description":"An open source benchmarking framework for IT automation","homepage_url":null,"license":"Apache-2.0","open_issues":36,"watchers":12,"ai_summary":"ITBench is an open-source tool designed to help in the benchmarking of IT automation solutions.","readme_excerpt":"# ITBench\n\n**[Paper](./it_bench_arxiv.pdf) | [Leaderboard](#leaderboard) | [Scenarios](#scenarios) | [Agents](#agents) | [Related Benchmarks](#related-benchmarks) | [How to Cite](#how-to-cite) | [Contributors](./CONTRIBUTORS.md) | [Contacts](#contacts)**\n\n---\n\n## 📢 Announcements\n\n### Latest Updates\n- **[May 27, 2026]** Artificial Analysis and IBM Research launched **ITBench-AA**, the first in a new series of benchmarks evaluating frontier models on agentic enterprise IT tasks—starting with 59 SRE tasks where all evaluated models score below 50%, with FinOps and CISO tasks to follow. [View the evaluation](https://artificialanalysis.ai/evaluations/itbench-aa).\n- **[January 21, 2026]** IBM Research has published the **Enterprise Agents and Benchmarks** collection on Hugging Face, featuring ITBench alongside other enterprise AI agent ecosystems and benchmarks. [View the collection](https://huggingface.co/collections/ibm-research/enterprise-agents-and-benchmarks).\n- **[December 19, 2025]** UC Berkeley's MAST team published a blog post analyzing ITBench SRE agent traces using MAST (Multi-Agent System Failure Taxonomy), revealing structured failure signatures that explain *how* and *why* agents fail. [Read more](https://ucb-mast.notion.site/).\n- **[December 2, 2025]** ITBench is now available on Kaggle! IBM has partnered with Kaggle to launch new AI leaderboards for enterprise tasks, including ITBench. [Read more](https://research.ibm.com/blog/ibm-kaggle-leaderboards-enterprise-ai).\n- **[November 30, 2025]** A big shoutout to [@phylisscity](https://github.com/phylisscity), [@preespp](https://github.com/preespp), [@tylrnguyen](https://github.com/tylrnguyen), [@VincentCCandela](https://github.com/VincentCCandela), and [@RMalone8](https://github.com/RMalone8) from Boston University for their contributions to ITBench, and to [@Red-GV](https://github.com/Red-GV) for mentoring them!\n- **[September 18, 2025]** Our paper **STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds** was accepted at NeurIPS 2025. [Read the paper](https://arxiv.org/abs/2506.02009).\n- **[July 17, 2025]** ITBench was presented as an **oral** at ICML 2025 (Oral 6A: Applications in Agents and Coding) in Vancouver. [View the talk](https://icml.cc/virtual/2025/oral/47199).\n- **[June 13, 2025]** Identified 25+ additional scenarios to be developed over the summer.\n- **[May 2, 2025]** 🚀 ITBench now provides **fully-managed scenario environments** for everyone! Our platform handles the complete workflow—from scenario deployment to agent evaluation and leaderboard updates. Visit our GitHub repository [here](#leaderboard) for guidelines and get started today.\n- **[February 28, 2025]** 🏆 **Limited Access Beta**: Invite-only access to the ITBench hosted scenario environments. ITBench handles scenario deployment, agent evaluation, and leaderboard updates. To request access, e-mail us [here](mailto:agent-bench-automation@ibm.com).\n- **[February 7, 2025]** 🎉 **Initial release!** Includes research paper, self-hosted environment setup tooling, sample scenarios, and baseline agents.\n\n---\n\n## Overview\n\nITBench measures the performance of AI agents across a wide variety of **complex and real-world inspired IT automation tasks** targeting three key use cases:\n\n| Use Case | Focus Area |\n|----------|------------|\n| **SRE** (Site Reliability Engineering) | Availability and resiliency |\n| **CISO** (Compliance & Security Operations) | Compliance and security enforcement |\n| **FinOps** (Financial Operations) | Cost efficiencies and ROI optimization |\n\n\n\n### Key Features\n\n- **Real-world representation** of IT environments and incident scenarios\n- **Open, extensible framework** with comprehensive IT coverage\n- **Push-button workflows** and interpretable metrics\n- **Kubernetes-based** scenario environments\n\n### What's Included\n\nITBench enables researchers and developers to replicate real-world incidents in Kubernetes environments and develop AI agents to addres","github_created_at":"2025-02-05T18:09:47+00:00","created_at":"2026-07-15T10:57:26.117149+00:00","updated_at":"2026-09-20T04:57:16.807379+00:00","categories":[{"slug":"developer-tools","name":"Developer Tools","url":"https://www.graphcanon.com/categories/developer-tools","markdown_url":"https://www.graphcanon.com/categories/developer-tools.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/developer-tools"}],"tags":[{"slug":"ai","name":"ai"},{"slug":"automation","name":"automation"},{"slug":"hacktoberfest","name":"hacktoberfest"},{"slug":"it-automation","name":"it-automation"},{"slug":"itops","name":"itops"}],"trust":{"provenance":{"is_fork":false,"github_id":927895865,"owner_type":"Organization","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-09-20T04:57:14.920Z","maintenance":{"label":"Very active","score":96,"methodology":"github_public_v1","releases_90d":0,"days_since_push":0,"last_release_at":"2026-01-13T02:57:15Z","stars_delta_30d":14,"open_issues_delta_30d":2},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-15T10:57:27.354Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-09-20T04:57:15.965Z"},"languages":{"value":["jinja","python"],"source":"github.language+pyproject.toml","observed_at":"2026-09-20T04:57:15.965Z"},"license_spdx":{"value":"Apache-2.0","source":"github.license","observed_at":"2026-09-20T04:57:15.965Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":null,"constraints":null,"when_to_use":["When your team uses Python and needs a framework that allows precise control over IT automation testing scenarios","To integrate with existing CI/CD pipelines requiring open-source software under the Apache-2.0 license for compliance or preference"],"when_not_to_use":["If you require integration support for languages other than Python, as ITBench's capabilities are specifically tied to this language","In environments where proprietary tools are preferred over open-source solutions due to existing software licensing policies"],"source":"enrich:decision_facts","observed_at":"2026-07-17T10:05:15.387Z"},"constraint_facets":null,"decision_summary":[{"label":"Adopt for","value":"ITBench is suitable for teams seeking an open-source Python-based benchmarking framework designed specifically to evaluate the effectiveness of IT automation solutions."}]}}