{"data":{"slug":"kaito-project-kaito","name":"kaito","tagline":"Kubernetes AI Toolchain Operator for managing and scaling inference workloads","github_url":"https://github.com/kaito-project/kaito","owner":"kaito-project","repo":"kaito","owner_avatar_url":"https://avatars.githubusercontent.com/u/186863079?v=4","primary_language":"Go","stars":992,"forks":176,"topics":["ai","gpu","kubernetes","operator"],"archived":false,"github_pushed_at":"2026-08-01T03:48:04+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/kaito-project-kaito","markdown_url":"https://www.graphcanon.com/tools/kaito-project-kaito.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/kaito-project-kaito","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=kaito-project-kaito","description":"Kubernetes AI Toolchain Operator","homepage_url":"https://kaito-project.github.io/kaito/docs/","license":"Other","open_issues":62,"watchers":10,"ai_summary":"A Kubernetes operator that enables the deployment, scaling, and management of AI models in a production environment using Helm or Terraform.","readme_excerpt":"## Getting Started \n- **Installation**: Please check the guidance [here](https://kaito-project.github.io/kaito/docs/installation) for installing core components (Workspace, InferenceSet) using helm and [here](https://github.com/kaito-project/kaito/blob/main/terraform/README.md) for installation using Terraform.\n- **Quick Start**: Please check the quick start guidance [here](https://kaito-project.github.io/kaito/docs/quick-start) for running your first model using KAITO!\n- **AutoScaling**: Please check this [doc](https://kaito-project.github.io/kaito/docs/keda-autoscaler-inference) for configuring KAITO and KEDA to enable autoscaling inference workload.\n- **BYO models using HuggingFace runtime**: If you plan to run any BYO models using the HuggingFace runtime, check this [doc](https://kaito-project.github.io/kaito/docs/custom-model). Note: KAITO only supports BYO models hosted in HuggingFace.\n- **CPU models**: Please check this [doc](https://kaito-project.github.io/kaito/docs/aikit) for running CPU models using [aikit](https://github.com/kaito-project/aikit/).\n- **RAGEngine**: Please check the installation guidance and usage documents [here](https://kaito-project.github.io/kaito/docs/rag).\n- **RAGEngine Output Guardrails**: Please check the current behavior, configuration, and limitations [here](https://kaito-project.github.io/kaito/docs/rag-output-guardrails).\n- **Prefill/Decode Disaggregation**: Please check this [doc](https://kaito-project.github.io/kaito/docs/prefill-decode-disaggregation) for deploying models with P/D disaggregation using MultiRoleInference and Gateway API Inference Extension.\n\n---\n\n## License\n\nSee [Apache License 2.0](LICENSE).","github_created_at":"2023-09-09T01:53:38+00:00","created_at":"2026-07-11T23:13:28.18533+00:00","updated_at":"2026-08-02T06:00:22.352197+00:00","categories":[{"slug":"inference-serving","name":"Inference & Serving","url":"https://www.graphcanon.com/categories/inference-serving","markdown_url":"https://www.graphcanon.com/categories/inference-serving.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/inference-serving"}],"tags":[{"slug":"ai","name":"ai"},{"slug":"autoscaling","name":"autoscaling"},{"slug":"gpu","name":"gpu"},{"slug":"helm","name":"helm"},{"slug":"huggingface-runtime","name":"huggingface-runtime"},{"slug":"kubernetes","name":"kubernetes"},{"slug":"operator","name":"operator"},{"slug":"terraform","name":"terraform"}],"trust":{"provenance":{"is_fork":false,"github_id":689172674,"owner_type":"Organization","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-02T06:00:21.615Z","maintenance":{"label":"Very active","score":96,"methodology":"github_public_v1","releases_90d":2,"days_since_push":1,"last_release_at":"2026-07-01T04:21:06Z"},"security_summary":{"status":"findings","scanner":"osv@v1","low_count":2,"high_count":0,"last_scan_at":"2026-07-11T23:13:30.158Z","medium_count":0,"scan_profile":"deps","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-02T06:00:22.067Z"},"languages":{"value":["go","python"],"source":"github.language+pyproject.toml","observed_at":"2026-08-02T06:00:22.067Z"},"license_spdx":{"value":"Other","source":"github.license","observed_at":"2026-08-02T06:00:22.067Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":{"min_ram_gb":null,"requires_docker":true},"constraints":{"min_ram_gb":null,"requires_docker":true},"when_to_use":["When you need to integrate HuggingFace runtime for BYO models within your Kubernetes environment, as KAITO specifically supports models hosted there.","If auto-scaling is required for inference workloads, leveraging KEDA integration provides optimized performance and resource management."],"when_not_to_use":["Avoid if your organization prefers open-source model hosting that does not include HuggingFace; KAITO mandates use of the HuggingFace ecosystem.","Do not use when a custom autoscaling solution outside of KEDA is needed, as KAITO integrates tightly with KEDA for its scaling capabilities."],"source":"enrich:decision_facts","observed_at":"2026-07-16T19:26:17.210Z"},"constraint_facets":{"min_ram_gb":null,"requires_docker":true},"decision_summary":[{"label":"Requirements","value":"Requires Docker"},{"label":"Adopt for","value":"Kaito is a Kubernetes AI Toolchain Operator that facilitates the deployment and scaling of AI models in production environments using Helm or Terraform."},{"label":"License detail","value":"Under Apache License 2.0"}]}}