{"data":{"slug":"last9-gpu-telemetry","name":"gpu-telemetry","tagline":"GPU Observability with Workload Attribution","github_url":"https://github.com/last9/gpu-telemetry","owner":"last9","repo":"gpu-telemetry","owner_avatar_url":"https://avatars.githubusercontent.com/u/53378302?v=4","primary_language":"Python","stars":66,"forks":8,"topics":["ai","amd","dcgm","gpu","gpu-monitoring","gpu-observability","helm","intel-gaudi-base-operator","kubernetes","llm-observability","mlops","nvidia","nvml-monitoring","opentelemetry","slurm","workload-attribution"],"archived":false,"github_pushed_at":"2026-08-02T09:12:57+00:00","maintenance_label":"Steady","stars_delta_30d":9,"url":"https://www.graphcanon.com/tools/last9-gpu-telemetry","markdown_url":"https://www.graphcanon.com/tools/last9-gpu-telemetry.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/last9-gpu-telemetry","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=last9-gpu-telemetry","description":"GPU Observability with workload attribution. One OTLP agent per node ties hardware metrics (NVIDIA, AMD, Intel Gaudi) to the K8s pod or Slurm job burning the GPU.","homepage_url":"https://last9.io/gpu-observability/","license":"MIT","open_issues":5,"watchers":3,"ai_summary":"Monitoring and observability for GPUs from NVIDIA, AMD, Intel Gaudi in Kubernetes and Slurm environments.","readme_excerpt":"## Quick Start — Bare Metal / systemd\n\n```bash\npip install l9gpu\nexport OTEL_EXPORTER_OTLP_ENDPOINT=<your-otlp-endpoint>\nexport OTEL_EXPORTER_OTLP_HEADERS=\"Authorization=Bearer <your-token>\"\n\nl9gpu nvml_monitor  --sink otel --cluster my-cluster  # NVIDIA\nl9gpu amd_monitor   --sink otel --cluster my-cluster  # AMD\nl9gpu gaudi_monitor --sink otel --cluster my-cluster  # Intel Gaudi\n```\n\nSanity-check without OTLP: `l9gpu nvml_monitor --sink stdout --once`.\n\nsystemd unit files: [`systemd/`](./systemd/).\n\n---\n\n---\n\n## Hardware support\n\nNVIDIA A100, H100 / H200, B200 / GB200, T4, A10, L4 (NVML + DCGM)  ·  AMD\nMI300X, MI325X (amdsmi)  ·  Intel Gaudi 2, Gaudi 3 (hl-smi).\n\nFull metric catalog with units and attributes: [`docs/METRICS.md`](./docs/METRICS.md).\n\n---\n\n---\n\n## License\n\nMIT for `l9gpu`, `k8shelper`, `k8sprocessor`. Apache-2.0 for `slurmprocessor`,\n`shelper`. Each subdirectory carries its own `LICENSE` where it differs from\nthe repo root.","github_created_at":"2026-04-19T17:21:55+00:00","created_at":"2026-07-15T10:41:54.818601+00:00","updated_at":"2026-09-20T04:25:38.521742+00:00","categories":[{"slug":"evaluation-observability","name":"Evaluation & Observability","url":"https://www.graphcanon.com/categories/evaluation-observability","markdown_url":"https://www.graphcanon.com/categories/evaluation-observability.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/evaluation-observability"}],"tags":[{"slug":"amd","name":"amd"},{"slug":"gpu-monitoring","name":"gpu-monitoring"},{"slug":"intel-gaudi-base-operator","name":"intel-gaudi-base-operator"},{"slug":"kubernetes","name":"kubernetes"},{"slug":"nvidia","name":"nvidia"},{"slug":"opentelemetry","name":"opentelemetry"},{"slug":"slurm","name":"slurm"}],"trust":{"provenance":{"is_fork":false,"github_id":1215254439,"owner_type":"Organization","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-09-11T06:00:19.673Z","maintenance":{"label":"Steady","score":60,"methodology":"github_public_v1","releases_90d":0,"days_since_push":39,"last_release_at":"2026-06-02T18:41:26Z","stars_delta_30d":9,"open_issues_delta_30d":0},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-15T10:41:56.246Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-09-11T06:00:20.169Z"},"has_cli":{"value":true,"source":"pyproject.toml:[project.scripts]","observed_at":"2026-09-11T06:00:20.169Z"},"languages":{"value":["python"],"source":"github.language+pyproject.toml","observed_at":"2026-09-11T06:00:20.169Z"},"license_spdx":{"value":"MIT","source":"github.license","observed_at":"2026-09-11T06:00:20.169Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":null,"constraints":null,"when_to_use":["When monitoring NVIDIA, AMD, or Intel Gaudi GPUs in Kubernetes clusters.","For workload attribution of GPU usage in Slurm clusters.","To integrate with OpenTelemetry for comprehensive telemetry."],"when_not_to_use":["If your infrastructure is not based on Kubernetes or Slurm.","When you prefer tools that do not require per-node OTLP agents.","For environments without support for NVIDIA, AMD, or Intel Gaudi GPUs."],"source":"enrich:decision_facts","observed_at":"2026-07-17T06:18:26.512Z"},"constraint_facets":null,"decision_summary":[{"label":"Adopt for","value":"gpu-telemetry provides comprehensive GPU observability in Kubernetes and Slurm environments by tying hardware metrics to the workload causing them."}]}}