GraphCanon updated Sep 11, 2026 · GitHub synced Sep 11, 2026
33views this month
Decision brief
gpu-telemetry provides comprehensive GPU observability in Kubernetes and Slurm environments by tying hardware metrics to the workload causing them.
Good fit when
- When monitoring NVIDIA, AMD, or Intel Gaudi GPUs in Kubernetes clusters.
- For workload attribution of GPU usage in Slurm clusters.
Avoid when
- If your infrastructure is not based on Kubernetes or Slurm.
- When you prefer tools that do not require per-node OTLP agents.
Observed Jul 17, 2026 · Source: enrich:decision_facts
Verify the decision
Maintenance and security
Full trust report- Maintenance
- Steady (39d since push)
- As of Sep 11, 2026
- Provenance
- Not a fork · Organization account
- As of Sep 11, 2026
- Security (OSV)
- No lockfile
- As of Jul 15, 2026
Public GitHub metadata and optional OSV scans. Signals, not a guarantee. Trust methodology.
Install
pip install gpu-telemetry PyPISimilar tools
Same-category neighbours. No typed graph edges are catalogued for this tool yet.
Evidence and technical details
Sourced facts, taxonomy, compatibility claims, README excerpt, and machine-readable endpoints.
Overview
Monitoring and observability for GPUs from NVIDIA, AMD, Intel Gaudi in Kubernetes and Slurm environments.
Capability facts
- CLI
- CLI entrypoint
Source: pyproject.toml:[project.scripts] · Sep 11, 2026
- Languages
- python
Source: github.language+pyproject.toml · Sep 11, 2026
Categories
Compatibility
Sourced claims from the README excerpt - not unsourced marketing copy.
Tags
README
Quick Start — Bare Metal / systemd Sanity check without OTLP: . systemd unit files: . Hardware support NVIDIA A100, H100 / H200, B200 / GB200, T4, A10, L4 (NVML + DCGM) · AMD MI300X, MI325X (amdsmi) · Intel Gaudi 2, Gaudi 3 (hl smi). Full metric catalog with units and attributes:...
For agents
This page has a .md twin and JSON over the API.