Enable GPU Metrics

Turn on GPU monitoring when you need utilization, memory, temperature, power, and health telemetry for GPU workloads in ACP.

Decide the metrics owner

SituationDo this
Alauda GPU Management manages the GPU nodesDCGM-Exporter is enabled by default in the ClusterPolicy. No separate install is needed.
A HAMi backend exposes the GPUsInstall Alauda Build of DCGM-Exporter as a standalone plugin for those nodes.
Metrics are currently disabledEnable the DCGM-Exporter component in the ClusterPolicy and reconcile.

For the install and verification mechanics in each case, see DCGM-Exporter.

Confirm metrics are collected

Check that the exporter Pods are running on the GPU nodes:

kubectl get pods -l app=nvidia-dcgm-exporter -A

View dashboards

After the exporter has run for a short time, open Administrator > Operations Center > Monitor > Dashboards and switch to the GPU node and pod panels. ACP owns the dashboards; this product provides the metrics source.

For dashboard management in ACP, see Manage dashboards.