Skip to main content

Check Overall GPU Status

The Overview page in AI Insight provides a summary of GPU resources and their status. In Overview, you can quickly identify GPUs in Warning, Critical, Idle, or Agent Missing status, then move to the GPU Explorer detail page to check the cause.

Note

To check metrics in AI Insight, Metric Exporter or the monitoring agent must be installed on the target resource. If it is not installed or is not working properly, the resource may appear in Agent Missing status.

Go to the Overview Menu

  1. Go to the KakaoCloud Console.
  2. Go to AI Service > AI Insight.
  3. Click Overview in the left menu.

Check Resource Summary

In Resource Summary, you can check summary information for all GPU resources.

ItemDescription
Total GPUsTotal number of GPUs in the query scope
Total ClustersTotal number of clusters in the query scope
Total NodesTotal number of nodes in the query scope
Average GPU LoadAverage load across all GPUs
Average GPU Memory UsageAverage memory usage across all GPUs
Average GPU TemperatureAverage temperature across all GPUs
ECC ErrorsNumber of ECC Errors that occurred in the last 24 hours

Check GPU Status

Check the number of GPUs by status in the status cards under Resource Summary.

StatusDescription
Active GPUGPU currently in normal use
Warning GPUGPU that requires attention
Critical GPUGPU in a severe abnormal state that requires inspection
Pending GPUGPU waiting because the node is in an inactive lifecycle state such as stopped, booting, rebooting, or resizing
Idle GPUGPU determined to be idle
Agent MissingGPU whose metric collection component is missing or not working properly
Caution

Metrics may not be collected properly for GPUs in Agent Missing status. In this case, check Metric Exporter Installation first.

Use GPU Map

In GPU Map, you can visually explore GPU resources.

  1. Select one of the GPU, Cluster, or Node tabs.
  2. Check Active, Idle, Pending, Warning, Critical, and Agent Missing statuses using the color legend.
  3. Click a resource in the map.
  4. In the details panel on the right, check status, GPU Flavor, GPU load, GPU memory usage, GPU temperature, ECC Error, XID Event Code, and Throttle Event.
  5. To go to the details page, click the move icon in the right panel.
FeatureDescription
Zoom in/outAdjusts the GPU Map display scale
Fit to screenFits the GPU Map to the screen size
Select resourceDisplays details for the selected GPU or MIG instance
Switch tabsChanges the map display unit by GPU, cluster, or node

Check GPU List

In the list at the bottom of Overview, you can check each GPU's status and key metrics.

ItemDescription
NameGPU or node name
StatusGPU status
GPU FlavorGPU instance or resource type
LoadGPU load
TempGPU temperature
MemoryGPU memory usage
XID Event CodeLast detected XID Event Code

Refresh Data

Overview data can be refreshed manually or automatically.

  1. Select a refresh interval from the Auto refresh dropdown at the top of the page.
  2. To refresh immediately, click the refresh icon.
  3. Check Last updated to verify that data has been refreshed.

If Data Is Not Displayed

If No data to display appears on the Overview page, check the following:

CategoryCheck ItemDescription
CommonProject and regionCheck whether the selected project and region are correct
CommonGPU resourcesCheck whether GPU resources exist in the target environment
CommonMetric Exporter or monitoring agentCheck the installation status of Metric Exporter or the monitoring agent
CommonAgent Missing statusCheck whether the target resource is in Agent Missing status
CommonTime range and refreshChange the time range or refresh the page, then query again

For detailed checks by cause, see AI Insight Troubleshooting.