Check Overall GPU Status
The Overview page in AI Insight provides a summary of GPU resources and their status. In Overview, you can quickly identify GPUs in Warning, Critical, Idle, or Agent Missing status, then move to the GPU Explorer detail page to check the cause.
To check metrics in AI Insight, Metric Exporter or the monitoring agent must be installed on the target resource. If it is not installed or is not working properly, the resource may appear in Agent Missing status.
Go to the Overview Menu
- Go to the KakaoCloud Console.
- Go to AI Service > AI Insight.
- Click Overview in the left menu.
Check Resource Summary
In Resource Summary, you can check summary information for all GPU resources.
| Item | Description |
|---|---|
| Total GPUs | Total number of GPUs in the query scope |
| Total Clusters | Total number of clusters in the query scope |
| Total Nodes | Total number of nodes in the query scope |
| Average GPU Load | Average load across all GPUs |
| Average GPU Memory Usage | Average memory usage across all GPUs |
| Average GPU Temperature | Average temperature across all GPUs |
| ECC Errors | Number of ECC Errors that occurred in the last 24 hours |
Check GPU Status
Check the number of GPUs by status in the status cards under Resource Summary.
| Status | Description |
|---|---|
| Active GPU | GPU currently in normal use |
| Warning GPU | GPU that requires attention |
| Critical GPU | GPU in a severe abnormal state that requires inspection |
| Pending GPU | GPU waiting because the node is in an inactive lifecycle state such as stopped, booting, rebooting, or resizing |
| Idle GPU | GPU determined to be idle |
| Agent Missing | GPU whose metric collection component is missing or not working properly |
Metrics may not be collected properly for GPUs in Agent Missing status. In this case, check Metric Exporter Installation first.
Use GPU Map
In GPU Map, you can visually explore GPU resources.
- Select one of the GPU, Cluster, or Node tabs.
- Check Active, Idle, Pending, Warning, Critical, and Agent Missing statuses using the color legend.
- Click a resource in the map.
- In the details panel on the right, check status, GPU Flavor, GPU load, GPU memory usage, GPU temperature, ECC Error, XID Event Code, and Throttle Event.
- To go to the details page, click the move icon in the right panel.
| Feature | Description |
|---|---|
| Zoom in/out | Adjusts the GPU Map display scale |
| Fit to screen | Fits the GPU Map to the screen size |
| Select resource | Displays details for the selected GPU or MIG instance |
| Switch tabs | Changes the map display unit by GPU, cluster, or node |
Check GPU List
In the list at the bottom of Overview, you can check each GPU's status and key metrics.
| Item | Description |
|---|---|
| Name | GPU or node name |
| Status | GPU status |
| GPU Flavor | GPU instance or resource type |
| Load | GPU load |
| Temp | GPU temperature |
| Memory | GPU memory usage |
| XID Event Code | Last detected XID Event Code |
Refresh Data
Overview data can be refreshed manually or automatically.
- Select a refresh interval from the Auto refresh dropdown at the top of the page.
- To refresh immediately, click the refresh icon.
- Check Last updated to verify that data has been refreshed.
If Data Is Not Displayed
If No data to display appears on the Overview page, check the following:
| Category | Check Item | Description |
|---|---|---|
| Common | Project and region | Check whether the selected project and region are correct |
| Common | GPU resources | Check whether GPU resources exist in the target environment |
| Common | Metric Exporter or monitoring agent | Check the installation status of Metric Exporter or the monitoring agent |
| Common | Agent Missing status | Check whether the target resource is in Agent Missing status |
| Common | Time range and refresh | Change the time range or refresh the page, then query again |
For detailed checks by cause, see AI Insight Troubleshooting.