View endpoint details and monitoring
The endpoint details page provides basic information, logs, detailed settings, request and token usage, and system resource metrics.
-
Go to AI Service > Advanced AI Space.
-
Click Endpoints.
-
Click the name of the endpoint you want to inspect.
-
Select the section you want to view on the endpoint details page.
Section Description Details Status, model information, domain ID, endpoint URL, instance type, profile information, replica count, and creation date and time. Logs Logs from the Pods that make up the endpoint and from model execution. Detailed settings Whether logging is enabled and the API limit settings (Tiers) registered on the endpoint. Request/token usage Request counts, token usage, and quota usage by endpoint and API key. Monitoring GPU, CPU, memory, and network metrics for each Pod.
View logs
Select a Pod from the Pod list. The most recent 1,000 log lines are displayed, and you can view subsequent logs in real time.
| Log | Description |
|---|---|
| Pod logs | Logs for model server startup and inference request processing. |
| Storage Initializer | Logs for preparing model files, container images, and other items required to run the endpoint during startup. |
For endpoints that use Object Storage logging, you can check the logging configuration and bucket information under Detailed settings.
View detailed settings
Check the following under Detailed settings:
- Whether Object Storage logging is enabled
- API limit settings (Tiers) registered on the endpoint
- Total request, input token, and output token limits and reset settings for each Tier
To change the settings, click [Edit] in the endpoint's more options menu.
Request/token usage
Select a time range to view the following metrics:
- Input Token
- Output Token
- Total Token
- Total Request
- Error Rate
- Cache Hit Rate
Under Usage by API key, you can view the names of API keys connected to the endpoint, their applied Tiers, token usage, total request counts, and quota usage. Quota usage is not displayed for connections without a Tier.
Monitoring
Select a Pod, time range, and refresh settings to view the following resource metrics:
- GPU usage and GPU memory
- CPU usage
- Memory usage
- Network traffic sent and received
Change the time range or refresh the metrics to check the endpoint's current status and trends.