Skip to main content

View endpoint details and monitoring

The endpoint details page provides basic information, logs, detailed settings, request and token usage, and system resource metrics.

  1. Go to AI Service > Advanced AI Space.

  2. Click Endpoints.

  3. Click the name of the endpoint you want to inspect.

  4. Select the section you want to view on the endpoint details page.

    SectionDescription
    DetailsStatus, model information, domain ID, endpoint URL, instance type, profile information, replica count, and creation date and time.
    LogsLogs from the Pods that make up the endpoint and from model execution.
    Detailed settingsWhether logging is enabled and the API limit settings (Tiers) registered on the endpoint.
    Request/token usageRequest counts, token usage, and quota usage by endpoint and API key.
    MonitoringGPU, CPU, memory, and network metrics for each Pod.

View logs​

Select a Pod from the Pod list. The most recent 1,000 log lines are displayed, and you can view subsequent logs in real time.

LogDescription
Pod logsLogs for model server startup and inference request processing.
Storage InitializerLogs for preparing model files, container images, and other items required to run the endpoint during startup.

For endpoints that use Object Storage logging, you can check the logging configuration and bucket information under Detailed settings.

View detailed settings​

Check the following under Detailed settings:

  • Whether Object Storage logging is enabled
  • API limit settings (Tiers) registered on the endpoint
  • Total request, input token, and output token limits and reset settings for each Tier

To change the settings, click [Edit] in the endpoint's more options menu.

Request/token usage​

Select a time range to view the following metrics:

  • Input Token
  • Output Token
  • Total Token
  • Total Request
  • Error Rate
  • Cache Hit Rate

Under Usage by API key, you can view the names of API keys connected to the endpoint, their applied Tiers, token usage, total request counts, and quota usage. Quota usage is not displayed for connections without a Tier.

Monitoring​

Select a Pod, time range, and refresh settings to view the following resource metrics:

  • GPU usage and GPU memory
  • CPU usage
  • Memory usage
  • Network traffic sent and received

Change the time range or refresh the metrics to check the endpoint's current status and trends.