Skip to main content

Create an endpoint

You can deploy a built-in or custom model from the model catalog to an endpoint. Configure the resources required to run the model, logging, and API limit policies when creating the endpoint.

Prerequisites​

  • Start the Advanced AI Space service in your project.
  • Identify the model to deploy. To use your own model, create a custom model first.
  • Determine the instance type, resource profile, and number of replicas required for your workload.

Open the creation page​

  1. Go to AI Service > Advanced AI Space.

  2. Open the endpoint creation page using one of the following methods.

    • From the Endpoints menu: Click [Create endpoint] in Endpoints, then select the model to deploy.
    • From the model list: Select a model in Model catalog > Built-in models or Model catalog > Custom models, then click [Create endpoint].
    • From the model details page: Click a model name, then click [Create endpoint] on its details page. The selected model is automatically populated on the creation page.

Configure basic information​

ItemDescription
ModelSelect the built-in or custom model to deploy.
The model is automatically populated when you open the creation page from a model list or details page.
Endpoint nameEnter 2–30 characters using English letters, numbers, hyphens (-), or underscores (_).
Domain IDThe domain ID to use in the endpoint URL.
Enter 1–20 characters using lowercase English letters, numbers, or hyphens (-).
A random DNS-safe suffix is added if the ID duplicates an existing endpoint's ID.

Configure resources​

ItemDescription
Instance typeSelect the instance type to run the model.
Replica countSelect the number of replicas for the endpoint.
Zero-downtime replica deploymentWhen enabled, existing replicas are terminated after new replicas are ready. Temporary additional resource usage during deployment may incur charges.
Resource profileSelect the resource profile to use on the selected instance.
Resource countSelect the number of resources per replica.
Check the maximum vCPU and memory for your selection.

Configure logging​

To save logs, enable Logging settings and select a bucket under Object Storage bucket information. Logs are automatically saved under the predictor path in the selected bucket. Daily folder names use UTC dates.

Configure API limits (Tier)​

API limits are optional. You can add multiple Tiers to an endpoint and configure the required limits for each Tier.

  1. Enable API limit settings.

  2. Enter the limits for each Tier.

    ItemDescription
    Tier nameA name to identify the Tier when connecting an API key.
    Total RequestThe total number of requests allowed by the Tier.
    Input TokenThe number of input tokens allowed by the Tier.
    Output TokenThe number of output tokens allowed by the Tier.
    Reset PeriodThe period for aggregating and resetting usage against the limits.
  3. Click [Add Tier] if you need additional Tiers.

You can also create a Tier without limits. In this case, do not configure a usage aggregation period.

Apply a Tier to an API key

A Tier is a policy registered on an endpoint. When connecting an API key to that endpoint, you can select one of its registered Tiers or connect without a Tier.

Complete creation​

  1. Review the settings and click [Create].
  2. Check the status in the endpoint list.
  3. When the status changes to Active, you can find the URL on the endpoint details page and send inference requests.

After creating the endpoint, you must create an API key and connect it to the endpoint before you can call it.