Create an endpoint
You can deploy a built-in or custom model from the model catalog to an endpoint. Configure the resources required to run the model, logging, and API limit policies when creating the endpoint.
Prerequisites
- Start the Advanced AI Space service in your project.
- Identify the model to deploy. To use your own model, create a custom model first.
- Determine the instance type, resource profile, and number of replicas required for your workload.
Open the creation page
-
Go to AI Service > Advanced AI Space.
-
Open the endpoint creation page using one of the following methods.
- From the Endpoints menu: Click [Create endpoint] in Endpoints, then select the model to deploy.
- From the model list: Select a model in Model catalog > Built-in models or Model catalog > Custom models, then click [Create endpoint].
- From the model details page: Click a model name, then click [Create endpoint] on its details page. The selected model is automatically populated on the creation page.
Configure basic information
| Item | Description |
|---|---|
| Model | Select the built-in or custom model to deploy. The model is automatically populated when you open the creation page from a model list or details page. |
| Endpoint name | Enter 2–30 characters using English letters, numbers, hyphens (-), or underscores (_). |
| Domain ID | The domain ID to use in the endpoint URL. Enter 1–20 characters using lowercase English letters, numbers, or hyphens ( -). A random DNS-safe suffix is added if the ID duplicates an existing endpoint's ID. |
Configure resources
| Item | Description |
|---|---|
| Instance type | Select the instance type to run the model. |
| Replica count | Select the number of replicas for the endpoint. |
| Zero-downtime replica deployment | When enabled, existing replicas are terminated after new replicas are ready. Temporary additional resource usage during deployment may incur charges. |
| Resource profile | Select the resource profile to use on the selected instance. |
| Resource count | Select the number of resources per replica. Check the maximum vCPU and memory for your selection. |
Configure logging
To save logs, enable Logging settings and select a bucket under Object Storage bucket information. Logs are automatically saved under the predictor path in the selected bucket. Daily folder names use UTC dates.
Configure API limits (Tier)
API limits are optional. You can add multiple Tiers to an endpoint and configure the required limits for each Tier.
-
Enable API limit settings.
-
Enter the limits for each Tier.
Item Description Tier name A name to identify the Tier when connecting an API key. Total Request The total number of requests allowed by the Tier. Input Token The number of input tokens allowed by the Tier. Output Token The number of output tokens allowed by the Tier. Reset Period The period for aggregating and resetting usage against the limits. -
Click [Add Tier] if you need additional Tiers.
You can also create a Tier without limits. In this case, do not configure a usage aggregation period.
A Tier is a policy registered on an endpoint. When connecting an API key to that endpoint, you can select one of its registered Tiers or connect without a Tier.
Complete creation
- Review the settings and click [Create].
- Check the status in the endpoint list.
- When the status changes to
Active, you can find the URL on the endpoint details page and send inference requests.
After creating the endpoint, you must create an API key and connect it to the endpoint before you can call it.