Call the inference API
Call the inference API using the URL displayed on the endpoint details page and an API key connected to that endpoint.
Prerequisites
- Verify that the endpoint is
Active. - Create an API key and connect it to the endpoint you want to call.
- To apply usage limits to the API key, select a Tier for the endpoint connection.
- Check the API Schema configured for the model Revision.
Check endpoint information
- Go to AI Service > Advanced AI Space.
- Click Endpoints.
- Click the name of the endpoint you want to call.
- Find and copy the Endpoint URL under Details.
- Go to Credentials > API keys and check the Connected endpoints count for the API key you want to use.
- Click [Endpoint settings] in the API key's more options menu and verify that the endpoint you want to call is connected.
Send a request
Send a request to the URL shown on the endpoint details page, including the API key connected to that endpoint as the authentication credential. Follow the request format defined in API Schema(endpoint) for the model Revision.
Check the response
- If authentication fails, check whether the API key has expired and whether it is connected to the endpoint.
- If the request is rejected, check the Tier applied to the API key connection and the current quota usage.
- If the endpoint does not respond, verify that its status is
Active. - Check the model server logs on the endpoint's Logs tab.
- Check the request count and token usage on the Request/token usage tab.