Skip to main content

Configure Gateway

Configure API key authentication, model routing, traffic policies, and rate limits for the Gateway provided in your Namespace. You can apply only the configurations that you need, independently of one another.

Before you begin
  • A Gateway with the same name as your Namespace has been created in advance.
  • Replace values in angle brackets (<...>) with values for your environment.

Components​

ConfigurationResourceFunction
AuthenticationSecret, SecurityPolicyAPI key authentication
Inbound IP restrictionSecurityPolicyRestricts IP addresses for each endpoint
External service connectionHTTPRouteRoutes a custom API externally
AI GatewayBackend, AIServiceBackend, AIGatewayRouteRoutes a serving Service based on the hostname and model name
TrafficClientTrafficPolicy, BackendTrafficPolicyRequest buffering, timeouts, load balancing, circuit breaking, and retries
Rate limitAIGatewayRoute, BackendTrafficPolicyLimits based on the number of requests or token usage

Configure API key authentication​

Validate API keys on requests entering the Gateway. Create the API keys in advance, and then configure them as follows.

Create a Secret to store API keys
apiVersion: v1
kind: Secret
type: Opaque
metadata:
name: apikey-secret
namespace: <user-namespace>
stringData:
<key>: <value>
<key>: <value>

After creating the Secret as shown above, create a SecurityPolicy to apply API key authentication.

You can apply a SecurityPolicy only to an HTTPRoute. A Gateway-wide configuration is not supported.

Run the following command to list the HTTPRoutes that you can use.

List HTTPRoutes
kubectl get httproute -n <user-namespace>

Create a SecurityPolicy.

Create a SecurityPolicy
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: SecurityPolicy
metadata:
name: custom-sp
namespace: <user-namespace>
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: HTTPRoute
name: <http-route-name>
apiKeyAuth:
credentialRefs:
- group: ""
kind: Secret
name: <secret-name>
extractFrom:
- headers:
- authorization # Header name

If the API key is missing or does not match a registered value, the Gateway rejects the request with 401 Unauthorized.

curl -s https://<hostname>/v1/chat/completions \
-H "authorization: Bearer <api-key>" \
-H "content-type: application/json" \
-d '{"model":"<model-name>","messages":[{"role":"user","content":"Hello"}],"max_tokens":16}'
Caution

Do not commit YAML that contains secrets to a source repository. Use a separate API key for each user and rotate keys regularly.

Configure inbound IP restrictions​

You can configure IP restrictions through a SecurityPolicy only for an HTTPRoute.

Add the authorization configuration to the existing SecurityPolicy.

apiVersion: gateway.envoyproxy.io/v1alpha1
kind: SecurityPolicy
metadata:
name: custom-sp # Reuse the existing SecurityPolicy
namespace: <user-namespace>
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: HTTPRoute
name: <http-route-name>
apiKeyAuth: ... # Existing configuration
authorization: # Configuration to add
defaultAction: Deny
rules:
- action: Allow
principal:
clientCIDRs:
- <allowed-ip>/32

Configure HTTPRoute​

To expose a separate API Service externally, create an HTTPRoute. An HTTPRoute connects the Service through the LB → Gateway → Service path.

Create an HTTPRoute
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: <route-name>
namespace: <user-namespace>
spec:
hostnames:
- <hostname>
parentRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: <user-namespace>
namespace: kserve
sectionName: https
rules:
- backendRefs:
- group: ""
kind: Service
name: <service-name>
port: <service-port>
matches:
- path:
type: PathPrefix
value: /
timeouts:
request: 600s

After you create the HTTPRoute, the separately configured Service is connected to the Gateway and exposed externally. For <hostname>, enter a hostname in your assigned *.moduai.kakaocloud.com domain. If necessary, also apply Configure API key authentication and Configure inbound IP restrictions.

Configure AI Gateway​

Expose a serving Service through an external hostname and route requests based on the model value in the request body. Create a Backend, AIServiceBackend, and AIGatewayRoute together.

Create Backend​

A Backend defines the internal cluster address and port of the serving Service that receives requests.

backend.yaml
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: Backend
metadata:
name: <backend-name>
namespace: <user-namespace>
spec:
endpoints:
- fqdn:
hostname: <serving-service-name>.<user-namespace>.svc.cluster.local
port: <serving-service-port>

Create AIServiceBackend​

Define the API schema of a Backend that provides an OpenAI-compatible API.

ai-service-backend.yaml
apiVersion: aigateway.envoyproxy.io/v1beta1
kind: AIServiceBackend
metadata:
name: <ai-service-backend-name>
namespace: <user-namespace>
spec:
schema:
name: OpenAI
backendRef:
group: gateway.envoyproxy.io
kind: Backend
name: <backend-name>

Create AIGatewayRoute​

An AIGatewayRoute routes requests to an AIServiceBackend based on the hostname and model name. The model value in the request body is extracted into the x-ai-eg-model header and used as a routing condition.

ai-gateway-route.yaml
apiVersion: aigateway.envoyproxy.io/v1beta1
kind: AIGatewayRoute
metadata:
name: <ai-gateway-route-name>
namespace: <user-namespace>
spec:
hostnames:
- <hostname>
parentRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: <user-namespace>
namespace: kserve
rules:
- matches:
- headers:
- type: Exact
name: x-ai-eg-model
value: <model-name>
backendRefs:
- name: <ai-service-backend-name>
timeouts:
request: 1800s

When you create an AIGatewayRoute, an HTTPRoute with the same name is created automatically. The value in hostnames must match the hostname pattern of the Gateway listener.

Configure traffic​

Because LLM requests can have large bodies and long response generation times, configure the request buffer and backend timeouts for the workload.

Configure backend traffic policy​

BackendTrafficPolicy targets the HTTPRoute that is created automatically by AIGatewayRoute.

backend-traffic-policy.yaml
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: BackendTrafficPolicy
metadata:
name: <backend-traffic-policy-name>
namespace: <user-namespace>
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: HTTPRoute
name: <ai-gateway-route-name>
timeout:
http:
requestTimeout: 1800s
streamIdleTimeout: 1800s
maxStreamDuration: 3600s
tcp:
connectTimeout: 30s
loadBalancer:
type: LeastRequest
circuitBreaker:
maxConnections: 4096
maxPendingRequests: 4096
maxParallelRequests: 4096
maxParallelRetries: 16
retry:
numRetries: 1
perRetry:
timeout: 1800s
backOff:
baseInterval: 500ms
maxInterval: 5s
retryOn:
triggers:
- connect-failure
- reset
- unavailable

Configure rate limits​

Use rateLimit.global in BackendTrafficPolicy to limit the number of requests or token usage for each client. Requests that exceed the limit are rejected with 429 Too Many Requests and the x-envoy-ratelimited: true header.

Request-based limit​

request-rate-limit.yaml
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: BackendTrafficPolicy
metadata:
name: <backend-traffic-policy-name>
namespace: <user-namespace>
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: HTTPRoute
name: <ai-gateway-route-name>
rateLimit:
type: Global
global:
rules:
- clientSelectors:
- headers:
- name: <client-identifier-header>
type: Distinct
limit:
requests: <allowed-request-count>
unit: <time-unit>

Token-based limit​

Use llmRequestCosts in AIGatewayRoute to record token usage in metadata.

spec:
llmRequestCosts:
- metadataKey: llm_input_token
type: InputToken
- metadataKey: llm_output_token
type: OutputToken
- metadataKey: llm_total_token
type: TotalToken

Reference the same metadata key from BackendTrafficPolicy to deduct from the token budget.

spec:
rateLimit:
type: Global
global:
rules:
- clientSelectors:
- headers:
- name: <client-identifier-header>
type: Distinct
limit:
requests: <allowed-token-count>
unit: <time-unit>
cost:
request:
from: Number
number: 0
response:
from: Metadata
metadata:
namespace: io.envoy.ai_gateway
key: llm_total_token
When tokens are deducted

Token usage is deducted after the response completes. For a streaming request, tokens are deducted after the stream ends, and an active stream is not terminated when the limit is reached.

Check status​

kubectl get gateway,httproute -n <user-namespace>
kubectl get backend,aiservicebackend,aigatewayroute -n <user-namespace>
kubectl get securitypolicy,clienttrafficpolicy,backendtrafficpolicy -n <user-namespace>

Verify that the Accepted and ResolvedRefs conditions of the HTTPRoute are True.

kubectl describe httproute <ai-gateway-route-name> -n <user-namespace>