Configure Gateway
Configure API key authentication, model routing, traffic policies, and rate limits for the Gateway provided in your Namespace. You can apply only the configurations that you need, independently of one another.
- A Gateway with the same name as your Namespace has been created in advance.
- Replace values in angle brackets (
<...>) with values for your environment.
Components
| Configuration | Resource | Function |
|---|---|---|
| Authentication | Secret, SecurityPolicy | API key authentication |
| Inbound IP restriction | SecurityPolicy | Restricts IP addresses for each endpoint |
| External service connection | HTTPRoute | Routes a custom API externally |
| AI Gateway | Backend, AIServiceBackend, AIGatewayRoute | Routes a serving Service based on the hostname and model name |
| Traffic | ClientTrafficPolicy, BackendTrafficPolicy | Request buffering, timeouts, load balancing, circuit breaking, and retries |
| Rate limit | AIGatewayRoute, BackendTrafficPolicy | Limits based on the number of requests or token usage |
Configure API key authentication
Validate API keys on requests entering the Gateway. Create the API keys in advance, and then configure them as follows.
apiVersion: v1
kind: Secret
type: Opaque
metadata:
name: apikey-secret
namespace: <user-namespace>
stringData:
<key>: <value>
<key>: <value>
After creating the Secret as shown above, create a SecurityPolicy to apply API key authentication.
You can apply a SecurityPolicy only to an HTTPRoute. A Gateway-wide configuration is not supported.
Run the following command to list the HTTPRoutes that you can use.
kubectl get httproute -n <user-namespace>
Create a SecurityPolicy.
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: SecurityPolicy
metadata:
name: custom-sp
namespace: <user-namespace>
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: HTTPRoute
name: <http-route-name>
apiKeyAuth:
credentialRefs:
- group: ""
kind: Secret
name: <secret-name>
extractFrom:
- headers:
- authorization # Header name
If the API key is missing or does not match a registered value, the Gateway rejects the request with 401 Unauthorized.
curl -s https://<hostname>/v1/chat/completions \
-H "authorization: Bearer <api-key>" \
-H "content-type: application/json" \
-d '{"model":"<model-name>","messages":[{"role":"user","content":"Hello"}],"max_tokens":16}'
Do not commit YAML that contains secrets to a source repository. Use a separate API key for each user and rotate keys regularly.
Configure inbound IP restrictions
You can configure IP restrictions through a SecurityPolicy only for an HTTPRoute.
Add the authorization configuration to the existing SecurityPolicy.
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: SecurityPolicy
metadata:
name: custom-sp # Reuse the existing SecurityPolicy
namespace: <user-namespace>
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: HTTPRoute
name: <http-route-name>
apiKeyAuth: ... # Existing configuration
authorization: # Configuration to add
defaultAction: Deny
rules:
- action: Allow
principal:
clientCIDRs:
- <allowed-ip>/32
Configure HTTPRoute
To expose a separate API Service externally, create an HTTPRoute. An HTTPRoute connects the Service through the LB → Gateway → Service path.
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: <route-name>
namespace: <user-namespace>
spec:
hostnames:
- <hostname>
parentRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: <user-namespace>
namespace: kserve
sectionName: https
rules:
- backendRefs:
- group: ""
kind: Service
name: <service-name>
port: <service-port>
matches:
- path:
type: PathPrefix
value: /
timeouts:
request: 600s
After you create the HTTPRoute, the separately configured Service is connected to the Gateway and exposed externally. For <hostname>, enter a hostname in your assigned *.moduai.kakaocloud.com domain. If necessary, also apply Configure API key authentication and Configure inbound IP restrictions.
Configure AI Gateway
Expose a serving Service through an external hostname and route requests based on the model value in the request body. Create a Backend, AIServiceBackend, and AIGatewayRoute together.
Create Backend
A Backend defines the internal cluster address and port of the serving Service that receives requests.
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: Backend
metadata:
name: <backend-name>
namespace: <user-namespace>
spec:
endpoints:
- fqdn:
hostname: <serving-service-name>.<user-namespace>.svc.cluster.local
port: <serving-service-port>
Create AIServiceBackend
Define the API schema of a Backend that provides an OpenAI-compatible API.
apiVersion: aigateway.envoyproxy.io/v1beta1
kind: AIServiceBackend
metadata:
name: <ai-service-backend-name>
namespace: <user-namespace>
spec:
schema:
name: OpenAI
backendRef:
group: gateway.envoyproxy.io
kind: Backend
name: <backend-name>
Create AIGatewayRoute
An AIGatewayRoute routes requests to an AIServiceBackend based on the hostname and model name. The model value in the request body is extracted into the x-ai-eg-model header and used as a routing condition.
apiVersion: aigateway.envoyproxy.io/v1beta1
kind: AIGatewayRoute
metadata:
name: <ai-gateway-route-name>
namespace: <user-namespace>
spec:
hostnames:
- <hostname>
parentRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: <user-namespace>
namespace: kserve
rules:
- matches:
- headers:
- type: Exact
name: x-ai-eg-model
value: <model-name>
backendRefs:
- name: <ai-service-backend-name>
timeouts:
request: 1800s
When you create an AIGatewayRoute, an HTTPRoute with the same name is created automatically. The value in hostnames must match the hostname pattern of the Gateway listener.
Configure traffic
Because LLM requests can have large bodies and long response generation times, configure the request buffer and backend timeouts for the workload.
Configure backend traffic policy
BackendTrafficPolicy targets the HTTPRoute that is created automatically by AIGatewayRoute.
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: BackendTrafficPolicy
metadata:
name: <backend-traffic-policy-name>
namespace: <user-namespace>
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: HTTPRoute
name: <ai-gateway-route-name>
timeout:
http:
requestTimeout: 1800s
streamIdleTimeout: 1800s
maxStreamDuration: 3600s
tcp:
connectTimeout: 30s
loadBalancer:
type: LeastRequest
circuitBreaker:
maxConnections: 4096
maxPendingRequests: 4096
maxParallelRequests: 4096
maxParallelRetries: 16
retry:
numRetries: 1
perRetry:
timeout: 1800s
backOff:
baseInterval: 500ms
maxInterval: 5s
retryOn:
triggers:
- connect-failure
- reset
- unavailable
Configure rate limits
Use rateLimit.global in BackendTrafficPolicy to limit the number of requests or token usage for each client. Requests that exceed the limit are rejected with 429 Too Many Requests and the x-envoy-ratelimited: true header.
Request-based limit
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: BackendTrafficPolicy
metadata:
name: <backend-traffic-policy-name>
namespace: <user-namespace>
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: HTTPRoute
name: <ai-gateway-route-name>
rateLimit:
type: Global
global:
rules:
- clientSelectors:
- headers:
- name: <client-identifier-header>
type: Distinct
limit:
requests: <allowed-request-count>
unit: <time-unit>
Token-based limit
Use llmRequestCosts in AIGatewayRoute to record token usage in metadata.
spec:
llmRequestCosts:
- metadataKey: llm_input_token
type: InputToken
- metadataKey: llm_output_token
type: OutputToken
- metadataKey: llm_total_token
type: TotalToken
Reference the same metadata key from BackendTrafficPolicy to deduct from the token budget.
spec:
rateLimit:
type: Global
global:
rules:
- clientSelectors:
- headers:
- name: <client-identifier-header>
type: Distinct
limit:
requests: <allowed-token-count>
unit: <time-unit>
cost:
request:
from: Number
number: 0
response:
from: Metadata
metadata:
namespace: io.envoy.ai_gateway
key: llm_total_token
Token usage is deducted after the response completes. For a streaming request, tokens are deducted after the stream ends, and an active stream is not terminated when the limit is reached.
Check status
kubectl get gateway,httproute -n <user-namespace>
kubectl get backend,aiservicebackend,aigatewayroute -n <user-namespace>
kubectl get securitypolicy,clienttrafficpolicy,backendtrafficpolicy -n <user-namespace>
Verify that the Accepted and ResolvedRefs conditions of the HTTPRoute are True.
kubectl describe httproute <ai-gateway-route-name> -n <user-namespace>