Gateway 구성
네임스페이스에 제공된 Gateway에 API 키 인증, 모델 라우팅, 트래픽 정책과 Rate Limit을 구성합니다. 각 구성은 필요한 항목만 독립적으로 적용할 수 있습니다.
- Gateway는 네임스페이스와 같은 이름으로 미리 생성되어 있습니다.
- 예시의
<...>값은 실제 환경에 맞게 변경합니다.
구성 요소
| 구성 | 리소스 | 기능 |
|---|---|---|
| 인증 | Secret, SecurityPolicy | API 키 인증 |
| Inbound IP 제한 | SecurityPolicy | Endpoint별 IP 제한 |
| 외부 서비스 연결 | HTTPRoute | 커스텀 API를 외부에 라우팅 |
| AI Gateway | Backend, AIServiceBackend, AIGatewayRoute | 서빙 Service를 호스트명과 모델 이름으로 라우팅 |
| 트래픽 | ClientTrafficPolicy, BackendTrafficPolicy | 요청 버퍼, 타임아웃, 로드밸런싱, 서킷브레이커, 재시도 |
| Rate Limit | AIGatewayRoute, BackendTrafficPolicy | 요청 수 또는 토큰 사용량 기반 제한 |
API 키 인증 구성
Gateway로 들어오는 요청의 API 키를 검증합니다. API 키를 미리 생성한 후 다음과 같이 설정합니다.
apiVersion: v1
kind: Secret
type: Opaque
metadata:
name: apikey-secret
namespace: <사용자 Namespace>
stringData:
<key>: <value>
<key>: <value>
위와 같이 생성한 이후, 해당 Key 인증을 적용할 SecurityPolicy를 생성합니다.
SecurityPolicy는 HTTPRoute에 한해서 설정이 가능하며 Gateway 전역 설정은 불가능합니다.
아래의 명령어로 사용할 HTTPRoute를 조회합니다.
kubectl get httproute -n <사용자 Namespace>
SecurityPolicy를 생성합니다.
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: SecurityPolicy
metadata:
name: custom-sp
namespace: <사용자 Namespace>
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: HTTPRoute
name: <적용할 HTTPRoute Name>
apiKeyAuth:
credentialRefs:
- group: ""
kind: Secret
name: <Secret 명>
extractFrom:
- headers:
- authorization # 설정할 헤더명
API 키가 없거나 등록된 값과 일치하지 않으면 Gateway가 401 Unauthorized로 요청을 거부합니다.
curl -s https://<호스트명>/v1/chat/completions \
-H "authorization: Bearer <API 키>" \
-H "content-type: application/json" \
-d '{"model":"<모델 이름>","messages":[{"role":"user","content":"안녕"}],"max_tokens":16}'
Secret이 포함된 YAML을 소스 저장소에 커밋하지 마세요. API 키는 사용자별로 구분하고 정기적으로 교체하는 것을 권장합니다.
Inbound IP 제한 구성
HTTPRoute에 한하여 SecurityPolicy를 통해 IP 제한을 구성할 수 있습니다.
단, 기존에 생성해놓은 SecurityPolicy에 authorization 설정을 추가하여 구성합니다.
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: SecurityPolicy
metadata:
name: custom-sp # 기존 SecurityPolicy 재사용
namespace: <사용자 Namespace>
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: HTTPRoute
name: <적용할 HTTPRoute Name>
apiKeyAuth: ... # 기존 설정 값
authorization: # 추가할 내용
defaultAction: Deny
rules:
- action: Allow
principal:
clientCIDRs:
- <허용할 IP>/32
HTTPRoute 구성
별도의 API Service를 외부에 공개하려면 HTTPRoute를 생성합니다. HTTPRoute를 사용하면 LB → Gateway → Service 경로로 연결됩니다.
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: <route-name>
namespace: <사용자 Namespace>
spec:
hostnames:
- <호스트명>
parentRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: <사용자 Namespace>
namespace: kserve
sectionName: https
rules:
- backendRefs:
- group: ""
kind: Service
name: <연결할 Service>
port: <연결할 Port>
matches:
- path:
type: PathPrefix
value: /
timeouts:
request: 600s
HTTPRoute를 생성하면 별도로 구성한 Service를 Gateway에 연결하여 외부에 공개할 수 있습니다. <호스트명>에는 할당받은 *.moduai.kakaocloud.com 도메인의 호스트명을 입력합니다. 필요한 경우 API 키 인증 구성과 Inbound IP 제한 구성을 추가로 적용합니다.
AI Gateway 구성
서빙 Service를 외부 호스트명으로 노출하고 요청 본문의 model 값에 따라 라우팅합니다. Backend, AIServiceBackend, AIGatewayRoute를 함께 생성합니다.
Backend 생성
Backend는 요청을 전달할 서빙 Service의 클러스터 내부 주소와 포트를 정의합니다.
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: Backend
metadata:
name: <Backend 이름>
namespace: <사용자 namespace>
spec:
endpoints:
- fqdn:
hostname: <서빙 Service 이름>.<사용자 namespace>.svc.cluster.local
port: <서빙 Service 포트>
AIServiceBackend 생성
OpenAI 호환 API를 제공하는 Backend의 API 스키마를 정의합니다.
apiVersion: aigateway.envoyproxy.io/v1beta1
kind: AIServiceBackend
metadata:
name: <AIServiceBackend 이름>
namespace: <사용자 namespace>
spec:
schema:
name: OpenAI
backendRef:
group: gateway.envoyproxy.io
kind: Backend
name: <Backend 이름>
AIGatewayRoute 생성
AIGatewayRoute는 호스트명과 모델 이름을 기준으로 요청을 AIServiceBackend에 전달합니다. 요청 본문의 model 값은 x-ai-eg-model 헤더로 추출되며 라우팅 조건에 사용됩니다.
apiVersion: aigateway.envoyproxy.io/v1beta1
kind: AIGatewayRoute
metadata:
name: <AIGatewayRoute 이름>
namespace: <사용자 namespace>
spec:
hostnames:
- <호스트명>
parentRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: <사용자 namespace>
namespace: kserve
rules:
- matches:
- headers:
- type: Exact
name: x-ai-eg-model
value: <모델 이름>
backendRefs:
- name: <AIServiceBackend 이름>
timeouts:
request: 1800s
AIGatewayRoute를 생성하면 같은 이름의 HTTPRoute가 자동으로 생성됩니다. hostnames는 Gateway 리스너의 호스트명 패턴에 포함되어야 합니다.
트래픽 구성
LLM 요청은 본문이 크고 응답 생성 시간이 길기 때문에 요청 버퍼와 백엔드 타임아웃을 워크로드 특성에 맞게 설정합니다.
백엔드 트래픽 정책 설정
BackendTrafficPolicy는 AIGatewayRoute가 자동으로 생성한 HTTPRoute를 대상으로 합니다.
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: BackendTrafficPolicy
metadata:
name: <BackendTrafficPolicy 이름>
namespace: <사용자 namespace>
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: HTTPRoute
name: <AIGatewayRoute 이름>
timeout:
http:
requestTimeout: 1800s
streamIdleTimeout: 1800s
maxStreamDuration: 3600s
tcp:
connectTimeout: 30s
loadBalancer:
type: LeastRequest
circuitBreaker:
maxConnections: 4096
maxPendingRequests: 4096
maxParallelRequests: 4096
maxParallelRetries: 16
retry:
numRetries: 1
perRetry:
timeout: 1800s
backOff:
baseInterval: 500ms
maxInterval: 5s
retryOn:
triggers:
- connect-failure
- reset
- unavailable
Rate Limit 구성
BackendTrafficPolicy의 rateLimit.global로 클라이언트별 요청 수 또는 토큰 사용량을 제한합니다. 제한을 초과한 요청은 429 Too Many Requests와 x-envoy-ratelimited: true 헤더로 거부됩니다.
요청 수 기반 제한
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: BackendTrafficPolicy
metadata:
name: <BackendTrafficPolicy 이름>
namespace: <사용자 namespace>
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: HTTPRoute
name: <AIGatewayRoute 이름>
rateLimit:
type: Global
global:
rules:
- clientSelectors:
- headers:
- name: <클라이언트 식별 헤더>
type: Distinct
limit:
requests: <허용 요청 수>
unit: <단위 시간>
토큰 사용량 기반 제한
AIGatewayRoute의 llmRequestCosts로 토큰 사용량을 메타데이터에 기록합니다.
spec:
llmRequestCosts:
- metadataKey: llm_input_token
type: InputToken
- metadataKey: llm_output_token
type: OutputToken
- metadataKey: llm_total_token
type: TotalToken
BackendTrafficPolicy에서 같은 메타데이터 키를 참조하여 토큰 예산을 차감합니다.
spec:
rateLimit:
type: Global
global:
rules:
- clientSelectors:
- headers:
- name: <클라이언트 식별 헤더>
type: Distinct
limit:
requests: <허용 토큰 수>
unit: <단위 시간>
cost:
request:
from: Number
number: 0
response:
from: Metadata
metadata:
namespace: io.envoy.ai_gateway
key: llm_total_token
토큰 사용량은 응답이 완료된 후 차감됩니다. 스트리밍 요청은 스트림이 끝난 후 차감되며, 진행 중인 스트림을 중간에 종료하지 않습니다.
상태 확인
kubectl get gateway,httproute -n <사용자 namespace>
kubectl get backend,aiservicebackend,aigatewayroute -n <사용자 namespace>
kubectl get securitypolicy,clienttrafficpolicy,backendtrafficpolicy -n <사용자 namespace>
HTTPRoute의 Accepted와 ResolvedRefs 조건이 True인지 확인합니다.
kubectl describe httproute <AIGatewayRoute 이름> -n <사용자 namespace>