본문으로 건너뛰기

Gateway 구성

네임스페이스에 제공된 Gateway에 API 키 인증, 모델 라우팅, 트래픽 정책과 Rate Limit을 구성합니다. 각 구성은 필요한 항목만 독립적으로 적용할 수 있습니다.

시작하기 전에
  • Gateway는 네임스페이스와 같은 이름으로 미리 생성되어 있습니다.
  • 예시의 <...> 값은 실제 환경에 맞게 변경합니다.

구성 요소​

구성리소스기능
인증Secret, SecurityPolicyAPI 키 인증
Inbound IP 제한SecurityPolicyEndpoint별 IP 제한
외부 서비스 연결HTTPRoute커스텀 API를 외부에 라우팅
AI GatewayBackend, AIServiceBackend, AIGatewayRoute서빙 Service를 호스트명과 모델 이름으로 라우팅
트래픽ClientTrafficPolicy, BackendTrafficPolicy요청 버퍼, 타임아웃, 로드밸런싱, 서킷브레이커, 재시도
Rate LimitAIGatewayRoute, BackendTrafficPolicy요청 수 또는 토큰 사용량 기반 제한

API 키 인증 구성​

Gateway로 들어오는 요청의 API 키를 검증합니다. API 키를 미리 생성한 후 다음과 같이 설정합니다.

API-Key 저장을 위한 Secret 생성
apiVersion: v1
kind: Secret
type: Opaque
metadata:
name: apikey-secret
namespace: <사용자 Namespace>
stringData:
<key>: <value>
<key>: <value>

위와 같이 생성한 이후, 해당 Key 인증을 적용할 SecurityPolicy를 생성합니다.

SecurityPolicy는 HTTPRoute에 한해서 설정이 가능하며 Gateway 전역 설정은 불가능합니다.

아래의 명령어로 사용할 HTTPRoute를 조회합니다.

HTTPRoute 조회
kubectl get httproute -n <사용자 Namespace>

SecurityPolicy를 생성합니다.

SecurityPolicy 생성
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: SecurityPolicy
metadata:
name: custom-sp
namespace: <사용자 Namespace>
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: HTTPRoute
name: <적용할 HTTPRoute Name>
apiKeyAuth:
credentialRefs:
- group: ""
kind: Secret
name: <Secret 명>
extractFrom:
- headers:
- authorization # 설정할 헤더명

API 키가 없거나 등록된 값과 일치하지 않으면 Gateway가 401 Unauthorized로 요청을 거부합니다.

curl -s https://<호스트명>/v1/chat/completions \
-H "authorization: Bearer <API 키>" \
-H "content-type: application/json" \
-d '{"model":"<모델 이름>","messages":[{"role":"user","content":"안녕"}],"max_tokens":16}'
주의

Secret이 포함된 YAML을 소스 저장소에 커밋하지 마세요. API 키는 사용자별로 구분하고 정기적으로 교체하는 것을 권장합니다.

Inbound IP 제한 구성​

HTTPRoute에 한하여 SecurityPolicy를 통해 IP 제한을 구성할 수 있습니다.

단, 기존에 생성해놓은 SecurityPolicy에 authorization 설정을 추가하여 구성합니다.

apiVersion: gateway.envoyproxy.io/v1alpha1
kind: SecurityPolicy
metadata:
name: custom-sp # 기존 SecurityPolicy 재사용
namespace: <사용자 Namespace>
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: HTTPRoute
name: <적용할 HTTPRoute Name>
apiKeyAuth: ... # 기존 설정 값
authorization: # 추가할 내용
defaultAction: Deny
rules:
- action: Allow
principal:
clientCIDRs:
- <허용할 IP>/32

HTTPRoute 구성​

별도의 API Service를 외부에 공개하려면 HTTPRoute를 생성합니다. HTTPRoute를 사용하면 LB → Gateway → Service 경로로 연결됩니다.

HTTPRoute 생성
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: <route-name>
namespace: <사용자 Namespace>
spec:
hostnames:
- <호스트명>
parentRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: <사용자 Namespace>
namespace: kserve
sectionName: https
rules:
- backendRefs:
- group: ""
kind: Service
name: <연결할 Service>
port: <연결할 Port>
matches:
- path:
type: PathPrefix
value: /
timeouts:
request: 600s

HTTPRoute를 생성하면 별도로 구성한 Service를 Gateway에 연결하여 외부에 공개할 수 있습니다. <호스트명>에는 할당받은 *.moduai.kakaocloud.com 도메인의 호스트명을 입력합니다. 필요한 경우 API 키 인증 구성과 Inbound IP 제한 구성을 추가로 적용합니다.

AI Gateway 구성​

서빙 Service를 외부 호스트명으로 노출하고 요청 본문의 model 값에 따라 라우팅합니다. Backend, AIServiceBackend, AIGatewayRoute를 함께 생성합니다.

Backend 생성​

Backend는 요청을 전달할 서빙 Service의 클러스터 내부 주소와 포트를 정의합니다.

backend.yaml
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: Backend
metadata:
name: <Backend 이름>
namespace: <사용자 namespace>
spec:
endpoints:
- fqdn:
hostname: <서빙 Service 이름>.<사용자 namespace>.svc.cluster.local
port: <서빙 Service 포트>

AIServiceBackend 생성​

OpenAI 호환 API를 제공하는 Backend의 API 스키마를 정의합니다.

ai-service-backend.yaml
apiVersion: aigateway.envoyproxy.io/v1beta1
kind: AIServiceBackend
metadata:
name: <AIServiceBackend 이름>
namespace: <사용자 namespace>
spec:
schema:
name: OpenAI
backendRef:
group: gateway.envoyproxy.io
kind: Backend
name: <Backend 이름>

AIGatewayRoute 생성​

AIGatewayRoute는 호스트명과 모델 이름을 기준으로 요청을 AIServiceBackend에 전달합니다. 요청 본문의 model 값은 x-ai-eg-model 헤더로 추출되며 라우팅 조건에 사용됩니다.

ai-gateway-route.yaml
apiVersion: aigateway.envoyproxy.io/v1beta1
kind: AIGatewayRoute
metadata:
name: <AIGatewayRoute 이름>
namespace: <사용자 namespace>
spec:
hostnames:
- <호스트명>
parentRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: <사용자 namespace>
namespace: kserve
rules:
- matches:
- headers:
- type: Exact
name: x-ai-eg-model
value: <모델 이름>
backendRefs:
- name: <AIServiceBackend 이름>
timeouts:
request: 1800s

AIGatewayRoute를 생성하면 같은 이름의 HTTPRoute가 자동으로 생성됩니다. hostnames는 Gateway 리스너의 호스트명 패턴에 포함되어야 합니다.

트래픽 구성​

LLM 요청은 본문이 크고 응답 생성 시간이 길기 때문에 요청 버퍼와 백엔드 타임아웃을 워크로드 특성에 맞게 설정합니다.

백엔드 트래픽 정책 설정​

BackendTrafficPolicy는 AIGatewayRoute가 자동으로 생성한 HTTPRoute를 대상으로 합니다.

backend-traffic-policy.yaml
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: BackendTrafficPolicy
metadata:
name: <BackendTrafficPolicy 이름>
namespace: <사용자 namespace>
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: HTTPRoute
name: <AIGatewayRoute 이름>
timeout:
http:
requestTimeout: 1800s
streamIdleTimeout: 1800s
maxStreamDuration: 3600s
tcp:
connectTimeout: 30s
loadBalancer:
type: LeastRequest
circuitBreaker:
maxConnections: 4096
maxPendingRequests: 4096
maxParallelRequests: 4096
maxParallelRetries: 16
retry:
numRetries: 1
perRetry:
timeout: 1800s
backOff:
baseInterval: 500ms
maxInterval: 5s
retryOn:
triggers:
- connect-failure
- reset
- unavailable

Rate Limit 구성​

BackendTrafficPolicy의 rateLimit.global로 클라이언트별 요청 수 또는 토큰 사용량을 제한합니다. 제한을 초과한 요청은 429 Too Many Requests와 x-envoy-ratelimited: true 헤더로 거부됩니다.

요청 수 기반 제한​

request-rate-limit.yaml
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: BackendTrafficPolicy
metadata:
name: <BackendTrafficPolicy 이름>
namespace: <사용자 namespace>
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: HTTPRoute
name: <AIGatewayRoute 이름>
rateLimit:
type: Global
global:
rules:
- clientSelectors:
- headers:
- name: <클라이언트 식별 헤더>
type: Distinct
limit:
requests: <허용 요청 수>
unit: <단위 시간>

토큰 사용량 기반 제한​

AIGatewayRoute의 llmRequestCosts로 토큰 사용량을 메타데이터에 기록합니다.

spec:
llmRequestCosts:
- metadataKey: llm_input_token
type: InputToken
- metadataKey: llm_output_token
type: OutputToken
- metadataKey: llm_total_token
type: TotalToken

BackendTrafficPolicy에서 같은 메타데이터 키를 참조하여 토큰 예산을 차감합니다.

spec:
rateLimit:
type: Global
global:
rules:
- clientSelectors:
- headers:
- name: <클라이언트 식별 헤더>
type: Distinct
limit:
requests: <허용 토큰 수>
unit: <단위 시간>
cost:
request:
from: Number
number: 0
response:
from: Metadata
metadata:
namespace: io.envoy.ai_gateway
key: llm_total_token
토큰 차감 시점

토큰 사용량은 응답이 완료된 후 차감됩니다. 스트리밍 요청은 스트림이 끝난 후 차감되며, 진행 중인 스트림을 중간에 종료하지 않습니다.

상태 확인​

kubectl get gateway,httproute -n <사용자 namespace>
kubectl get backend,aiservicebackend,aigatewayroute -n <사용자 namespace>
kubectl get securitypolicy,clienttrafficpolicy,backendtrafficpolicy -n <사용자 namespace>

HTTPRoute의 Accepted와 ResolvedRefs 조건이 True인지 확인합니다.

kubectl describe httproute <AIGatewayRoute 이름> -n <사용자 namespace>