Serve model with KServe LLMInferenceService
Use KServe LLMInferenceService to deploy GLM-5.3 with disaggregated Prefill and Decode. Prefill runs as a single-node TP8 Deployment, and Decode runs as a LeaderWorkerSet that spans two nodes.
Prerequisites
- See Before you begin to verify the Namespace,
modelsPVC, Gateway, and NetworkAttachmentDefinition. - See Use Queues and Schedulers to prepare the
servingLocalQueue. - Prepare the model weights at
<model-path>in themodelsPVC. - Prepare different values under
*.moduai.kakaocloud.comfor the service hostname and AIGatewayRoute hostname.
Architecture
A request passes through the Gateway, HTTPRoute, and InferencePool to a Decode Pod. The Decode Pod receives the KV cache generated by a Prefill Pod and generates tokens.
| Resource | Created by | Function |
|---|---|---|
| LLMInferenceService | User | Creates a Prefill Deployment, Decode LeaderWorkerSet, EPP, InferencePool, and HTTPRoute |
| AIGatewayRoute | User | Exposes the same InferencePool through a separate hostname and routes requests based on the model name |
| BackendTrafficPolicy | User | Configures timeouts, retries, and rate limits for both HTTPRoutes |
<service-name>-inference-pool | KServe | Groups Prefill Pods and the Decode Leader Pod and connects them to the EPP |
<service-name>-kserve-route | KServe | HTTPRoute that receives requests through the service hostname |
| HTTPRoute with the same name as AIGatewayRoute | AI Gateway | HTTPRoute that receives requests through the AIGatewayRoute hostname |
The EPP selects Prefill Pods with the Prefill profile and Decode Pods with the Decode profile. The KV cache is transferred from Prefill to Decode through UCX and InfiniBand.
Configuration
Common settings
| Field | Description |
|---|---|
metadata.labels | Kueue Queue and priority. Only labels with the kueue.x-k8s.io/ prefix are propagated to created Pods. |
metadata.annotations | InfiniBand NetworkAttachmentDefinition. Repeat the Namespace name for the number of requested NICs. |
spec.model.uri | Model location in the pvc://models/<model-path> format |
spec.model.name | Model name exposed by vLLM. It must match the model value in the request body. |
spec.annotations | Disables Istio sidecar injection for Decode Pods |
spec.prefill.annotations | Disables Istio sidecar injection for Prefill Pods |
spec.router.scheduler.annotations | Disables Istio sidecar injection for the EPP Pod |
Decode settings
- With
tensor: 8,data: 2, anddataLocal: 1, one Decode group consists of one Leader and one Worker, for a total of two nodes and 16 GPUs. - The routing sidecar of the Leader receives requests on port 8000 and forwards them to vLLM on port 8001.
- Set
kv_roletokv_consumer. - Both the Leader and Worker use
volcano-scheduler.
Prefill settings
- Prefill is configured as a single-node TP8 Deployment.
- Set
kv_roletokv_producer. - Use
--max-num-batched-tokensand--enforce-eageras Prefill-specific options.
-
If you specify only
hostnamesinspec.router.route.http.specand omitrules, the InferencePool is not connected to the Gateway. You must specify<service-name>-inference-poolinrules[].backendRefs. -
Do not add
initContainerstospec.worker. The routing sidecar is required only on the Leader. If you add the same entry to the Worker, Worker Pod creation fails, and the Leader also remains inPendingbecause the Gang cannot be formed.
Step 1. Create LLMInferenceService
Save the following manifest as llm-inference-service.yaml.
apiVersion: serving.kserve.io/v1alpha2
kind: LLMInferenceService
metadata:
name: <service-name>
namespace: <user-namespace>
labels:
kueue.x-k8s.io/queue-name: serving
kueue.x-k8s.io/priority-class: high
annotations:
k8s.v1.cni.cncf.io/networks: |-
<user-namespace>,
<user-namespace>,
<user-namespace>,
<user-namespace>,
<user-namespace>,
<user-namespace>,
<user-namespace>,
<user-namespace>
spec:
model:
uri: pvc://models/<model-path>
name: zai-org/GLM-5.3
annotations:
sidecar.istio.io/inject: "false"
router:
gateway:
refs:
- name: <user-namespace>
namespace: <user-namespace>
route:
http:
spec:
hostnames:
- <service-hostname>.moduai.kakaocloud.com
rules:
- matches:
- path:
type: PathPrefix
value: /
backendRefs:
- group: inference.networking.k8s.io
kind: InferencePool
name: <service-name>-inference-pool
port: 8000
scheduler:
labels:
kueue.x-k8s.io/queue-name: serving
kueue.x-k8s.io/priority-class: high
annotations:
sidecar.istio.io/inject: "false"
config:
inline:
apiVersion: llm-d.ai/v1alpha1
kind: EndpointPickerConfig
plugins:
- type: disagg-headers-handler
- type: always-disagg-pd-decider
- type: disagg-profile-handler
parameters:
deciders:
prefill: always-disagg-pd-decider
- type: prefill-filter
- type: decode-filter
- type: approx-prefix-cache-producer
parameters:
maxPrefixTokensToMatch: 131072
- type: inflight-load-producer
- type: prefix-cache-affinity-filter
parameters:
peakPrefillThroughput: 33821
- type: token-load-scorer
- type: active-request-scorer
- type: max-score-picker
schedulingProfiles:
- name: prefill
plugins:
- pluginRef: prefill-filter
- pluginRef: prefix-cache-affinity-filter
- pluginRef: token-load-scorer
- pluginRef: max-score-picker
- name: decode
plugins:
- pluginRef: decode-filter
- pluginRef: active-request-scorer
- pluginRef: max-score-picker
replicas: 1
parallelism:
tensor: 8
data: 2
dataLocal: 1
template:
schedulerName: volcano-scheduler
initContainers:
- name: llm-d-routing-sidecar
args:
- --port=8000
- --vllm-port=8001
- --kv-connector=nixlv2
- --secure-proxy=false
- --enable-ssrf-protection=false
containers:
- name: main
image: docker.io/vllm/vllm-openai:v0.28.0
args:
- --kv-cache-dtype=fp8_e4m3
- --max-num-seqs=32
- --no-enable-prefix-caching
- --speculative-config='{"method":"mtp","num_speculative_tokens":5}'
- --tool-call-parser=glm47
- --reasoning-parser=glm45
- --enable-auto-tool-choice
- --kv-transfer-config='{"kv_connector":"NixlConnector","kv_role":"kv_consumer","kv_buffer_device":"cuda","kv_connector_extra_config":{"backends":["UCX"]}}'
env:
- name: UCX_TLS
value: rc_mlx5,rc,cuda_ipc,cuda_copy,sm,shm,self
- name: UCX_NET_DEVICES
value: all
- name: NCCL_SOCKET_IFNAME
value: eth0
- name: GLOO_SOCKET_IFNAME
value: eth0
- name: VLLM_NIXL_SIDE_CHANNEL_HOST
valueFrom:
fieldRef:
fieldPath: status.podIP
- name: VLLM_NIXL_SIDE_CHANNEL_PORT
value: "5600"
- name: VLLM_HTTP_TIMEOUT_KEEP_ALIVE
value: "120"
securityContext:
runAsNonRoot: false
resources:
limits:
nvidia.com/gpu: "8"
nvidia.com/mlnxnics: "8"
cpu: "200"
memory: 1800Gi
requests:
nvidia.com/gpu: "8"
nvidia.com/mlnxnics: "8"
cpu: "200"
memory: 1800Gi
volumes:
- name: dshm
emptyDir:
medium: Memory
sizeLimit: 64Gi
worker:
schedulerName: volcano-scheduler
containers:
- name: main
image: docker.io/vllm/vllm-openai:v0.28.0
args:
- --kv-cache-dtype=fp8_e4m3
- --max-num-seqs=32
- --no-enable-prefix-caching
- --speculative-config='{"method":"mtp","num_speculative_tokens":5}'
- --tool-call-parser=glm47
- --reasoning-parser=glm45
- --enable-auto-tool-choice
- --kv-transfer-config='{"kv_connector":"NixlConnector","kv_role":"kv_consumer","kv_buffer_device":"cuda","kv_connector_extra_config":{"backends":["UCX"]}}'
env:
- name: UCX_TLS
value: rc_mlx5,rc,cuda_ipc,cuda_copy,sm,shm,self
- name: UCX_NET_DEVICES
value: all
- name: NCCL_SOCKET_IFNAME
value: eth0
- name: GLOO_SOCKET_IFNAME
value: eth0
- name: VLLM_NIXL_SIDE_CHANNEL_HOST
valueFrom:
fieldRef:
fieldPath: status.podIP
- name: VLLM_NIXL_SIDE_CHANNEL_PORT
value: "5600"
- name: VLLM_HTTP_TIMEOUT_KEEP_ALIVE
value: "120"
securityContext:
runAsNonRoot: false
resources:
limits:
nvidia.com/gpu: "8"
nvidia.com/mlnxnics: "8"
cpu: "200"
memory: 1800Gi
requests:
nvidia.com/gpu: "8"
nvidia.com/mlnxnics: "8"
cpu: "200"
memory: 1800Gi
volumes:
- name: dshm
emptyDir:
medium: Memory
sizeLimit: 64Gi
prefill:
replicas: 1
parallelism:
tensor: 8
annotations:
sidecar.istio.io/inject: "false"
template:
schedulerName: volcano-scheduler
containers:
- name: main
image: docker.io/vllm/vllm-openai:v0.28.0
args:
- --kv-cache-dtype=fp8_e4m3
- --max-num-batched-tokens=16384
- --no-enable-prefix-caching
- --enforce-eager
- --speculative-config='{"method":"mtp","num_speculative_tokens":5}'
- --tool-call-parser=glm47
- --reasoning-parser=glm45
- --enable-auto-tool-choice
- --kv-transfer-config='{"kv_connector":"NixlConnector","kv_role":"kv_producer","kv_buffer_device":"cuda","kv_connector_extra_config":{"backends":["UCX"]}}'
env:
- name: UCX_TLS
value: rc_mlx5,rc,cuda_ipc,cuda_copy,sm,shm,self
- name: UCX_NET_DEVICES
value: all
- name: NCCL_SOCKET_IFNAME
value: eth0
- name: GLOO_SOCKET_IFNAME
value: eth0
- name: VLLM_NIXL_SIDE_CHANNEL_HOST
valueFrom:
fieldRef:
fieldPath: status.podIP
- name: VLLM_NIXL_SIDE_CHANNEL_PORT
value: "5600"
- name: VLLM_HTTP_TIMEOUT_KEEP_ALIVE
value: "120"
securityContext:
runAsNonRoot: false
resources:
limits:
nvidia.com/gpu: "8"
nvidia.com/mlnxnics: "8"
cpu: "200"
memory: 1800Gi
requests:
nvidia.com/gpu: "8"
nvidia.com/mlnxnics: "8"
cpu: "200"
memory: 1800Gi
volumes:
- name: dshm
emptyDir:
medium: Memory
sizeLimit: 64Gi
Apply the manifest.
kubectl apply -f llm-inference-service.yaml
If you change metadata.labels or spec.prefill.labels after deployment, the selector of the Prefill Deployment might change and cause a field is immutable error. In this case, delete the Prefill Deployment, and the KServe controller creates it again.
Step 2. Connect to Gateway
To expose the InferencePool created by LLMInferenceService through a separate hostname, see Configure a Gateway - Configure AI Gateway.
Step 3. Check deployment status
kubectl get llmisvc <service-name> -n <user-namespace>
kubectl get lws,deploy,pods -l app.kubernetes.io/name=<service-name> -n <user-namespace> -o wide
kubectl get podgroup,workload -n <user-namespace>
kubectl get inferencepool,httproute -n <user-namespace>
Verify the following conditions:
- The Prefill Deployment is in
1/1. - The Decode Leader and Worker Pods are in
Runningon different nodes. - The EPP Pod is in
Running. - The
ADMITTEDcondition of the Kueue Workload isTrue. - All status conditions of LLMInferenceService are
True.
kubectl get llmisvc <service-name> -n <user-namespace> \
-o jsonpath='{range .status.conditions[*]}{.type}{"\t"}{.status}{"\t"}{.reason}{"\n"}{end}'
Troubleshooting
| Symptom | What to check |
|---|---|
Pod is in SchedulingGated | Kueue has not admitted the Workload. Check the Queue label location and requested resources. |
Workload is admitted, but a Decode Pod is in Pending | Verify that two GPU nodes required for Gang scheduling are available, and check the Pod events. |
InferencePoolReady is WaitingForGateway | Verify that rules[].backendRefs in the HTTPRoute references the InferencePool. |
Worker Pod is not created, and the Leader is also in Pending | Check whether unnecessary initContainers were added to spec.worker. |
For detailed commands, see Troubleshooting.
Clean up resources
kubectl delete aigatewayroute <ai-gateway-route-name> -n <user-namespace>
kubectl delete llminferenceservice <service-name> -n <user-namespace>
Verify that the automatically created Prefill Deployment, Decode LeaderWorkerSet, EPP, InferencePool, and HTTPRoute are also deleted.