Skip to main content

Serve model with KServe LLMInferenceService

Use KServe LLMInferenceService to deploy GLM-5.3 with disaggregated Prefill and Decode. Prefill runs as a single-node TP8 Deployment, and Decode runs as a LeaderWorkerSet that spans two nodes.

Prerequisites​

  • See Before you begin to verify the Namespace, models PVC, Gateway, and NetworkAttachmentDefinition.
  • See Use Queues and Schedulers to prepare the serving LocalQueue.
  • Prepare the model weights at <model-path> in the models PVC.
  • Prepare different values under *.moduai.kakaocloud.com for the service hostname and AIGatewayRoute hostname.

Architecture​

A request passes through the Gateway, HTTPRoute, and InferencePool to a Decode Pod. The Decode Pod receives the KV cache generated by a Prefill Pod and generates tokens.

ResourceCreated byFunction
LLMInferenceServiceUserCreates a Prefill Deployment, Decode LeaderWorkerSet, EPP, InferencePool, and HTTPRoute
AIGatewayRouteUserExposes the same InferencePool through a separate hostname and routes requests based on the model name
BackendTrafficPolicyUserConfigures timeouts, retries, and rate limits for both HTTPRoutes
<service-name>-inference-poolKServeGroups Prefill Pods and the Decode Leader Pod and connects them to the EPP
<service-name>-kserve-routeKServeHTTPRoute that receives requests through the service hostname
HTTPRoute with the same name as AIGatewayRouteAI GatewayHTTPRoute that receives requests through the AIGatewayRoute hostname

The EPP selects Prefill Pods with the Prefill profile and Decode Pods with the Decode profile. The KV cache is transferred from Prefill to Decode through UCX and InfiniBand.

Configuration​

Common settings​

FieldDescription
metadata.labelsKueue Queue and priority. Only labels with the kueue.x-k8s.io/ prefix are propagated to created Pods.
metadata.annotationsInfiniBand NetworkAttachmentDefinition. Repeat the Namespace name for the number of requested NICs.
spec.model.uriModel location in the pvc://models/<model-path> format
spec.model.nameModel name exposed by vLLM. It must match the model value in the request body.
spec.annotationsDisables Istio sidecar injection for Decode Pods
spec.prefill.annotationsDisables Istio sidecar injection for Prefill Pods
spec.router.scheduler.annotationsDisables Istio sidecar injection for the EPP Pod

Decode settings​

  • With tensor: 8, data: 2, and dataLocal: 1, one Decode group consists of one Leader and one Worker, for a total of two nodes and 16 GPUs.
  • The routing sidecar of the Leader receives requests on port 8000 and forwards them to vLLM on port 8001.
  • Set kv_role to kv_consumer.
  • Both the Leader and Worker use volcano-scheduler.

Prefill settings​

  • Prefill is configured as a single-node TP8 Deployment.
  • Set kv_role to kv_producer.
  • Use --max-num-batched-tokens and --enforce-eager as Prefill-specific options.
Caution
  • If you specify only hostnames in spec.router.route.http.spec and omit rules, the InferencePool is not connected to the Gateway. You must specify <service-name>-inference-pool in rules[].backendRefs.

  • Do not add initContainers to spec.worker. The routing sidecar is required only on the Leader. If you add the same entry to the Worker, Worker Pod creation fails, and the Leader also remains in Pending because the Gang cannot be formed.

Step 1. Create LLMInferenceService​

Save the following manifest as llm-inference-service.yaml.

llm-inference-service.yaml
apiVersion: serving.kserve.io/v1alpha2
kind: LLMInferenceService
metadata:
name: <service-name>
namespace: <user-namespace>
labels:
kueue.x-k8s.io/queue-name: serving
kueue.x-k8s.io/priority-class: high
annotations:
k8s.v1.cni.cncf.io/networks: |-
<user-namespace>,
<user-namespace>,
<user-namespace>,
<user-namespace>,
<user-namespace>,
<user-namespace>,
<user-namespace>,
<user-namespace>
spec:
model:
uri: pvc://models/<model-path>
name: zai-org/GLM-5.3
annotations:
sidecar.istio.io/inject: "false"

router:
gateway:
refs:
- name: <user-namespace>
namespace: <user-namespace>
route:
http:
spec:
hostnames:
- <service-hostname>.moduai.kakaocloud.com
rules:
- matches:
- path:
type: PathPrefix
value: /
backendRefs:
- group: inference.networking.k8s.io
kind: InferencePool
name: <service-name>-inference-pool
port: 8000
scheduler:
labels:
kueue.x-k8s.io/queue-name: serving
kueue.x-k8s.io/priority-class: high
annotations:
sidecar.istio.io/inject: "false"
config:
inline:
apiVersion: llm-d.ai/v1alpha1
kind: EndpointPickerConfig
plugins:
- type: disagg-headers-handler
- type: always-disagg-pd-decider
- type: disagg-profile-handler
parameters:
deciders:
prefill: always-disagg-pd-decider
- type: prefill-filter
- type: decode-filter
- type: approx-prefix-cache-producer
parameters:
maxPrefixTokensToMatch: 131072
- type: inflight-load-producer
- type: prefix-cache-affinity-filter
parameters:
peakPrefillThroughput: 33821
- type: token-load-scorer
- type: active-request-scorer
- type: max-score-picker
schedulingProfiles:
- name: prefill
plugins:
- pluginRef: prefill-filter
- pluginRef: prefix-cache-affinity-filter
- pluginRef: token-load-scorer
- pluginRef: max-score-picker
- name: decode
plugins:
- pluginRef: decode-filter
- pluginRef: active-request-scorer
- pluginRef: max-score-picker

replicas: 1
parallelism:
tensor: 8
data: 2
dataLocal: 1
template:
schedulerName: volcano-scheduler
initContainers:
- name: llm-d-routing-sidecar
args:
- --port=8000
- --vllm-port=8001
- --kv-connector=nixlv2
- --secure-proxy=false
- --enable-ssrf-protection=false
containers:
- name: main
image: docker.io/vllm/vllm-openai:v0.28.0
args:
- --kv-cache-dtype=fp8_e4m3
- --max-num-seqs=32
- --no-enable-prefix-caching
- --speculative-config='{"method":"mtp","num_speculative_tokens":5}'
- --tool-call-parser=glm47
- --reasoning-parser=glm45
- --enable-auto-tool-choice
- --kv-transfer-config='{"kv_connector":"NixlConnector","kv_role":"kv_consumer","kv_buffer_device":"cuda","kv_connector_extra_config":{"backends":["UCX"]}}'
env:
- name: UCX_TLS
value: rc_mlx5,rc,cuda_ipc,cuda_copy,sm,shm,self
- name: UCX_NET_DEVICES
value: all
- name: NCCL_SOCKET_IFNAME
value: eth0
- name: GLOO_SOCKET_IFNAME
value: eth0
- name: VLLM_NIXL_SIDE_CHANNEL_HOST
valueFrom:
fieldRef:
fieldPath: status.podIP
- name: VLLM_NIXL_SIDE_CHANNEL_PORT
value: "5600"
- name: VLLM_HTTP_TIMEOUT_KEEP_ALIVE
value: "120"
securityContext:
runAsNonRoot: false
resources:
limits:
nvidia.com/gpu: "8"
nvidia.com/mlnxnics: "8"
cpu: "200"
memory: 1800Gi
requests:
nvidia.com/gpu: "8"
nvidia.com/mlnxnics: "8"
cpu: "200"
memory: 1800Gi
volumes:
- name: dshm
emptyDir:
medium: Memory
sizeLimit: 64Gi

worker:
schedulerName: volcano-scheduler
containers:
- name: main
image: docker.io/vllm/vllm-openai:v0.28.0
args:
- --kv-cache-dtype=fp8_e4m3
- --max-num-seqs=32
- --no-enable-prefix-caching
- --speculative-config='{"method":"mtp","num_speculative_tokens":5}'
- --tool-call-parser=glm47
- --reasoning-parser=glm45
- --enable-auto-tool-choice
- --kv-transfer-config='{"kv_connector":"NixlConnector","kv_role":"kv_consumer","kv_buffer_device":"cuda","kv_connector_extra_config":{"backends":["UCX"]}}'
env:
- name: UCX_TLS
value: rc_mlx5,rc,cuda_ipc,cuda_copy,sm,shm,self
- name: UCX_NET_DEVICES
value: all
- name: NCCL_SOCKET_IFNAME
value: eth0
- name: GLOO_SOCKET_IFNAME
value: eth0
- name: VLLM_NIXL_SIDE_CHANNEL_HOST
valueFrom:
fieldRef:
fieldPath: status.podIP
- name: VLLM_NIXL_SIDE_CHANNEL_PORT
value: "5600"
- name: VLLM_HTTP_TIMEOUT_KEEP_ALIVE
value: "120"
securityContext:
runAsNonRoot: false
resources:
limits:
nvidia.com/gpu: "8"
nvidia.com/mlnxnics: "8"
cpu: "200"
memory: 1800Gi
requests:
nvidia.com/gpu: "8"
nvidia.com/mlnxnics: "8"
cpu: "200"
memory: 1800Gi
volumes:
- name: dshm
emptyDir:
medium: Memory
sizeLimit: 64Gi

prefill:
replicas: 1
parallelism:
tensor: 8
annotations:
sidecar.istio.io/inject: "false"
template:
schedulerName: volcano-scheduler
containers:
- name: main
image: docker.io/vllm/vllm-openai:v0.28.0
args:
- --kv-cache-dtype=fp8_e4m3
- --max-num-batched-tokens=16384
- --no-enable-prefix-caching
- --enforce-eager
- --speculative-config='{"method":"mtp","num_speculative_tokens":5}'
- --tool-call-parser=glm47
- --reasoning-parser=glm45
- --enable-auto-tool-choice
- --kv-transfer-config='{"kv_connector":"NixlConnector","kv_role":"kv_producer","kv_buffer_device":"cuda","kv_connector_extra_config":{"backends":["UCX"]}}'
env:
- name: UCX_TLS
value: rc_mlx5,rc,cuda_ipc,cuda_copy,sm,shm,self
- name: UCX_NET_DEVICES
value: all
- name: NCCL_SOCKET_IFNAME
value: eth0
- name: GLOO_SOCKET_IFNAME
value: eth0
- name: VLLM_NIXL_SIDE_CHANNEL_HOST
valueFrom:
fieldRef:
fieldPath: status.podIP
- name: VLLM_NIXL_SIDE_CHANNEL_PORT
value: "5600"
- name: VLLM_HTTP_TIMEOUT_KEEP_ALIVE
value: "120"
securityContext:
runAsNonRoot: false
resources:
limits:
nvidia.com/gpu: "8"
nvidia.com/mlnxnics: "8"
cpu: "200"
memory: 1800Gi
requests:
nvidia.com/gpu: "8"
nvidia.com/mlnxnics: "8"
cpu: "200"
memory: 1800Gi
volumes:
- name: dshm
emptyDir:
medium: Memory
sizeLimit: 64Gi

Apply the manifest.

kubectl apply -f llm-inference-service.yaml
Caution

If you change metadata.labels or spec.prefill.labels after deployment, the selector of the Prefill Deployment might change and cause a field is immutable error. In this case, delete the Prefill Deployment, and the KServe controller creates it again.

Step 2. Connect to Gateway​

To expose the InferencePool created by LLMInferenceService through a separate hostname, see Configure a Gateway - Configure AI Gateway.

Step 3. Check deployment status​

kubectl get llmisvc <service-name> -n <user-namespace>
kubectl get lws,deploy,pods -l app.kubernetes.io/name=<service-name> -n <user-namespace> -o wide
kubectl get podgroup,workload -n <user-namespace>
kubectl get inferencepool,httproute -n <user-namespace>

Verify the following conditions:

  • The Prefill Deployment is in 1/1.
  • The Decode Leader and Worker Pods are in Running on different nodes.
  • The EPP Pod is in Running.
  • The ADMITTED condition of the Kueue Workload is True.
  • All status conditions of LLMInferenceService are True.
kubectl get llmisvc <service-name> -n <user-namespace> \
-o jsonpath='{range .status.conditions[*]}{.type}{"\t"}{.status}{"\t"}{.reason}{"\n"}{end}'

Troubleshooting​

SymptomWhat to check
Pod is in SchedulingGatedKueue has not admitted the Workload. Check the Queue label location and requested resources.
Workload is admitted, but a Decode Pod is in PendingVerify that two GPU nodes required for Gang scheduling are available, and check the Pod events.
InferencePoolReady is WaitingForGatewayVerify that rules[].backendRefs in the HTTPRoute references the InferencePool.
Worker Pod is not created, and the Leader is also in PendingCheck whether unnecessary initContainers were added to spec.worker.

For detailed commands, see Troubleshooting.

Clean up resources​

kubectl delete aigatewayroute <ai-gateway-route-name> -n <user-namespace>
kubectl delete llminferenceservice <service-name> -n <user-namespace>

Verify that the automatically created Prefill Deployment, Decode LeaderWorkerSet, EPP, InferencePool, and HTTPRoute are also deleted.