Serve a model with NVIDIA Dynamo
Use NVIDIA Dynamo's DynamoGraphDeployment to deploy GLM-5.3 on a vLLM backend and connect its OpenAI-compatible API to the Gateway in your Namespace.
Prerequisites
- See Before you begin to verify the Namespace and common resources.
- See Use Queues and Schedulers to prepare the
servingLocalQueue. - Prepare the model files in the
GLM-5.3directory of themodelsPVC. - To apply authentication or traffic policies to the external endpoint, first see Configure a Gateway.
Components
| Configuration | Resource | Function |
|---|---|---|
| Serving | DynamoGraphDeployment | Deploys the Frontend and vLLM Workers and provides an OpenAI-compatible API |
| Gateway connection | Backend, AIServiceBackend, AIGatewayRoute | Routes the Dynamo Frontend Service based on the hostname and model name |
| Traffic | BackendTrafficPolicy | Configures timeouts, load balancing, circuit breaking, and retries |
Step 1. Create a DynamoGraphDeployment
A DynamoGraphDeployment consists of a Frontend that receives requests and Workers that run the model. The Worker replicas value is the number of model instances.
The main settings are as follows.
| Field | Description |
|---|---|
spec.backendFramework | Inference backend. This example uses vllm. |
spec.components[].type | Component role: frontend or worker |
spec.components[].runtimeVersionOverride | Dynamo runtime version. Specify the same value as the Worker image tag. |
podTemplate.spec.schedulerName | Scheduler used for Workers. This example uses kai-scheduler. |
--served-model-name | model value in the API request body and the model-matching value used by the Gateway |
--tensor-parallel-size | Number of tensor-parallel GPUs. Specify the same value as the nvidia.com/gpu request. |
--gpu-memory-utilization | Percentage of GPU memory used by vLLM |
Save the following manifest as dynamo-graph-deployment.yaml.
apiVersion: nvidia.com/v1beta1
kind: DynamoGraphDeployment
metadata:
name: glm-5-3
namespace: <user-namespace>
spec:
backendFramework: vllm
components:
- name: Frontend
type: frontend
replicas: 1
podTemplate:
metadata:
labels:
kai.scheduler/enabled: "true"
kueue.x-k8s.io/queue-name: serving
kueue.x-k8s.io/priority-class: high
annotations:
sidecar.istio.io/inject: "false"
spec:
containers:
- name: main
image: nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.4.2
resources:
requests:
cpu: "2"
memory: 4Gi
limits:
cpu: "4"
memory: 8Gi
- name: VllmWorker
type: worker
replicas: 2
runtimeVersionOverride: "1.4.2"
podTemplate:
metadata:
labels:
kai.scheduler/enabled: "true"
kueue.x-k8s.io/queue-name: serving
kueue.x-k8s.io/priority-class: high
annotations:
sidecar.istio.io/inject: "false"
spec:
schedulerName: kai-scheduler
containers:
- name: main
image: nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.4.2
command: ["sh", "-c"]
args:
- |
exec python3 -m dynamo.vllm \
--model=/mnt/models/GLM-5.3 \
--served-model-name=zai-org/GLM-5.3 \
--dyn-tool-call-parser=glm47 \
--dyn-reasoning-parser=glm45 \
--tensor-parallel-size=8 \
--enable-expert-parallel \
--kv-cache-dtype=fp8_e4m3 \
--max-model-len=131072 \
--max-num-seqs=128 \
--max-num-batched-tokens=32768 \
'--speculative-config={"method":"mtp","num_speculative_tokens":5}' \
--gpu-memory-utilization=0.95 \
--safetensors-load-strategy=prefetch
env:
- name: VLLM_DEEP_GEMM_WARMUP
value: skip
- name: NCCL_SOCKET_IFNAME
value: eth0
- name: GLOO_SOCKET_IFNAME
value: eth0
resources:
limits:
nvidia.com/gpu: "8"
cpu: "200"
memory: 1900Gi
requests:
nvidia.com/gpu: "8"
cpu: "200"
memory: 1900Gi
volumeMounts:
- name: models
mountPath: /mnt
- name: devshm
mountPath: /dev/shm
volumes:
- name: models
persistentVolumeClaim:
claimName: models
- name: devshm
emptyDir:
medium: Memory
sizeLimit: 64Gi
The CPU, memory, and GPU requests for a Worker are configured per node. If your GPU node specification or quota differs from the example, adjust both requests and limits.
Apply the manifest.
kubectl apply -f dynamo-graph-deployment.yaml
After the resource is created, a Service named <DynamoGraphDeployment-name>-frontend is created and provides an OpenAI-compatible API on port 8000. In this example, the Service name is glm-5-3-frontend.
Step 2. Connect to the Gateway
To connect the Frontend Service to the Gateway, create a Backend, AIServiceBackend, and AIGatewayRoute. Configure timeouts and retries with BackendTrafficPolicy.
For field descriptions and example YAML, see Configure a Gateway - Configure AI Gateway and Configure a Gateway - Configure a backend traffic policy. Use the following values for a Dynamo deployment.
| Resource | Field | Value |
|---|---|---|
| Backend | spec.endpoints[].fqdn.hostname | glm-5-3-frontend.<user-namespace>.svc.cluster.local |
| Backend | spec.endpoints[].fqdn.port | 8000 |
| AIServiceBackend | spec.schema.name | OpenAI |
| AIGatewayRoute | spec.rules[].matches[].headers[].value | zai-org/GLM-5.3 |
| AIGatewayRoute | spec.rules[].timeouts.request | 1800s |
| BackendTrafficPolicy | spec.targetRefs[].name | AIGatewayRoute name |
The model-matching value used by the Gateway must match the Worker's --served-model-name value.
If you need authentication and rate limits, also apply Configure a Gateway - Configure API key authentication and Configure a Gateway - Configure rate limits.
Step 3. Check the deployment status
kubectl get dynamographdeployment glm-5-3 -n <user-namespace>
kubectl get pods,service -n <user-namespace> -o wide
kubectl get workload -n <user-namespace>
Verify that the Worker Pods have loaded the model and are in Running. If a Pod is in SchedulingGated, check the Queue and priority. If a Pod is in Pending, check the KAI Scheduler events and GPU availability.
kubectl describe pod <worker-pod-name> -n <user-namespace>
kubectl logs <worker-pod-name> -n <user-namespace> -c main
Step 4. Invoke the model
If you applied authentication, include the API key in the authorization header.
curl -s https://<hostname>/v1/chat/completions \
-H "authorization: Bearer <api-key>" \
-H "content-type: application/json" \
-d '{"model":"zai-org/GLM-5.3","messages":[{"role":"user","content":"Hello"}],"max_tokens":16}'
Clean up resources
Delete the DynamoGraphDeployment when you no longer need it.
kubectl delete dynamographdeployment glm-5-3 -n <user-namespace>
To also remove the Gateway resources, delete them in the following order: BackendTrafficPolicy, AIGatewayRoute, AIServiceBackend, and Backend.