Skip to main content

Serve a model with NVIDIA Dynamo

Use NVIDIA Dynamo's DynamoGraphDeployment to deploy GLM-5.3 on a vLLM backend and connect its OpenAI-compatible API to the Gateway in your Namespace.

Prerequisites​

  • See Before you begin to verify the Namespace and common resources.
  • See Use Queues and Schedulers to prepare the serving LocalQueue.
  • Prepare the model files in the GLM-5.3 directory of the models PVC.
  • To apply authentication or traffic policies to the external endpoint, first see Configure a Gateway.

Components​

ConfigurationResourceFunction
ServingDynamoGraphDeploymentDeploys the Frontend and vLLM Workers and provides an OpenAI-compatible API
Gateway connectionBackend, AIServiceBackend, AIGatewayRouteRoutes the Dynamo Frontend Service based on the hostname and model name
TrafficBackendTrafficPolicyConfigures timeouts, load balancing, circuit breaking, and retries

Step 1. Create a DynamoGraphDeployment​

A DynamoGraphDeployment consists of a Frontend that receives requests and Workers that run the model. The Worker replicas value is the number of model instances.

The main settings are as follows.

FieldDescription
spec.backendFrameworkInference backend. This example uses vllm.
spec.components[].typeComponent role: frontend or worker
spec.components[].runtimeVersionOverrideDynamo runtime version. Specify the same value as the Worker image tag.
podTemplate.spec.schedulerNameScheduler used for Workers. This example uses kai-scheduler.
--served-model-namemodel value in the API request body and the model-matching value used by the Gateway
--tensor-parallel-sizeNumber of tensor-parallel GPUs. Specify the same value as the nvidia.com/gpu request.
--gpu-memory-utilizationPercentage of GPU memory used by vLLM

Save the following manifest as dynamo-graph-deployment.yaml.

dynamo-graph-deployment.yaml
apiVersion: nvidia.com/v1beta1
kind: DynamoGraphDeployment
metadata:
name: glm-5-3
namespace: <user-namespace>
spec:
backendFramework: vllm
components:
- name: Frontend
type: frontend
replicas: 1
podTemplate:
metadata:
labels:
kai.scheduler/enabled: "true"
kueue.x-k8s.io/queue-name: serving
kueue.x-k8s.io/priority-class: high
annotations:
sidecar.istio.io/inject: "false"
spec:
containers:
- name: main
image: nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.4.2
resources:
requests:
cpu: "2"
memory: 4Gi
limits:
cpu: "4"
memory: 8Gi

- name: VllmWorker
type: worker
replicas: 2
runtimeVersionOverride: "1.4.2"
podTemplate:
metadata:
labels:
kai.scheduler/enabled: "true"
kueue.x-k8s.io/queue-name: serving
kueue.x-k8s.io/priority-class: high
annotations:
sidecar.istio.io/inject: "false"
spec:
schedulerName: kai-scheduler
containers:
- name: main
image: nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.4.2
command: ["sh", "-c"]
args:
- |
exec python3 -m dynamo.vllm \
--model=/mnt/models/GLM-5.3 \
--served-model-name=zai-org/GLM-5.3 \
--dyn-tool-call-parser=glm47 \
--dyn-reasoning-parser=glm45 \
--tensor-parallel-size=8 \
--enable-expert-parallel \
--kv-cache-dtype=fp8_e4m3 \
--max-model-len=131072 \
--max-num-seqs=128 \
--max-num-batched-tokens=32768 \
'--speculative-config={"method":"mtp","num_speculative_tokens":5}' \
--gpu-memory-utilization=0.95 \
--safetensors-load-strategy=prefetch
env:
- name: VLLM_DEEP_GEMM_WARMUP
value: skip
- name: NCCL_SOCKET_IFNAME
value: eth0
- name: GLOO_SOCKET_IFNAME
value: eth0
resources:
limits:
nvidia.com/gpu: "8"
cpu: "200"
memory: 1900Gi
requests:
nvidia.com/gpu: "8"
cpu: "200"
memory: 1900Gi
volumeMounts:
- name: models
mountPath: /mnt
- name: devshm
mountPath: /dev/shm
volumes:
- name: models
persistentVolumeClaim:
claimName: models
- name: devshm
emptyDir:
medium: Memory
sizeLimit: 64Gi
Caution

The CPU, memory, and GPU requests for a Worker are configured per node. If your GPU node specification or quota differs from the example, adjust both requests and limits.

Apply the manifest.

kubectl apply -f dynamo-graph-deployment.yaml

After the resource is created, a Service named <DynamoGraphDeployment-name>-frontend is created and provides an OpenAI-compatible API on port 8000. In this example, the Service name is glm-5-3-frontend.

Step 2. Connect to the Gateway​

To connect the Frontend Service to the Gateway, create a Backend, AIServiceBackend, and AIGatewayRoute. Configure timeouts and retries with BackendTrafficPolicy.

For field descriptions and example YAML, see Configure a Gateway - Configure AI Gateway and Configure a Gateway - Configure a backend traffic policy. Use the following values for a Dynamo deployment.

ResourceFieldValue
Backendspec.endpoints[].fqdn.hostnameglm-5-3-frontend.<user-namespace>.svc.cluster.local
Backendspec.endpoints[].fqdn.port8000
AIServiceBackendspec.schema.nameOpenAI
AIGatewayRoutespec.rules[].matches[].headers[].valuezai-org/GLM-5.3
AIGatewayRoutespec.rules[].timeouts.request1800s
BackendTrafficPolicyspec.targetRefs[].nameAIGatewayRoute name

The model-matching value used by the Gateway must match the Worker's --served-model-name value.

If you need authentication and rate limits, also apply Configure a Gateway - Configure API key authentication and Configure a Gateway - Configure rate limits.

Step 3. Check the deployment status​

kubectl get dynamographdeployment glm-5-3 -n <user-namespace>
kubectl get pods,service -n <user-namespace> -o wide
kubectl get workload -n <user-namespace>

Verify that the Worker Pods have loaded the model and are in Running. If a Pod is in SchedulingGated, check the Queue and priority. If a Pod is in Pending, check the KAI Scheduler events and GPU availability.

kubectl describe pod <worker-pod-name> -n <user-namespace>
kubectl logs <worker-pod-name> -n <user-namespace> -c main

Step 4. Invoke the model​

If you applied authentication, include the API key in the authorization header.

curl -s https://<hostname>/v1/chat/completions \
-H "authorization: Bearer <api-key>" \
-H "content-type: application/json" \
-d '{"model":"zai-org/GLM-5.3","messages":[{"role":"user","content":"Hello"}],"max_tokens":16}'

Clean up resources​

Delete the DynamoGraphDeployment when you no longer need it.

kubectl delete dynamographdeployment glm-5-3 -n <user-namespace>

To also remove the Gateway resources, delete them in the following order: BackendTrafficPolicy, AIGatewayRoute, AIServiceBackend, and Backend.