Use Queues and Schedulers
Use Kueue to manage queueing and priority for GPU resource requests, and use KAI Scheduler or Volcano to schedule GPU workloads.
Supported components
| Component | Role |
|---|---|
| Kueue | Manages resource queues, allocation, and preemption |
| KAI Scheduler | Schedules workloads based on Gang Scheduling and Network Topology |
| Volcano | Schedules workloads based on Gang Scheduling and Network Topology |
To use a Queue, submit a request through your designated customer support channel in advance. After Queue usage is enabled, all resource creation and deletion is controlled through the Queue.
Create a LocalQueue
Create a LocalQueue for each purpose in your Namespace. In spec.clusterQueue, specify the ClusterQueue that has the same name as your assigned Namespace.
apiVersion: kueue.x-k8s.io/v1beta2
kind: LocalQueue
metadata:
name: <queue-name>
namespace: <user-namespace>
spec:
clusterQueue: <user-namespace>
fairSharing:
weight: "2"
kubectl apply -f localqueue.yaml
kubectl get localqueue -n <user-namespace>
fairSharing.weight is the relative weight applied when Queues compete for resources. A higher value receives more resources under otherwise equal conditions.
Assign a Queue to a workload
Add kueue.x-k8s.io/queue-name to the workload's metadata.labels.
metadata:
labels:
kueue.x-k8s.io/queue-name: serving
Specify the Queue name in metadata.labels, not in an annotation. If you specify a LocalQueue that does not exist, the Workload cannot receive resources.
Set priority
Specify the WorkloadPriorityClass appropriate for the purpose of the workload.
| Priority name | Weight |
|---|---|
low | 100 |
normal | 500 |
high | 1000 |
critical | 1500 |
metadata:
labels:
kueue.x-k8s.io/priority-class: normal
When a high-priority workload is admitted, an existing lower-priority workload might be preempted. Use high and critical only when operationally necessary.
Use KAI Scheduler
For workloads that require KAI Scheduler integration, such as NVIDIA Dynamo, specify both the enablement label and schedulerName.
metadata:
labels:
kai.scheduler/enabled: "true"
spec:
schedulerName: kai-scheduler
By default, Network Topology-based scheduling prioritizes GPU nodes connected to the same Spine Switch.
Use Volcano
For workloads that require Gang Scheduling, set schedulerName to volcano-scheduler.
spec:
schedulerName: volcano-scheduler
You can set Network Topology constraints to soft or hard mode.
spec:
podGroupPolicy:
volcano:
networkTopology:
mode: hard
highestTierAllowed: 1
| Mode | Behavior |
|---|---|
soft | Schedules Pods within a nearby network topology when possible, but can schedule them on other nodes if the condition cannot be met |
hard | Schedules Pods only when the specified topology condition is met |
In the current Spine topology, highestTierAllowed: 1 prevents scheduling when Pods cannot be placed within the same Spine. Specify highestTierAllowed only in hard mode.
Check status
Check the LocalQueue and Kueue Workload status.
kubectl get localqueue -n <user-namespace>
kubectl get workload -n <user-namespace>
kubectl describe workload <workload-name> -n <user-namespace>
Check the PodGroup and Pod events created by Volcano.
kubectl get podgroup -n <user-namespace>
kubectl get pod -n <user-namespace> -o wide
kubectl describe pod <pod-name> -n <user-namespace>
If a Workload remains in SchedulingGated, check the Queue name, priority, requested resources, and available resources in the Queue. If Kueue has admitted the Workload but a Pod remains in Pending, check the Scheduler events and resource availability across all nodes in the Gang.