GPU workload overview
In the KakaoCloud GPU environment, you can run GPU workloads in your assigned Kubernetes Namespace, schedule resources by using Queues and Schedulers, and serve large language models with NVIDIA Dynamo or KServe.
Main components
| Component | Role |
|---|---|
| Kueue | Queues GPU resource requests and runs or preempts workloads based on available resources and priority |
| KAI Scheduler | Schedules GPU workloads based on Gang Scheduling and Network Topology |
| Volcano | Schedules workloads based on Gang Scheduling and Network Topology |
| Gateway | Exposes model-serving endpoints in a Namespace through external hostnames and applies authentication and traffic policies |
| NVIDIA Dynamo | Builds a distributed inference graph and provides an OpenAI-compatible API |
| KServe | Uses LLMInferenceService to build an inference environment with disaggregated Prefill and Decode |
Workflow
- See Before you begin to verify common resources such as the Namespace, PVC, and Gateway.
- See Use Queues and Schedulers to configure queueing and scheduling for GPU workloads.
- See Configure a Gateway to prepare authentication and external routing policies.
- Depending on your workload, deploy the model by using NVIDIA Dynamo or KServe LLMInferenceService.
Note
To use a Queue, submit a request through your designated customer support channel in advance.