Skip to main content

GPU workload overview

In the KakaoCloud GPU environment, you can run GPU workloads in your assigned Kubernetes Namespace, schedule resources by using Queues and Schedulers, and serve large language models with NVIDIA Dynamo or KServe.

Main components​

ComponentRole
KueueQueues GPU resource requests and runs or preempts workloads based on available resources and priority
KAI SchedulerSchedules GPU workloads based on Gang Scheduling and Network Topology
VolcanoSchedules workloads based on Gang Scheduling and Network Topology
GatewayExposes model-serving endpoints in a Namespace through external hostnames and applies authentication and traffic policies
NVIDIA DynamoBuilds a distributed inference graph and provides an OpenAI-compatible API
KServeUses LLMInferenceService to build an inference environment with disaggregated Prefill and Decode

Workflow​

  1. See Before you begin to verify common resources such as the Namespace, PVC, and Gateway.
  2. See Use Queues and Schedulers to configure queueing and scheduling for GPU workloads.
  3. See Configure a Gateway to prepare authentication and external routing policies.
  4. Depending on your workload, deploy the model by using NVIDIA Dynamo or KServe LLMInferenceService.
Note

To use a Queue, submit a request through your designated customer support channel in advance.