Before you begin
Before deploying GPU workloads, verify common resources such as the Namespace, Queue, model files, and Gateway.
Requirements
| Item | Description |
|---|---|
| Kubernetes Namespace | Dedicated user Namespace where GPU workloads and serving resources are created |
kubectl | Kubernetes CLI configured to access your assigned cluster and Namespace |
| LocalQueue | Namespace-level Queue that Kueue uses to manage GPU resource requests |
| Model PVC | PVC that stores model weights. This guide uses models as an example. |
| Gateway | Gateway created in advance with the same name as the Namespace |
| NetworkAttachmentDefinition | Network definition attached to workloads that use InfiniBand. It uses the same name as the Namespace. |
| External hostname | Hostname under the *.moduai.kakaocloud.com domain used for model serving |
Check resources
-
Check the current context and default Namespace.
kubectl config current-context
kubectl config view --minify --output 'jsonpath={..namespace}' -
Check the common resources provided in the Namespace.
kubectl get localqueue
kubectl get pvc models
kubectl get gateway
kubectl get network-attachment-definitions -
Verify that the model files are available in the model PVC. The model path must match the
modelorurisetting of the serving runtime that you use.
Placeholders in examples
Replace values in angle brackets (<...>) with values for your environment.
| Placeholder | Description | Example |
|---|---|---|
<user-namespace> | Assigned Kubernetes Namespace | gpu-team-a |
<queue-name> | Name of a LocalQueue created in the Namespace | serving |
<service-name> | Name of the model serving resource | glm-5-3 |
<model-path> | Model directory in the models PVC | GLM-5.3 |
<hostname> | External hostname used by the Gateway | glm-5-3.moduai.kakaocloud.com |
<api-key> | API key used for Gateway authentication | A value issued for your environment |
Apply YAML
Save the example YAML to a file, and then apply it.
kubectl apply -f <filename>.yaml
To manage multiple files together, create a kustomization.yaml file in the same directory, and then apply it.
kubectl apply -k <directory>
The CPU, memory, GPU, and InfiniBand NIC counts in the examples are based on GLM-5.3. When using a different model or node specification, adjust the values to match the model requirements and your quota.