Skip to main content

Before you begin

Before deploying GPU workloads, verify common resources such as the Namespace, Queue, model files, and Gateway.

Requirements​

ItemDescription
Kubernetes NamespaceDedicated user Namespace where GPU workloads and serving resources are created
kubectlKubernetes CLI configured to access your assigned cluster and Namespace
LocalQueueNamespace-level Queue that Kueue uses to manage GPU resource requests
Model PVCPVC that stores model weights. This guide uses models as an example.
GatewayGateway created in advance with the same name as the Namespace
NetworkAttachmentDefinitionNetwork definition attached to workloads that use InfiniBand. It uses the same name as the Namespace.
External hostnameHostname under the *.moduai.kakaocloud.com domain used for model serving

Check resources​

  1. Check the current context and default Namespace.

    kubectl config current-context
    kubectl config view --minify --output 'jsonpath={..namespace}'
  2. Check the common resources provided in the Namespace.

    kubectl get localqueue
    kubectl get pvc models
    kubectl get gateway
    kubectl get network-attachment-definitions
  3. Verify that the model files are available in the model PVC. The model path must match the model or uri setting of the serving runtime that you use.

Placeholders in examples​

Replace values in angle brackets (<...>) with values for your environment.

PlaceholderDescriptionExample
<user-namespace>Assigned Kubernetes Namespacegpu-team-a
<queue-name>Name of a LocalQueue created in the Namespaceserving
<service-name>Name of the model serving resourceglm-5-3
<model-path>Model directory in the models PVCGLM-5.3
<hostname>External hostname used by the Gatewayglm-5-3.moduai.kakaocloud.com
<api-key>API key used for Gateway authenticationA value issued for your environment

Apply YAML​

Save the example YAML to a file, and then apply it.

kubectl apply -f <filename>.yaml

To manage multiple files together, create a kustomization.yaml file in the same directory, and then apply it.

kubectl apply -k <directory>
Caution

The CPU, memory, GPU, and InfiniBand NIC counts in the examples are based on GLM-5.3. When using a different model or node specification, adjust the values to match the model requirements and your quota.