Skip to main content

Build a predictive model with Kubeflow Notebooks

Practice with a sample dataset in Jupyter Notebook.

Step 1. Prerequisites​

Download the training data and use the dataset to complete the exercise in Jupyter Notebook.

  • Check the minimum specifications for the node pool used in the exercise.
  • Minimum node pool specifications: at least 4 vCPUs and 8 GB of memory
  • For option 2: a GPU notebook node pool is required
  • To create a GPU pipeline: a GPU pipeline node pool is required
caution

The exercise might not run properly if the node pool does not meet the minimum specifications.

Prepare training data​

In this exercise, you build a taxi fare prediction model using a public dataset from New York City.

Source dataset information

ItemDescription
GoalBuild a taxi fare prediction model
DatasetNew York City Yellow Taxi fare data from 2009–2015
- Pickup and drop-off locations, passenger count, pickup time, and more

Step 2. Create a notebook for the exercise​

info

For more information about creating notebooks, see Use notebooks.

Create a CPU image-based notebook​

The following example creates a notebook that uses only a CPU instance type. Configure the notebook as required for the exercise.

  1. On the Kubeflow dashboard, click the Notebooks tab, and then click New Notebook.

  2. On the New Notebook page, enter the required information and click LAUNCH to create a CPU image-based notebook instance.

    Create a CPU image-based notebook (option 1) Create a CPU image-based notebook (option 1)

    ItemFieldDescription
    NameNameName used to identify the notebook instance on the Kubeflow dashboard
    NamespaceKubernetes namespace where the notebook instance is created
    Docker ImageImagemlops-pipelines/jupyter-pyspark-pytorch:v1.0.1
    - Docker image to use
    CPU / RAMRequested CPUs2
    - Number of CPU cores allocated to the notebook instance
    Requested memory in Gi6
    - Amount of memory in GiB allocated to the notebook instance
    GPUsNumber of GPUsNone
    - GPU resources allocated to the notebook instance
    Affinity / TolerationsAffinity ConfigCPU node pool where the notebook is created
    - Specifies the node on which the notebook instance runs
    Tolerations GroupNone
    - Configuration that permits specific node taints
caution

The Requested CPUs and Requested memory in Gi values must be less than the vCPU count and memory capacity of the node pool selected under Affinity Config.

Create a GPU image-based notebook (option 2)​

The following example creates a notebook that uses only a GPU instance type. Configure the notebook as required for the exercise.

  1. On the Kubeflow dashboard, click the Notebooks tab, and then click New Notebook.

  2. On the New Notebook page, enter the required information and click LAUNCH to create a GPU image-based notebook instance.

    Create a GPU image-based notebook (option 2) Create a GPU image-based notebook (option 2)

    ItemFieldDescription
    NameNameName used to identify the notebook instance on the Kubeflow dashboard
    NamespaceKubernetes namespace where the notebook instance is created
    Docker ImageImagemlops-pipelines/jupyter-pyspark-pytorch-cuda:v1.0.1
    - Docker image to use
    CPU / RAMRequested CPUs2
    - Number of CPU cores allocated to the notebook instance
    Requested memory in Gi6
    - Amount of memory in GiB allocated to the notebook instance
    GPUsNumber of GPUsAt least 1
    - GPU resources allocated to the notebook instance
    GPU VendorNVIDIA MIG - 1g.10gb
    - GPU driver and software toolkit to use
    Affinity / TolerationsAffinity ConfigGPU node pool where the notebook is created
    - Specifies the node on which the notebook instance runs
    Tolerations GroupNone
    - Configuration that permits specific node taints

Step 3. Train a predictive model in a notebook (exercise 1)​

  1. Download the following sample file.

  2. Access the notebook instance you created and upload the file in the browser.

    Upload a file in the Jupyter Notebook console Upload a file in the Jupyter Notebook console

  3. After the upload is complete, review the exercise in the right pane.

    Uploaded sample file Uploaded sample file

  4. Follow the sample instructions to train the model.

Step 4. Create a pipeline and train a predictive model in a notebook (exercise 2)​

  1. Download the following sample files.

    Upload a file in the Jupyter Notebook console Upload a file in the Jupyter Notebook console

  2. After the upload is complete, review the exercise in the right pane.

    Uploaded sample file for exercise 2 Uploaded sample file for exercise 2

  3. Follow the sample instructions to train the model.

    info

    We recommend deleting completed or unused runs as described below to conserve resources.

  4. After completing the exercise, delete the completed run. Move the run to Archived, and then click Delete.

    Delete a run Delete a run

  5. After deleting the run, verify that its pods were also deleted.

    Verify run deletion Verify run deletion