Skip to main content
In this lesson we build a compact Kubeflow Pipelines (KFP) example that represents a realistic machine learning workflow: load data, preprocess it, then train and evaluate a model. We use the breast cancer dataset from scikit-learn and train a logistic regression model to predict whether a tumor is malignant or benign. This example demonstrates KFP component inputs/outputs, artifact passing, logging metrics, and simple hyperparameter search with parallelism. Key pipeline steps
  1. load_cancer_data — load the dataset and export it as a CSV artifact.
  2. preprocess_data — read the CSV, split into train/test, scale features, and write train/test CSVs.
  3. train_and_evaluate — train a logistic regression model with a tunable C hyperparameter, log metrics, and save the trained model.
Table: components at a glance

Component: load_cancer_data

This component loads the built-in breast cancer dataset and writes it to the component output dataset. The component installs pandas and scikit-learn in the container at runtime so it can operate independently.
What to notice
  • packages_to_install ensures required Python packages are present inside the component container.
  • Writing to dataset.path produces an artifact that downstream components can access via their Input[Dataset] parameters.
Using as_frame=True in scikit-learn simplifies CSV export because the dataset is already a pandas DataFrame. This reduces manual column handling when writing the primary dataset artifact.

Component: preprocess_data

This component reads the CSV created by load_cancer_data, splits it into train/test sets, standardizes features with StandardScaler, and writes train_data and test_data as CSV artifacts for downstream steps.
Implementation details
  • We create stable column names (f0, f1, …) for the scaled features so CSV schemas remain consistent across pipeline runs.
  • Both outputs are written to .path, exposing them as artifact files to later components.

Component: train_and_evaluate

This component consumes the preprocessed CSVs, trains a logistic regression model with tunable hyperparameter C, evaluates accuracy, logs metrics via KFP Metrics, and serializes the trained model to the model output. The component installs pandas and scikit-learn at runtime.
Corrections and clarifications
  • Ensure training uses X_train and y_train in .fit(...) (not the test data).
  • metrics.log_metric values must be serializable primitives; casting to float ensures compatibility.
  • pickle is included in the Python standard library, so no extra package installation is required.

Pipeline: cancer_pipeline

This pipeline composes the three components and demonstrates a simple hyperparameter sweep using ParallelFor to run multiple train_and_evaluate tasks in parallel for different C values.
Best practices and usage notes
  • Always call a component to create a task (e.g., task = component_fn()), otherwise the returned object will be the component definition rather than a runnable task and .outputs won’t exist.
  • Use artifact names (like "dataset", "train_data", "test_data") consistently between components and when accessing .outputs[...].
  • Parallelism with ParallelFor lets you efficiently evaluate multiple hyperparameter values; each iteration receives the loop variable (here C) and schedules its own task.
A common error is forgetting the parentheses when calling a component. For example, writing load_task = load_cancer_data (missing ()) will assign the component definition itself — not a task — and accessing load_task.outputs will raise: AttributeError: ‘PythonComponent’ object has no attribute ‘outputs’. Always call the component (e.g., load_cancer_data()) to obtain a task object.
Links and references That’s it — this example shows how to build a clear KFP workflow that loads a dataset, preprocesses it, trains and evaluates a model, logs metrics, and performs a simple hyperparameter sweep using parallelism.

Watch Video

Practice Lab