- load_cancer_data — load the dataset and export it as a CSV artifact.
- preprocess_data — read the CSV, split into train/test, scale features, and write train/test CSVs.
- train_and_evaluate — train a logistic regression model with a tunable
Chyperparameter, log metrics, and save the trained model.
Component: load_cancer_data
This component loads the built-in breast cancer dataset and writes it to the component outputdataset. The component installs pandas and scikit-learn in the container at runtime so it can operate independently.
packages_to_installensures required Python packages are present inside the component container.- Writing to
dataset.pathproduces an artifact that downstream components can access via theirInput[Dataset]parameters.
Using
as_frame=True in scikit-learn simplifies CSV export because the dataset is already a pandas DataFrame. This reduces manual column handling when writing the primary dataset artifact.Component: preprocess_data
This component reads the CSV created byload_cancer_data, splits it into train/test sets, standardizes features with StandardScaler, and writes train_data and test_data as CSV artifacts for downstream steps.
- We create stable column names (
f0,f1, …) for the scaled features so CSV schemas remain consistent across pipeline runs. - Both outputs are written to
.path, exposing them as artifact files to later components.
Component: train_and_evaluate
This component consumes the preprocessed CSVs, trains a logistic regression model with tunable hyperparameterC, evaluates accuracy, logs metrics via KFP Metrics, and serializes the trained model to the model output. The component installs pandas and scikit-learn at runtime.
- Ensure training uses
X_trainandy_trainin.fit(...)(not the test data). metrics.log_metricvalues must be serializable primitives; casting tofloatensures compatibility.pickleis included in the Python standard library, so no extra package installation is required.
Pipeline: cancer_pipeline
This pipeline composes the three components and demonstrates a simple hyperparameter sweep usingParallelFor to run multiple train_and_evaluate tasks in parallel for different C values.
- Always call a component to create a task (e.g.,
task = component_fn()), otherwise the returned object will be the component definition rather than a runnable task and.outputswon’t exist. - Use artifact names (like
"dataset","train_data","test_data") consistently between components and when accessing.outputs[...]. - Parallelism with
ParallelForlets you efficiently evaluate multiple hyperparameter values; each iteration receives the loop variable (hereC) and schedules its own task.
A common error is forgetting the parentheses when calling a component. For example, writing
load_task = load_cancer_data (missing ()) will assign the component definition itself — not a task — and accessing load_task.outputs will raise: AttributeError: ‘PythonComponent’ object has no attribute ‘outputs’. Always call the component (e.g., load_cancer_data()) to obtain a task object.- Kubeflow Pipelines (KFP) Python SDK: https://github.com/kubeflow/pipelines
- scikit-learn: https://scikit-learn.org/
- pandas: https://pandas.pydata.org/
- KFP component authoring: https://www.kubeflow.org/docs/components/pipelines/sdk/sdk-overview/