> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Demo Section Project

> Compact Kubeflow Pipelines example that loads breast cancer data, preprocesses it, trains and evaluates a logistic regression, logs metrics, and runs a parallel hyperparameter sweep.

In this lesson we build a compact Kubeflow Pipelines (KFP) example that represents a realistic machine learning workflow: load data, preprocess it, then train and evaluate a model. We use the breast cancer dataset from scikit-learn and train a logistic regression model to predict whether a tumor is malignant or benign. This example demonstrates KFP component inputs/outputs, artifact passing, logging metrics, and simple hyperparameter search with parallelism.

Key pipeline steps

1. load\_cancer\_data — load the dataset and export it as a CSV artifact.
2. preprocess\_data — read the CSV, split into train/test, scale features, and write train/test CSVs.
3. train\_and\_evaluate — train a logistic regression model with a tunable `C` hyperparameter, log metrics, and save the trained model.

Table: components at a glance

| Component | Purpose | Outputs |
| -: | - | - |
| `load_cancer_data` | Load scikit-learn breast cancer dataset and export as CSV | `dataset` (CSV) |
| `preprocess_data` | Split and standardize features; produce train/test CSVs | `train_data`, `test_data` (CSV) |
| `train_and_evaluate` | Train LogisticRegression, evaluate accuracy, save model | `metrics`, `model` (pickle) |

## Component: load\_cancer\_data

This component loads the built-in breast cancer dataset and writes it to the component output `dataset`. The component installs `pandas` and `scikit-learn` in the container at runtime so it can operate independently.

```python theme={null}
from kfp.dsl import component, Output, Dataset

@component(
    packages_to_install=["pandas", "scikit-learn"]
)
def load_cancer_data(dataset: Output[Dataset]):
    from sklearn.datasets import load_breast_cancer
    import pandas as pd

    data = load_breast_cancer(as_frame=True)
    df = data.frame

    df.to_csv(dataset.path, index=False)
```

What to notice

* `packages_to_install` ensures required Python packages are present inside the component container.
* Writing to `dataset.path` produces an artifact that downstream components can access via their `Input[Dataset]` parameters.

<Callout icon="lightbulb" color="#1CB2FE">
  Using `as_frame=True` in scikit-learn simplifies CSV export because the dataset is already a pandas DataFrame. This reduces manual column handling when writing the primary dataset artifact.
</Callout>

## Component: preprocess\_data

This component reads the CSV created by `load_cancer_data`, splits it into train/test sets, standardizes features with `StandardScaler`, and writes `train_data` and `test_data` as CSV artifacts for downstream steps.

```python theme={null}
from kfp.dsl import component, Input, Output, Dataset

@component(packages_to_install=["pandas", "scikit-learn"])
def preprocess_data(
    dataset: Input[Dataset],
    train_data: Output[Dataset],
    test_data: Output[Dataset]
):
    import pandas as pd
    from sklearn.model_selection import train_test_split
    from sklearn.preprocessing import StandardScaler

    # Read input dataset
    df = pd.read_csv(dataset.path)

    # Features and target
    X = df.drop("target", axis=1)
    y = df["target"]

    # Split
    X_train, X_test, y_train, y_test = train_test_split(
        X, y, test_size=0.25, random_state=42, stratify=y
    )

    # Standardize features
    scaler = StandardScaler()
    X_train = scaler.fit_transform(X_train)
    X_test = scaler.transform(X_test)

    # Recreate DataFrames with target column for downstream components
    train_df = pd.DataFrame(X_train, columns=[f"f{i}" for i in range(X_train.shape[1])])
    train_df["target"] = y_train.values

    test_df = pd.DataFrame(X_test, columns=[f"f{i}" for i in range(X_test.shape[1])])
    test_df["target"] = y_test.values

    # Write outputs
    train_df.to_csv(train_data.path, index=False)
    test_df.to_csv(test_data.path, index=False)
```

Implementation details

* We create stable column names (`f0`, `f1`, ...) for the scaled features so CSV schemas remain consistent across pipeline runs.
* Both outputs are written to `.path`, exposing them as artifact files to later components.

## Component: train\_and\_evaluate

This component consumes the preprocessed CSVs, trains a logistic regression model with tunable hyperparameter `C`, evaluates accuracy, logs metrics via KFP `Metrics`, and serializes the trained model to the `model` output. The component installs `pandas` and `scikit-learn` at runtime.

```python theme={null}
from kfp.dsl import component, Input, Output, Dataset, Metrics, Model

@component(packages_to_install=["pandas", "scikit-learn"])
def train_and_evaluate(
    train_data: Input[Dataset],
    test_data: Input[Dataset],
    C: float,
    metrics: Output[Metrics],
    model: Output[Model],
):
    import pandas as pd
    from sklearn.linear_model import LogisticRegression
    from sklearn.metrics import accuracy_score
    import pickle

    # Load data
    train_df = pd.read_csv(train_data.path)
    test_df = pd.read_csv(test_data.path)

    X_train = train_df.drop("target", axis=1)
    y_train = train_df["target"]

    X_test = test_df.drop("target", axis=1)
    y_test = test_df["target"]

    # Train
    clf = LogisticRegression(C=C, max_iter=500, solver="liblinear")
    clf.fit(X_train, y_train)

    # Predict and evaluate
    preds = clf.predict(X_test)
    accuracy = float(accuracy_score(y_test, preds))

    # Log metrics
    metrics.log_metric("accuracy", accuracy)
    metrics.log_metric("C", float(C))

    # Save model (pickle)
    with open(model.path, "wb") as f:
        pickle.dump(clf, f)
```

Corrections and clarifications

* Ensure training uses `X_train` and `y_train` in `.fit(...)` (not the test data).
* `metrics.log_metric` values must be serializable primitives; casting to `float` ensures compatibility.
* `pickle` is included in the Python standard library, so no extra package installation is required.

## Pipeline: cancer\_pipeline

This pipeline composes the three components and demonstrates a simple hyperparameter sweep using `ParallelFor` to run multiple `train_and_evaluate` tasks in parallel for different `C` values.

```python theme={null}
from kfp.dsl import pipeline, ParallelFor
from kfp import compiler

@pipeline(name="breast-cancer-ml-pipeline")
def cancer_pipeline():
    # hyperparameter values to evaluate
    C_values = [0.01, 0.1, 1.0, 10.0, 100.0]

    # Execute components
    load_task = load_cancer_data()  # call component to create a task
    preprocess_task = preprocess_data(dataset=load_task.outputs["dataset"])

    # Parallelize training across C values
    with ParallelFor(C_values) as C:
        train_and_evaluate(
            train_data=preprocess_task.outputs["train_data"],
            test_data=preprocess_task.outputs["test_data"],
            C=C
        )

if __name__ == "__main__":
    compiler.Compiler().compile(
        pipeline_func=cancer_pipeline,
        package_path="cancer_pipeline.yaml"
    )
```

Best practices and usage notes

* Always call a component to create a task (e.g., `task = component_fn()`), otherwise the returned object will be the component definition rather than a runnable task and `.outputs` won't exist.
* Use artifact names (like `"dataset"`, `"train_data"`, `"test_data"`) consistently between components and when accessing `.outputs[...]`.
* Parallelism with `ParallelFor` lets you efficiently evaluate multiple hyperparameter values; each iteration receives the loop variable (here `C`) and schedules its own task.

<Callout icon="warning" color="#FF6B6B">
  A common error is forgetting the parentheses when calling a component. For example, writing `load_task = load_cancer_data` (missing `()`) will assign the component definition itself — not a task — and accessing `load_task.outputs` will raise: AttributeError: 'PythonComponent' object has no attribute 'outputs'. Always call the component (e.g., `load_cancer_data()`) to obtain a task object.
</Callout>

Links and references

* Kubeflow Pipelines (KFP) Python SDK: [https://github.com/kubeflow/pipelines](https://github.com/kubeflow/pipelines)
* scikit-learn: [https://scikit-learn.org/](https://scikit-learn.org/)
* pandas: [https://pandas.pydata.org/](https://pandas.pydata.org/)
* KFP component authoring: [https://www.kubeflow.org/docs/components/pipelines/sdk/sdk-overview/](https://www.kubeflow.org/docs/components/pipelines/sdk/sdk-overview/)

That's it — this example shows how to build a clear KFP workflow that loads a dataset, preprocesses it, trains and evaluates a model, logs metrics, and performs a simple hyperparameter sweep using parallelism.

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/kubeflow/module/bece8da9-953e-480e-8774-b25b66c3830f/lesson/52dbcdfd-3d54-4c41-8b85-151fa565e26c" />

  <Card title="Practice Lab" icon="flask-conical" cta="Learn more" href="https://learn.kodekloud.com/user/courses/kubeflow/module/bece8da9-953e-480e-8774-b25b66c3830f/lesson/b6ecd722-385d-41b3-9a04-c019aec6fb9d" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.