> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Passing Data Between Components

> Guide to passing data between Kubeflow pipeline components using parameters for small scalars and artifacts for file-based large outputs, with examples and best practices

In this lesson we’ll cover reliable ways to pass data between components in a pipeline. Typical pipelines include stages such as: fetch data → process data → train model → evaluate model. Each stage must hand off the right information to the next — for example, fetch provides raw data to process, process provides prepared data to training, and training produces a model for evaluation.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/MGkgrGfKHDtoCnUb/images/Kubeflow/Working-With-Kubeflow/Passing-Data-Between-Components/passing-data-fetch-process-train-evaluate.jpg?fit=max&auto=format&n=MGkgrGfKHDtoCnUb&q=85&s=b213e9bc1a399c9a0b2153a92301ba91" alt="A simple flowchart titled &#x22;Passing Data&#x22; with four boxes: Fetch Data → Process Data → Train Model → Evaluate Model. Arrows between the boxes are labeled &#x22;Unprocessed Data&#x22;, &#x22;Processed Data&#x22;, and &#x22;Model&#x22; to show the data flow." width="1920" height="1080" data-path="images/Kubeflow/Working-With-Kubeflow/Passing-Data-Between-Components/passing-data-fetch-process-train-evaluate.jpg" />
</Frame>

There are two common approaches to pass data between components: parameters and artifacts. Use parameters for small scalar values and artifacts for file-based or large outputs. The following sections explain both methods with examples and best practices.

## Overview: Parameters vs Artifacts

| Mechanism | Typical Use Cases | Passing Semantics | When to Use |
| - | -: | -: | - |
| Parameters | Hyperparameters, thresholds, small numeric or string values | Passed by value (inline) | Use for scalars, flags, or small JSON-serializable objects |
| Artifacts | Datasets, trained models, reports, large JSON/CSV files | Passed by reference (object storage paths) | Use for large files, binary objects, and reproducible artifacts |

## How to pass data — Parameters

Parameters are the simplest method. They pass scalar values by value between components. Supported types commonly include integers, floats, strings, booleans, and small JSON-serializable objects. Typical parameter uses are hyperparameters, evaluation thresholds, or small metadata.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/MGkgrGfKHDtoCnUb/images/Kubeflow/Working-With-Kubeflow/Passing-Data-Between-Components/pass-data-parameters-scalars-json.jpg?fit=max&auto=format&n=MGkgrGfKHDtoCnUb&q=85&s=0ff441e3cbdfb710913c5a91b8d92d29" alt="A presentation slide titled &#x22;How to Pass Data – Parameters&#x22; explaining that scalar values are passed between components. It lists supported types: Int, Float, Str, Bool and small JSON (serializable objects)." width="1920" height="1080" data-path="images/Kubeflow/Working-With-Kubeflow/Passing-Data-Between-Components/pass-data-parameters-scalars-json.jpg" />
</Frame>

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/MGkgrGfKHDtoCnUb/images/Kubeflow/Working-With-Kubeflow/Passing-Data-Between-Components/parameters-pass-data-size-limited-files.jpg?fit=max&auto=format&n=MGkgrGfKHDtoCnUb&q=85&s=b895c0fec3cf224af03cd3549cecb918" alt="A presentation slide titled &#x22;How to Pass Data – Parameters&#x22; that lists limitations, shown in two gray boxes: &#x22;Size – Limited (do not pass large datasets)&#x22; and &#x22;Not suitable for Files or large objects.&#x22;" width="1920" height="1080" data-path="images/Kubeflow/Working-With-Kubeflow/Passing-Data-Between-Components/parameters-pass-data-size-limited-files.jpg" />
</Frame>

Limitations:

* Size limits: parameters are not appropriate for large datasets or files.
* Not suited for binary objects or large structured outputs — use artifacts in those cases.

Example — single scalar parameter
Return the scalar from the producing component and consume it as an input in the next component. For a single return value you can access the value with `.output` on the producing task.

```python theme={null}
from kfp.dsl import component, pipeline

@component
def step1() -> int:
    stats = 400
    return stats

@component
def step2(stats: int) -> None:
    # Pretend training logic
    print(f"stats: {stats}")

@pipeline()
def ml_pipeline():
    s1 = step1()
    # Use the single scalar return via s1.output
    s2 = step2(stats=s1.output)
```

Example — multiple named outputs (NamedTuple)
When a component needs to return multiple scalar values, use a `NamedTuple` (or other typed multi-output). The component returns values in the declared order, and the pipeline can read them from the producing task’s `.outputs` dictionary.

```python theme={null}
from typing import NamedTuple
from kfp.dsl import component, pipeline

Step1Outputs = NamedTuple("Step1Outputs", [("data", int), ("stats", int)])

@component
def step1() -> Step1Outputs:
    data = 20
    stats = 400
    return (data, stats)

@component
def step2(data: int, stats: int) -> None:
    # Pretend training logic
    print(f"data: {data}, stats: {stats}")

@pipeline()
def ml_pipeline():
    s1 = step1()
    # Use named outputs via s1.outputs["<name>"]
    s2 = step2(data=s1.outputs["data"], stats=s1.outputs["stats"])
```

Note: Single-return components expose `.output` while multi-output components expose `.outputs["name"]`.

## How to pass data — Artifacts

Artifacts are file-based outputs stored in object storage (for example, [MinIO](https://min.io/) in many Kubeflow setups or [Amazon S3](https://aws.amazon.com/s3/)). Artifacts are passed by reference: the producing component writes files to an artifact path and the consuming component reads from that path. This makes artifacts ideal for large datasets, trained models, evaluation results, and reproducible outputs.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/MGkgrGfKHDtoCnUb/images/Kubeflow/Working-With-Kubeflow/Passing-Data-Between-Components/artifacts-object-storage-minio-s3.jpg?fit=max&auto=format&n=MGkgrGfKHDtoCnUb&q=85&s=c95b9d333d12236831fb5a7d800acda7" alt="A presentation slide titled &#x22;Passing Data – Artifacts&#x22; with two notes: &#x22;File-based output stored in object storage&#x22; and &#x22;Passed by reference, not by value.&#x22; Below are MinIO and S3-style object storage icons and a KodeKloud copyright." width="1920" height="1080" data-path="images/Kubeflow/Working-With-Kubeflow/Passing-Data-Between-Components/artifacts-object-storage-minio-s3.jpg" />
</Frame>

Common artifact types include Dataset, Model, Metrics, HTML, and a generic Artifact type:

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/MGkgrGfKHDtoCnUb/images/Kubeflow/Working-With-Kubeflow/Passing-Data-Between-Components/passing-data-artifacts-common-types.jpg?fit=max&auto=format&n=MGkgrGfKHDtoCnUb&q=85&s=634667816bd1d019f6ccbfd1342f4a85" alt="A presentation slide titled &#x22;Passing Data – Artifacts&#x22; showing a dotted box labeled &#x22;Common Artifact Types&#x22; with five rounded boxes: Dataset, Model, Metrics, HTML, and Artifact (generic). A small &#x22;© Copyright KodeKloud&#x22; appears in the bottom corner." width="1920" height="1080" data-path="images/Kubeflow/Working-With-Kubeflow/Passing-Data-Between-Components/passing-data-artifacts-common-types.jpg" />
</Frame>

Artifact usage pattern:

* The producing component declares an `Output[...]` typed artifact (for example, `Output[Dataset]`) and writes files to the provided `.path`.
* The consuming component declares an `Input[...]` typed artifact (for example, `Input[Dataset]`) and reads from `.path`.
* The pipeline wires producer to consumer by passing `producer_task.outputs["artifact_name"]` into the consumer argument.

Example pipeline using artifacts
This example demonstrates retrieving raw data, preprocessing it, training a model, and saving the trained model as an artifact. Components use `Output` and `Input` artifact typing and read/write files via the `.path` property.

```python theme={null}
from kfp.dsl import Input, Output, Dataset, Model, component, pipeline

@component(packages_to_install=["pandas"])
def retrieve_data(output_data: Output[Dataset]):
    import pandas as pd
    df = pd.DataFrame({
        "x": [1, 2, 3, 4, 5],
        "y": [2, 4, 6, 8, 10]
    })
    # Save CSV to the artifact path provided by the runtime
    df.to_csv(output_data.path, index=False)

@component(packages_to_install=["pandas"])
def preprocess_data(
    raw_data: Input[Dataset],
    processed_data: Output[Dataset]
):
    import pandas as pd

    # Read the CSV from the artifact path
    df = pd.read_csv(raw_data.path)

    # Simple "feature engineering"
    df["x_squared"] = df["x"] ** 2

    # Write processed CSV back to the artifact path
    df.to_csv(processed_data.path, index=False)

@component(packages_to_install=["pandas", "scikit-learn", "cloudpickle"])
def train_model(
    training_data: Input[Dataset],
    model: Output[Model]
):
    import pandas as pd
    from sklearn.linear_model import LinearRegression
    import cloudpickle

    # Read training data
    df = pd.read_csv(training_data.path)
    X = df[["x", "x_squared"]]
    y = df["y"]

    # Train a simple model
    reg = LinearRegression()
    reg.fit(X, y)

    # Save the trained model to the artifact path
    with open(model.path, "wb") as f:
        cloudpickle.dump(reg, f)

@component(packages_to_install=["cloudpickle"])
def evaluate_model(model: Input[Model]):
    import cloudpickle
    # Load model from artifact path and print a basic message
    with open(model.path, "rb") as f:
        reg = cloudpickle.load(f)
    print("Loaded model for evaluation:", reg)

@pipeline(name="dummy-artifact-pipeline")
def dummy_pipeline():
    data_task = retrieve_data()
    preprocess_task = preprocess_data(raw_data=data_task.outputs["output_data"])
    train_task = train_model(training_data=preprocess_task.outputs["processed_data"])
    evaluate_task = evaluate_model(model=train_task.outputs["model"])
```

Notes on wiring and names:

* The output parameter name in the component signature (for example, `output_data: Output[Dataset]`) becomes the key in the producing task’s `.outputs` dictionary.
* To read or write files for an artifact, use the provided `.path` attribute (for example, `raw_data.path` or `model.path`).
* Artifacts are stored in the cluster’s object storage (e.g., [MinIO](https://min.io/)), and consumers are given the path reference to access them.

<Callout icon="warning" color="#FF6B6B">
  Do not pass large datasets, binary files, or trained models as parameters. Parameters are size-limited and intended for scalars or small JSON-serializable content only.
</Callout>

<Callout icon="lightbulb" color="#1CB2FE">
  Use parameters for small scalar values (hyperparameters, flags, small JSON). Use artifacts for large files, datasets, models, and results — artifacts are stored in object storage and passed by reference.
</Callout>

## Quick Best-Practices

* Use parameters for hyperparameters, simple thresholds, or small metadata.
* Use artifacts for datasets, model files, images, or any outputs too large to inline.
* Name artifact outputs clearly in component signatures to create predictable keys in `.outputs`.
* Prefer artifact-based workflows for reproducibility and when working across clusters.

## Links and references

* [Kubeflow Pipelines (KFP) SDK documentation](https://www.kubeflow.org/docs/components/pipelines/)
* [MinIO object storage](https://min.io/)
* [Amazon S3](https://aws.amazon.com/s3/)

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/kubeflow/module/bece8da9-953e-480e-8774-b25b66c3830f/lesson/a08dda33-4a9c-463b-8cdb-f286d61fd954" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.