Skip to main content
In this lesson we’ll cover reliable ways to pass data between components in a pipeline. Typical pipelines include stages such as: fetch data → process data → train model → evaluate model. Each stage must hand off the right information to the next — for example, fetch provides raw data to process, process provides prepared data to training, and training produces a model for evaluation.
A simple flowchart titled "Passing Data" with four boxes: Fetch Data → Process Data → Train Model → Evaluate Model. Arrows between the boxes are labeled "Unprocessed Data", "Processed Data", and "Model" to show the data flow.
There are two common approaches to pass data between components: parameters and artifacts. Use parameters for small scalar values and artifacts for file-based or large outputs. The following sections explain both methods with examples and best practices.

Overview: Parameters vs Artifacts

How to pass data — Parameters

Parameters are the simplest method. They pass scalar values by value between components. Supported types commonly include integers, floats, strings, booleans, and small JSON-serializable objects. Typical parameter uses are hyperparameters, evaluation thresholds, or small metadata.
A presentation slide titled "How to Pass Data – Parameters" explaining that scalar values are passed between components. It lists supported types: Int, Float, Str, Bool and small JSON (serializable objects).
A presentation slide titled "How to Pass Data – Parameters" that lists limitations, shown in two gray boxes: "Size – Limited (do not pass large datasets)" and "Not suitable for Files or large objects."
Limitations:
  • Size limits: parameters are not appropriate for large datasets or files.
  • Not suited for binary objects or large structured outputs — use artifacts in those cases.
Example — single scalar parameter Return the scalar from the producing component and consume it as an input in the next component. For a single return value you can access the value with .output on the producing task.
Example — multiple named outputs (NamedTuple) When a component needs to return multiple scalar values, use a NamedTuple (or other typed multi-output). The component returns values in the declared order, and the pipeline can read them from the producing task’s .outputs dictionary.
Note: Single-return components expose .output while multi-output components expose .outputs["name"].

How to pass data — Artifacts

Artifacts are file-based outputs stored in object storage (for example, MinIO in many Kubeflow setups or Amazon S3). Artifacts are passed by reference: the producing component writes files to an artifact path and the consuming component reads from that path. This makes artifacts ideal for large datasets, trained models, evaluation results, and reproducible outputs.
A presentation slide titled "Passing Data – Artifacts" with two notes: "File-based output stored in object storage" and "Passed by reference, not by value." Below are MinIO and S3-style object storage icons and a KodeKloud copyright.
Common artifact types include Dataset, Model, Metrics, HTML, and a generic Artifact type:
A presentation slide titled "Passing Data – Artifacts" showing a dotted box labeled "Common Artifact Types" with five rounded boxes: Dataset, Model, Metrics, HTML, and Artifact (generic). A small "© Copyright KodeKloud" appears in the bottom corner.
Artifact usage pattern:
  • The producing component declares an Output[...] typed artifact (for example, Output[Dataset]) and writes files to the provided .path.
  • The consuming component declares an Input[...] typed artifact (for example, Input[Dataset]) and reads from .path.
  • The pipeline wires producer to consumer by passing producer_task.outputs["artifact_name"] into the consumer argument.
Example pipeline using artifacts This example demonstrates retrieving raw data, preprocessing it, training a model, and saving the trained model as an artifact. Components use Output and Input artifact typing and read/write files via the .path property.
Notes on wiring and names:
  • The output parameter name in the component signature (for example, output_data: Output[Dataset]) becomes the key in the producing task’s .outputs dictionary.
  • To read or write files for an artifact, use the provided .path attribute (for example, raw_data.path or model.path).
  • Artifacts are stored in the cluster’s object storage (e.g., MinIO), and consumers are given the path reference to access them.
Do not pass large datasets, binary files, or trained models as parameters. Parameters are size-limited and intended for scalars or small JSON-serializable content only.
Use parameters for small scalar values (hyperparameters, flags, small JSON). Use artifacts for large files, datasets, models, and results — artifacts are stored in object storage and passed by reference.

Quick Best-Practices

  • Use parameters for hyperparameters, simple thresholds, or small metadata.
  • Use artifacts for datasets, model files, images, or any outputs too large to inline.
  • Name artifact outputs clearly in component signatures to create predictable keys in .outputs.
  • Prefer artifact-based workflows for reproducibility and when working across clusters.

Watch Video