Skip to main content
In this lesson you’ll learn how to pass files and structured data between Kubeflow Pipelines components using artifacts. Artifacts are the recommended mechanism when you need to exchange datasets, trained models, or other files that are larger or structured. For small scalar values, pipeline parameters still make sense. This guide demonstrates a simple, end‑to‑end pipeline that uses artifacts to move:
  • a CSV dataset,
  • a processed CSV,
  • a pickled model file,
  • and logged metrics.
The typical flow in the example pipeline below is:
  1. retrieve_data — generate and write a CSV dataset (artifact).
  2. preprocess_data — read the CSV, add a feature, write a processed CSV.
  3. train_model — read the processed CSV, train a trivial model, write a pickled model artifact.
  4. evaluate_model — read the pickled model, compute and log a metric.
Use artifacts whenever you need to exchange files or structured data between components. Inside a component the artifact location is available via the .path attribute on Input[...] and Output[...] typed arguments.
When to use artifacts vs parameters:
  • Use artifacts for files, datasets, models, and any structured or large content.
  • Use parameters for scalars (strings, numbers, booleans) and small configuration values.
Quick reference table for commonly used artifact types: Complete example pipeline Below is a complete, cleaned example implementing the pipeline described above. Save it as pipeline_artifact.py and compile to YAML to upload to the Kubeflow Pipelines UI.
Key implementation details and patterns
  • Declare artifact inputs and outputs as typed function arguments:
    • Example outputs: output_data: Output[Dataset], model: Output[Model], metrics: Output[Metrics]
    • Example inputs: raw_data: Input[Dataset], training_data: Input[Dataset], model: Input[Model]
  • Use the .path attribute inside the component to read/write the artifact contents:
    • Write CSV: df.to_csv(processed_data.path, index=False)
    • Read CSV: df = pd.read_csv(raw_data.path)
    • Save model: with open(model.path, "wb") as f: pickle.dump(..., f)
    • Load model: with open(model.path, "rb") as f: model_data = pickle.load(f)
  • Wire components in the pipeline by referencing a producer task’s outputs by name:
    • preprocess_data(raw_data=data_task.outputs["output_data"])
    • train_model(training_data=preprocess_task.outputs["processed_data"])
    • evaluate_model(model=train_task.outputs["model"])
Avoid passing large files as parameters. Parameters are for small scalar values — use Input[...] / Output[...] artifacts for files and structured data to prevent serialization and size issues.
Compile, upload, and inspect
  1. Compile the pipeline: python pipeline_artifact.py will generate artifact_pipeline.yaml.
  2. Upload artifact_pipeline.yaml to the Kubeflow Pipelines UI and run the pipeline.
  3. In the UI you can inspect:
    • Input/output artifacts for each task,
    • Object storage locations (e.g., MinIO),
    • Previews of supported artifact types,
    • Logged metrics such as accuracy.
You should see artifacts in the run UI for the original CSV, the processed CSV (with x_squared), the pickled model file, and the metrics output.
A screenshot of the Kubeflow Central Dashboard showing a pipeline run graph on the left and an Artifact Visualization panel on the right. The panel lists scalar metrics including "accuracy" with a value of 1.
Links and references

Watch Video