Skip to main content
In this lesson we’ll explore Kubeflow Pipelines — the component of Kubeflow that automates and orchestrates end-to-end ML workflows. A typical ML workflow includes tasks such as data collection, validation, cleaning, feature engineering, model training, evaluation, versioning, deployment, monitoring, and retraining. Kubeflow Pipelines runs each step inside its own Kubernetes pod so tasks are isolated, scalable, and managed by Kubernetes while Kubeflow handles orchestration, retries, and artifact tracking.
A Kubeflow Pipeline diagram illustrating an automated ML workflow. It shows stages like data collection/ingestion, validation, cleaning/preprocessing, feature engineering, model training, evaluation, versioning, deployment and monitoring with Kubernetes pod icons and a retraining loop.
Kubeflow executes the first step (for example, data collection) inside one pod, waits for it to complete, then runs the next step (data validation) in a different pod, and so on. This stepwise execution enforces task dependencies and makes the workflow reproducible and auditable.

How pipelines are authored

Pipelines are defined in Python using the Kubeflow Pipelines SDK and then compiled into a YAML package that Kubeflow accepts. Writing pipelines in Python gives you access to the broader Python ecosystem (libraries, testing, packaging) while letting the SDK generate the YAML representation for runtime execution.
A slide titled "How to Create a Pipeline" illustrating a Python file icon labeled pipeline.py with an arrow pointing to a YAML file icon labeled pipeline.yaml, indicating conversion or generation of a YAML pipeline from Python.
After compilation you can upload the YAML package to Kubeflow either via the web UI or programmatically (for CI/CD). Kubeflow will render the pipeline graph, allow parameterization, and trigger runs.

Pipeline structure and components

Consider a simple pipeline consisting of three stages: data collection & ingestion, data cleaning & preprocessing, and model training. In Kubeflow Pipelines each stage is modeled as a component. Components are isolated units of work that run in their own pods and encapsulate one distinct step of the workflow.
A slide titled "Creating Pipelines" showing a vertical flowchart. It has three turquoise boxes labeled "Data Collection & Ingestion," "Data Cleaning & Preprocessing," and "Model Training," linked by arrows and small "component" tags.
Key concepts at a glance:

Install the Kubeflow Pipelines SDK

Install the SDK locally before authoring pipelines:
Check the Kubeflow Pipelines SDK version compatibility with your Kubeflow deployment. The SDK API can vary between major versions (for example, v1 vs v2), so install the version that matches your cluster’s Pipelines component.

Writing a simple pipeline

Each component can be defined as a Python function and decorated to become a Kubeflow component. The pipeline itself is a function decorated with @dsl.pipeline. Invoking component functions inside the pipeline function creates tasks; you can set ordering with .after() or rely on implicit dependencies when a task’s output is passed to another. Example pipeline code:
Notes on the example:
  • Each @dsl.component-decorated function becomes a reusable pipeline component that will be executed inside its own pod.
  • Calling a component (for example, data_collection()) returns a task object. Use .after() to enforce ordering for tasks that do not exchange data.
  • If a component returns a value or artifact that is consumed by another component, the SDK infers the dependency automatically — you do not need .after().
  • compiler.Compiler().compile(...) converts the Python pipeline into a YAML package (demo_pipeline.yaml) which you can upload to Kubeflow.
If your components exchange data (for example, one component returns a file path or artifact consumed by the next), the dependency is implicit and you do not need to call .after(). Use .after() when components have no direct data edges but you still need to enforce ordering.

Uploading and running the pipeline

Once the YAML package is generated, upload it to Kubeflow:
  1. In the Kubeflow UI go to Pipelines → Upload Pipeline.
  2. Choose Import from file or URL and provide a pipeline name.
  3. Select the compiled YAML file (demo_pipeline.yaml) and click Create.
  4. The UI displays the pipeline graph and lets you start a new run with parameters and experiment settings.
After you launch a run, Kubeflow creates pods for each task, shows logs per step, and stores artifacts and metadata for lineage and reproducibility.

Watch Video