Skip to main content
This guide shows how to wire individual components into a complete Kubeflow Pipeline using the KFP v2 DSL. You’ll learn how component calls without explicit dependencies run in parallel, how to enforce serial execution, how to compile the pipeline into a YAML package, and how to run and inspect the pipeline in Kubeflow Pipelines. In this walkthrough we use simple placeholder components that print messages. In production, components will perform data I/O, return typed artifacts or primitive outputs, and pass artifacts between steps. What you’ll do:
  • See how calling components without dependencies results in parallel execution
  • Enforce serial ordering so components run one after another
  • Compile the pipeline Python definition to a YAML package
  • Upload and run the pipeline in Kubeflow and inspect pods and logs

1) Components: minimal definitions

Below is a minimal, correct set of component definitions using the KFP v2 DSL. These example components return None for simplicity. In real pipelines, prefer typed Output and Input artifacts to pass data between components.
Tip: Define inputs/outputs for components when you need to pass artifacts (datasets, models, metrics) between steps. Relying only on execution ordering does not transfer data.

2) Calling components without dependencies (parallel execution)

If you call component functions inside a pipeline without specifying any ordering, KFP treats each call as an independent task. Without dependencies, the scheduler may run tasks concurrently (parallel execution). This is useful for independent tasks that can run simultaneously.
When visualized in the Kubeflow UI, these steps appear side-by-side in the graph, indicating potential parallel runs. Use this pattern only when tasks do not rely on each other’s results.

3) Specifying ordering (run in series)

To force tasks to run in a specific sequence, use the .after() method on the task object returned by a component invocation. .after() guarantees a task does not start until the specified predecessor completes.
This enforces the following execution order:
  1. data_collection_processing
  2. feature_engineering (after data_collection_processing)
  3. model_training (after feature_engineering)
  4. model_deployment (after model_training)
Use .after() when the start of a task depends on the completion of a previous task. For data transfer between steps, prefer explicit typed Output and Input artifact parameters; ordering alone does not pass artifacts.
Do not rely solely on .after() to move data between components. Define and use typed inputs/outputs so artifacts are materialized and passed correctly across steps and runs.

4) Compiling the pipeline to YAML

To submit a pipeline to Kubeflow Pipelines you first compile the pipeline function into a YAML package. Use the KFP v2 compiler and produce a demo_pipeline.yaml package that contains component and deployment specs.
Run this from the terminal in the directory containing your pipeline script:
After running, demo_pipeline.yaml will be created. The YAML includes component definitions and the deployment spec Kubeflow uses to generate pod templates. Compilation may produce benign warnings about pip or environment; check exit status to ensure the compile succeeded.

5) Uploading the YAML to Kubeflow Pipelines

Open the Kubeflow Central Dashboard and navigate to Pipelines → Upload (or Create pipeline). Upload the generated demo_pipeline.yaml, provide a pipeline name and optionally a version/description, and save.
A screenshot of the Kubeflow Central Dashboard displaying a "demo_pipeline" graph with nodes labeled data-collection-process, feature-engineering, model-deployment, and model-training. The left sidebar shows navigation items (Home, Notebooks, Pipelines) and the top bar has buttons like Create run, Upload version, and Create experiment.
If you upload the YAML for the parallel pipeline, the UI shows nodes laid out side-by-side. If you upload the YAML with .after() ordering, the UI displays the DAG in sequence (vertical or directed order). To run the pipeline, choose Create run, select or create an Experiment to group runs, choose run type (one-off or recurring), and start the run.
A screenshot of the Kubeflow Central Dashboard showing the "New Pipeline" form for creating or uploading a pipeline version, with fields for pipeline name (demo_pipeline), version name, description, and options to upload a file or import by URL. The left sidebar shows navigation items like Home, Notebooks, TensorBoards, Volumes and Pipelines.

6) Observing pods and runtime behavior

When a pipeline run starts, Kubernetes creates pods for orchestration and for each component invocation. Typical pod roles you will encounter for a KFP run: To list pods across namespaces and check their status:
Inspect logs from a specific pod:
(Replace your-namespace and your-pod-name with values from kubectl get pod -A.) Implementation (impl) pods typically start, run the component, and then terminate. Executor and driver pods provide orchestration and helper logic; their logs are valuable when debugging pipeline failures. When a run finishes successfully, Kubeflow shows a green checkmark in the UI. From the run details you can access per-task logs and artifact locations, and you can also view the uploaded YAML that describes the pipeline spec.
Screenshot of the Kubeflow Central Dashboard showing a pipeline run ("Run of demo_pipeline-v2") with a vertical workflow graph of steps labeled data-collection-process, feature-engineering, model-training, and model-deployment. The left sidebar shows navigation items like Home, Notebooks, Pipelines, and Runs.

7) Inspecting the pipeline spec (YAML) in the UI

From a run details page you can open the pipeline spec to see the generated YAML that Kubeflow used to create pods and containers. The spec shows components, executor labels, and the runtime command used to invoke your component function. A truncated example:
The executor wraps a small bootstrap script to ensure pip and runtime dependencies are available, then invokes the specified component function. Check executor logs when components fail—these logs often surface install or import errors or Python exceptions from your component code.

8) Summary

  • Invoking components without explicit dependencies allows potential parallel execution — useful for independent tasks.
  • Use .after(previous_task) to enforce serial ordering when tasks depend on earlier steps.
  • Prefer typed Input/Output artifacts for passing data between steps; ordering alone does not transfer artifacts.
  • Compile your pipeline with the KFP v2 compiler to generate a YAML package for upload to Kubeflow Pipelines.
  • Use the Kubeflow Central Dashboard to upload pipeline versions, create runs, and inspect run details, logs, and specs.
  • Inspect Kubernetes pods (kubectl get pod -A and kubectl logs) to understand orchestration (DAG driver), container drivers, and impl pods that execute component code.
This completes wiring a minimal end-to-end pipeline: data collection → feature engineering → model training → model deployment.

Watch Video