> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Demo Creating Our First Pipeline

> Guide to building, compiling, uploading, and running a simple Kubeflow Pipeline with KFP v2 DSL, demonstrating parallel versus serial task execution and inspecting runs and pods

This guide shows how to wire individual components into a complete Kubeflow Pipeline using the KFP v2 DSL. You'll learn how component calls without explicit dependencies run in parallel, how to enforce serial execution, how to compile the pipeline into a YAML package, and how to run and inspect the pipeline in Kubeflow Pipelines.

In this walkthrough we use simple placeholder components that print messages. In production, components will perform data I/O, return typed artifacts or primitive outputs, and pass artifacts between steps.

What you'll do:

* See how calling components without dependencies results in parallel execution
* Enforce serial ordering so components run one after another
* Compile the pipeline Python definition to a YAML package
* Upload and run the pipeline in Kubeflow and inspect pods and logs

## 1) Components: minimal definitions

Below is a minimal, correct set of component definitions using the KFP v2 DSL. These example components return `None` for simplicity. In real pipelines, prefer typed `Output` and `Input` artifacts to pass data between components.

```python theme={null}
# python
from kfp.v2 import dsl

@dsl.component
def data_collection_processing() -> None:
    print("collecting and processing data")

@dsl.component
def feature_engineering() -> None:
    print("feature engineering")

@dsl.component
def model_training() -> None:
    print("Training Model")

@dsl.component
def model_deployment() -> None:
    print("deploying Model")
```

Tip: Define inputs/outputs for components when you need to pass artifacts (datasets, models, metrics) between steps. Relying only on execution ordering does not transfer data.

## 2) Calling components without dependencies (parallel execution)

If you call component functions inside a pipeline without specifying any ordering, KFP treats each call as an independent task. Without dependencies, the scheduler may run tasks concurrently (parallel execution). This is useful for independent tasks that can run simultaneously.

```python theme={null}
# python
from kfp.v2 import dsl

@dsl.pipeline(name="demo-pipeline-parallel")
def demo_pipeline_parallel() -> None:
    dc = data_collection_processing()
    fe = feature_engineering()
    mt = model_training()
    md = model_deployment()
```

When visualized in the Kubeflow UI, these steps appear side-by-side in the graph, indicating potential parallel runs. Use this pattern only when tasks do not rely on each other's results.

## 3) Specifying ordering (run in series)

To force tasks to run in a specific sequence, use the `.after()` method on the task object returned by a component invocation. `.after()` guarantees a task does not start until the specified predecessor completes.

```python theme={null}
# python
from kfp.v2 import dsl

@dsl.pipeline(name="demo-pipeline")
def demo_pipeline() -> None:
    dc = data_collection_processing()
    fe = feature_engineering().after(dc)
    mt = model_training().after(fe)
    md = model_deployment().after(mt)
```

This enforces the following execution order:

1. `data_collection_processing`
2. `feature_engineering` (after `data_collection_processing`)
3. `model_training` (after `feature_engineering`)
4. `model_deployment` (after `model_training`)

<Callout icon="lightbulb" color="#1CB2FE">
  Use `.after()` when the start of a task depends on the completion of a previous task. For data transfer between steps, prefer explicit typed `Output` and `Input` artifact parameters; ordering alone does not pass artifacts.
</Callout>

<Callout icon="warning" color="#FF6B6B">
  Do not rely solely on `.after()` to move data between components. Define and use typed inputs/outputs so artifacts are materialized and passed correctly across steps and runs.
</Callout>

## 4) Compiling the pipeline to YAML

To submit a pipeline to Kubeflow Pipelines you first compile the pipeline function into a YAML package. Use the KFP v2 compiler and produce a `demo_pipeline.yaml` package that contains component and deployment specs.

```python theme={null}
# python
from kfp.v2 import compiler

if __name__ == "__main__":
    compiler.Compiler().compile(
        pipeline_func=demo_pipeline,
        package_path="demo_pipeline.yaml"
    )
```

Run this from the terminal in the directory containing your pipeline script:

```bash theme={null}
# bash
pip install kfp
python pipeline.py
```

After running, `demo_pipeline.yaml` will be created. The YAML includes component definitions and the deployment spec Kubeflow uses to generate pod templates. Compilation may produce benign warnings about pip or environment; check exit status to ensure the compile succeeded.

## 5) Uploading the YAML to Kubeflow Pipelines

Open the Kubeflow Central Dashboard and navigate to Pipelines → Upload (or Create pipeline). Upload the generated `demo_pipeline.yaml`, provide a pipeline name and optionally a version/description, and save.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/MGkgrGfKHDtoCnUb/images/Kubeflow/Working-With-Kubeflow/Demo-Creating-Our-First-Pipeline/kubeflow-central-dashboard-demo-pipeline-graph.jpg?fit=max&auto=format&n=MGkgrGfKHDtoCnUb&q=85&s=f56e33b7015f7f7b9819c52344a1ef21" alt="A screenshot of the Kubeflow Central Dashboard displaying a &#x22;demo_pipeline&#x22; graph with nodes labeled data-collection-process, feature-engineering, model-deployment, and model-training. The left sidebar shows navigation items (Home, Notebooks, Pipelines) and the top bar has buttons like Create run, Upload version, and Create experiment." width="1920" height="1080" data-path="images/Kubeflow/Working-With-Kubeflow/Demo-Creating-Our-First-Pipeline/kubeflow-central-dashboard-demo-pipeline-graph.jpg" />
</Frame>

If you upload the YAML for the parallel pipeline, the UI shows nodes laid out side-by-side. If you upload the YAML with `.after()` ordering, the UI displays the DAG in sequence (vertical or directed order). To run the pipeline, choose Create run, select or create an Experiment to group runs, choose run type (one-off or recurring), and start the run.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/MGkgrGfKHDtoCnUb/images/Kubeflow/Working-With-Kubeflow/Demo-Creating-Our-First-Pipeline/kubeflow-central-dashboard-new-pipeline.jpg?fit=max&auto=format&n=MGkgrGfKHDtoCnUb&q=85&s=ab0131d0ceab4383a74644f29e7eba63" alt="A screenshot of the Kubeflow Central Dashboard showing the &#x22;New Pipeline&#x22; form for creating or uploading a pipeline version, with fields for pipeline name (demo_pipeline), version name, description, and options to upload a file or import by URL. The left sidebar shows navigation items like Home, Notebooks, TensorBoards, Volumes and Pipelines." width="1920" height="1080" data-path="images/Kubeflow/Working-With-Kubeflow/Demo-Creating-Our-First-Pipeline/kubeflow-central-dashboard-new-pipeline.jpg" />
</Frame>

## 6) Observing pods and runtime behavior

When a pipeline run starts, Kubernetes creates pods for orchestration and for each component invocation. Typical pod roles you will encounter for a KFP run:

| Pod Role | Purpose | Notes / Example |
| - | - | - |
| DAG driver | Orchestrates pipeline DAG and coordinates task scheduling | Long-lived for the run (`system-dag-driver`) |
| Container driver | Manages lifecycle and containerized executor processes | `system-container-driver` commonly appears |
| Implementation pod (impl) | Runs the component container image and executes component code | One impl pod per component invocation |

To list pods across namespaces and check their status:

```bash theme={null}
# bash
kubectl get pod -A
```

Inspect logs from a specific pod:

```bash theme={null}
# bash
kubectl logs -n `your-namespace` `your-pod-name`
```

(Replace `your-namespace` and `your-pod-name` with values from `kubectl get pod -A`.)

Implementation (impl) pods typically start, run the component, and then terminate. Executor and driver pods provide orchestration and helper logic; their logs are valuable when debugging pipeline failures.

When a run finishes successfully, Kubeflow shows a green checkmark in the UI. From the run details you can access per-task logs and artifact locations, and you can also view the uploaded YAML that describes the pipeline spec.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/MGkgrGfKHDtoCnUb/images/Kubeflow/Working-With-Kubeflow/Demo-Creating-Our-First-Pipeline/kubeflow-pipeline-run-vertical-workflow.jpg?fit=max&auto=format&n=MGkgrGfKHDtoCnUb&q=85&s=57425eed4baa694ba6f14b539b25a943" alt="Screenshot of the Kubeflow Central Dashboard showing a pipeline run (&#x22;Run of demo_pipeline-v2&#x22;) with a vertical workflow graph of steps labeled data-collection-process, feature-engineering, model-training, and model-deployment. The left sidebar shows navigation items like Home, Notebooks, Pipelines, and Runs." width="1920" height="1080" data-path="images/Kubeflow/Working-With-Kubeflow/Demo-Creating-Our-First-Pipeline/kubeflow-pipeline-run-vertical-workflow.jpg" />
</Frame>

## 7) Inspecting the pipeline spec (YAML) in the UI

From a run details page you can open the pipeline spec to see the generated YAML that Kubeflow used to create pods and containers. The spec shows components, executor labels, and the runtime command used to invoke your component function. A truncated example:

```yaml theme={null}
# yaml
components:
  comp-data-collection-processing:
    executorLabel: exec-data-collection-processing
  comp-feature-engineering:
    executorLabel: exec-feature-engineering
deploymentSpec:
  executors:
    exec-data-collection-processing:
      container:
        args:
        - '--executor_input'
        - '{{}}'
        - '--function_to_execute'
        - data_collection_processing
        command:
        - sh
        - '-c'
        - |
          if ! [ -x "$(command -v pip)" ]; then
            python3 -m ensurepip || python3 -m ensurepip --user || apt-get install python3-pip
          fi

          PIP_DISABLE_PIP_VERSION_CHECK=1 python3 -m pip install --quiet \
            --no-warn-script-location 'kfp==2.15.2' --no-deps \
            'typing-extensions>=3.7.4,<5; python_version<"3.9"' && "$@"
```

The executor wraps a small bootstrap script to ensure `pip` and runtime dependencies are available, then invokes the specified component function. Check executor logs when components fail—these logs often surface install or import errors or Python exceptions from your component code.

## 8) Summary

* Invoking components without explicit dependencies allows potential parallel execution — useful for independent tasks.
* Use `.after(previous_task)` to enforce serial ordering when tasks depend on earlier steps.
* Prefer typed `Input`/`Output` artifacts for passing data between steps; ordering alone does not transfer artifacts.
* Compile your pipeline with the KFP v2 compiler to generate a YAML package for upload to Kubeflow Pipelines.
* Use the Kubeflow Central Dashboard to upload pipeline versions, create runs, and inspect run details, logs, and specs.
* Inspect Kubernetes pods (`kubectl get pod -A` and `kubectl logs`) to understand orchestration (DAG driver), container drivers, and impl pods that execute component code.

This completes wiring a minimal end-to-end pipeline: data collection → feature engineering → model training → model deployment.

## Links and references

* KFP v2 DSL and compiler docs: [https://www.kubeflow.org/docs/components/pipelines/sdk/v2/](https://www.kubeflow.org/docs/components/pipelines/sdk/v2/)
* Kubeflow Central Dashboard: [https://www.kubeflow.org/docs/components/central-dash/](https://www.kubeflow.org/docs/components/central-dash/)
* Kubernetes `kubectl` reference: [https://kubernetes.io/docs/reference/kubectl/overview/](https://kubernetes.io/docs/reference/kubectl/overview/)

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/kubeflow/module/bece8da9-953e-480e-8774-b25b66c3830f/lesson/cb229f7b-8a55-4c66-81d9-1a6c3494ab11" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.