> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Section Introduction

> Explains Kubeflow Pipelines and Notebooks for building, orchestrating, and reproducing end-to-end containerized ML workflows on Kubernetes.

Building machine learning solutions is more than just training a model. A complete ML workflow includes data collection and preprocessing, feature engineering, training, evaluation, and deployment. When these steps are executed manually, teams face reproducibility gaps, poor scaling, and inefficient collaboration.

ML pipelines solve these problems by automating end-to-end workflows, making components reusable, and providing reliable orchestration on Kubernetes. Kubeflow Pipelines (KFP) is a core Kubeflow component designed to build, run, and manage these ML pipelines.

KFP lets you represent workflows as pipelines composed of connected steps. Each step runs inside its own container, which makes components portable, language-agnostic, and reusable. KFP also provides orchestration features — scheduling, execution ordering, retries, and dependency management — so you can run workflows reliably at scale.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/MGkgrGfKHDtoCnUb/images/Kubeflow/Working-With-Kubeflow/Section-Introduction/kubeflow-vs-traditional-ml-infographic.jpg?fit=max&auto=format&n=MGkgrGfKHDtoCnUb&q=85&s=424e89a846c57b92e91c70cc400c6e17" alt="An infographic comparing Traditional ML (left) vs a Kubeflow Pipeline (right), showing Data Prep, Training, Evaluation and Deployment as manual on the left and automated (checked) on the right. The bottom contrasts drawbacks for traditional workflows (hard to reproduce, slow to scale, no tracking) with benefits of Kubeflow (reproducible, scalable, tracked)." width="1920" height="1080" data-path="images/Kubeflow/Working-With-Kubeflow/Section-Introduction/kubeflow-vs-traditional-ml-infographic.jpg" />
</Frame>

Why pipelines matter

* Reproducibility: Pipeline definitions and containerized components make runs repeatable.
* Portability: Containerized steps are language-agnostic and run anywhere Kubernetes is available.
* Observability: KFP integrates experiment tracking and artifact metadata so you can compare runs.
* Scalability: Orchestration and scheduling let you scale workloads across a cluster.

Key concepts

* Pipelines: Directed acyclic graphs (DAGs) that define an ML workflow.
* Components: Independent containers that perform discrete tasks (data prep, training, etc.).
* Artifacts & Parameters: Mechanisms for passing data, models, and configuration between components.
* Orchestration: Scheduling, retries, and dependency management handled by KFP.

Pipelines are represented as directed acyclic graphs (DAGs): components form nodes and edges represent data or control dependencies. Components commonly exchange parameters, artifacts, and metadata so later steps can build on earlier results. KFP also integrates experiment and artifact tracking so runs can be compared and reproduced.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/MGkgrGfKHDtoCnUb/images/Kubeflow/Working-With-Kubeflow/Section-Introduction/kubeflow-pipelines-steps-features.jpg?fit=max&auto=format&n=MGkgrGfKHDtoCnUb&q=85&s=3d97766687e8efdf544a35a8ba304440" alt="An infographic titled &#x22;Understanding Kubeflow Pipelines&#x22; showing a KFP pipeline with five containerized steps: Data Prep, Train, Evaluate, Register, and Deploy. Below it are colored blocks listing features KFP automatically manages, such as orchestration, scheduling, retries, and dependency management." width="1920" height="1080" data-path="images/Kubeflow/Working-With-Kubeflow/Section-Introduction/kubeflow-pipelines-steps-features.jpg" />
</Frame>

Typical progression when adopting Kubeflow Pipelines

* Create your first pipeline to learn structure and component composition.
* Add pipeline parameters to make workflows configurable and reusable.
* Pass artifacts and data between components so steps can share outputs.
* Implement control flow (conditionals, branching, and dependencies) to enable flexible execution.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/MGkgrGfKHDtoCnUb/images/Kubeflow/Working-With-Kubeflow/Section-Introduction/building-dynamic-ml-workflows-4-panels.jpg?fit=max&auto=format&n=MGkgrGfKHDtoCnUb&q=85&s=91ba4d19db6e363137351830fedf654d" alt="A slide titled &#x22;Building Dynamic ML Workflows&#x22; showing four colored, numbered panels connected by arrows. The panels read: 01 First Pipeline (End-to-end structure), 02 Parameters (Configurable inputs), 03 Data Passing (Shared outputs), and 04 Control Flow (Conditional logic)." width="1920" height="1080" data-path="images/Kubeflow/Working-With-Kubeflow/Section-Introduction/building-dynamic-ml-workflows-4-panels.jpg" />
</Frame>

Prototyping and development workflow
Teams commonly prototype pipelines inside interactive notebooks before packaging components. Kubeflow Notebooks provide notebook servers (for example, Jupyter or VS Code) running inside Kubernetes so data scientists and ML engineers can write pipeline code, test components, explore datasets, and iterate on models — all integrated with Kubeflow for a smooth path from experimentation to scalable execution.

<Callout icon="lightbulb" color="#1CB2FE">
  Use notebooks to prototype components and debug data flows quickly. Once a component works, package it as a containerized KFP component for reuse and CI/CD integration.
</Callout>

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/MGkgrGfKHDtoCnUb/images/Kubeflow/Working-With-Kubeflow/Section-Introduction/kubeflow-notebooks-pipelines-kubernetes-jupyter-vscode.jpg?fit=max&auto=format&n=MGkgrGfKHDtoCnUb&q=85&s=13ad13634614762236bdd926c0dd659b" alt="A slide diagram titled &#x22;Developing Pipelines With Kubeflow Notebooks&#x22; showing a Kubernetes cluster containing a &#x22;Notebooks&#x22; panel with Jupyter and Visual Studio Code logos and tasks like &#x22;Write Code,&#x22; &#x22;Experiment,&#x22; &#x22;Explore Data,&#x22; and &#x22;Test Components.&#x22; The image is branded © Copyright KodeKloud." width="1920" height="1080" data-path="images/Kubeflow/Working-With-Kubeflow/Section-Introduction/kubeflow-notebooks-pipelines-kubernetes-jupyter-vscode.jpg" />
</Frame>

Quick reference: pipeline stages and examples

| Stage | Purpose | Example |
| - | - | - |
| Data prep | Clean and transform raw data | `python preprocess.py --input gs://bucket/raw --output /data/clean` |
| Training | Fit model using prepared data | `python train.py --data /data/clean --epochs 50` |
| Evaluation | Validate model performance | `python eval.py --model /models/latest --metrics_out /metrics` |
| Registration | Store model artifact and metadata | `kubectl apply -f model-registry.yaml` |
| Deployment | Serve model in production | `kubectl expose deployment model-server --port=80` |

Getting started tips

* Start small: build a simple end-to-end pipeline (preprocess → train → evaluate) to learn component composition.
* Parameterize: expose key configuration as pipeline parameters to reuse pipelines across datasets and experiments.
* Capture artifacts: ensure outputs (models, metrics) are stored in a persistable artifact store for lineage and reproducibility.
* CI/CD: containerize components early so they can be tested and promoted through CI/CD pipelines.

Links and references

* [Kubeflow Pipelines documentation](https://www.kubeflow.org/docs/components/pipelines/)
* [Kubeflow Notebooks documentation](https://www.kubeflow.org/docs/components/notebooks/)
* [Kubernetes documentation](https://kubernetes.io)
* [Jupyter](https://jupyter.org) and [Visual Studio Code](https://code.visualstudio.com)

Together, Kubeflow Notebooks and Kubeflow Pipelines provide an integrated development-to-production path for reproducible, scalable ML workflows on Kubernetes.

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/kubeflow/module/bece8da9-953e-480e-8774-b25b66c3830f/lesson/ec81aa76-eb38-4d5d-9fbc-c4e7877589bd" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.