Skip to main content
This lesson covers control flow patterns in Kubeflow Pipelines (KFP): conditional execution (If / Else) and parallel execution with result aggregation (ParallelFor + Collected). These primitives let you build robust ML pipelines that make decisions (promote or notify) and scale tasks (hyperparameter sweeps) while keeping downstream logic simple.

Conditional execution (If / Else)

Use conditional blocks when some pipeline stages should only run when a condition is met. A typical example: after training and evaluation, register and deploy a model only if its evaluation accuracy meets a threshold; otherwise, notify the team of failure. The example below demonstrates this pattern. It assumes the components train_model, evaluate_model, register_model, deploy_model, and notify_failure are available in your environment.
How it works:
  • train_model() runs first, producing a model output.
  • evaluate_model() consumes that model and produces an accuracy output.
  • The If block checks eval_task.outputs["accuracy"] against accuracy_threshold. If true, register_model and deploy_model run. If false, the Else block runs notify_failure.
Make sure your components expose the output names you reference (for example: "model", "accuracy", "registered_model"). The If / Else control flow evaluates expressions based on component outputs.

Parallel runs and collecting results (ParallelFor / Collected)

ParallelFor runs the same component multiple times (one iteration per input value) and Collected aggregates simple outputs (scalars, small JSON-serializable values) across those parallel iterations. This pattern is ideal for hyperparameter sweeps and parameter search. The diagram below illustrates running the “Train Model” component three times with different learning rates, then passing results to a “Select Best Model” component.
A slide titled "ParallelFor/Collected" showing three colored rounded boxes labeled "Train Model" with learning rates lr=0.1, lr=0.01, and lr=0.001, and a gray rounded box below labeled "Select Best Model." The slide has a © KodeKloud mark in the corner.
Example: run train_and_evaluate in parallel for multiple learning rates, collect the outputs, and select the best learning rate.
Key points:
  • ParallelFor(learning_rates) launches one task for each learning rate in parallel.
  • Each iteration returns a task reference (here run), representing the per-iteration outputs.
  • Collected(run.outputs["accuracy"]) and Collected(run.outputs["learning_rate"]) gather scalar outputs across iterations into lists for downstream consumption by pick_best_learning_rate.
Use Collected() only for small, JSON-serializable outputs (metrics, scalars, short strings). Do not use Collected() for large artifacts (models, large files). For artifacts, store them in artifact storage (e.g., MinIO, GCS) and pass references instead.

Comparison: If / Else vs ParallelFor / Collected

Common pitfalls and tips

  • Ensure component outputs are named and typed consistently; mismatched names are a common source of runtime errors.
  • Keep Collected payloads small to avoid serialization/memory issues.
  • When running many parallel tasks, watch cluster resource limits (CPU, memory, GPU) and set concurrency or resource requests/limits appropriately.
  • For reproducible experiments, fix random seeds or use deterministic training where possible.
Use these control-flow patterns to make your ML pipelines more modular, scalable, and production-ready.

Watch Video