Conditional execution (If / Else)
Use conditional blocks when some pipeline stages should only run when a condition is met. A typical example: after training and evaluation, register and deploy a model only if its evaluation accuracy meets a threshold; otherwise, notify the team of failure. The example below demonstrates this pattern. It assumes the componentstrain_model, evaluate_model, register_model, deploy_model, and notify_failure are available in your environment.
train_model()runs first, producing amodeloutput.evaluate_model()consumes that model and produces anaccuracyoutput.- The
Ifblock checkseval_task.outputs["accuracy"]againstaccuracy_threshold. If true,register_modelanddeploy_modelrun. If false, theElseblock runsnotify_failure.
Make sure your components expose the output names you reference (for example:
"model", "accuracy", "registered_model"). The If / Else control flow evaluates expressions based on component outputs.Parallel runs and collecting results (ParallelFor / Collected)
ParallelFor runs the same component multiple times (one iteration per input value) and Collected aggregates simple outputs (scalars, small JSON-serializable values) across those parallel iterations. This pattern is ideal for hyperparameter sweeps and parameter search. The diagram below illustrates running the “Train Model” component three times with different learning rates, then passing results to a “Select Best Model” component.
train_and_evaluate in parallel for multiple learning rates, collect the outputs, and select the best learning rate.
ParallelFor(learning_rates)launches one task for each learning rate in parallel.- Each iteration returns a task reference (here
run), representing the per-iteration outputs. Collected(run.outputs["accuracy"])andCollected(run.outputs["learning_rate"])gather scalar outputs across iterations into lists for downstream consumption bypick_best_learning_rate.
Use
Collected() only for small, JSON-serializable outputs (metrics, scalars, short strings). Do not use Collected() for large artifacts (models, large files). For artifacts, store them in artifact storage (e.g., MinIO, GCS) and pass references instead.Comparison: If / Else vs ParallelFor / Collected
Common pitfalls and tips
- Ensure component outputs are named and typed consistently; mismatched names are a common source of runtime errors.
- Keep Collected payloads small to avoid serialization/memory issues.
- When running many parallel tasks, watch cluster resource limits (CPU, memory, GPU) and set concurrency or resource requests/limits appropriately.
- For reproducible experiments, fix random seeds or use deterministic training where possible.
Links and references
- Kubeflow Pipelines SDK documentation: https://kubeflow-pipelines.readthedocs.io/
- Kubeflow official docs: https://www.kubeflow.org/docs/
- KFP control-flow examples and patterns: https://github.com/kubeflow/pipelines/tree/master/samples