Skip to main content
This lesson demonstrates how to execute a component (task) multiple times in parallel inside a Kubeflow Pipeline, collect outputs from each parallel run, and then select the best result. Use case summary:
  • Run a train_and_evaluate component for several learning_rate values in parallel.
  • Aggregate accuracies and learning rates from each parallel iteration.
  • Use a selector component (pick_best_learning_rate) to choose the learning rate that produced the highest accuracy.
Key ideas used:
  • ParallelFor to iterate and launch parallel runs.
  • Collected to gather outputs from parallel iterations into Python lists.
  • NamedTuple return type so component outputs are named and accessible by downstream steps.
Quick overview of tools in this example: Below is a compact, corrected, and working pipeline example that demonstrates this pattern.
How it works — step-by-step:
  1. Define train_and_evaluate to return a NamedTuple with ("accuracy", float) and ("learning_rate", float).
  2. In the pipeline, prepare the list of learning rates to try: [0.001, 0.01, 0.1].
  3. Use with ParallelFor(learning_rates) as lr: to execute train_and_evaluate(learning_rate=lr) for each value concurrently.
  4. Use Collected(run.outputs["accuracy"]) and Collected(run.outputs["learning_rate"]) to aggregate each output across all parallel iterations into lists.
  5. Call pick_best_learning_rate with the collected lists; it zips them, finds the maximum accuracy, and returns the corresponding learning rate.
  6. Compile the pipeline to a package (here par.yaml) and upload it to Kubeflow Pipelines.
When executed in the Kubeflow UI you will see a DAG representing each parallel iteration for each learning rate; each iteration runs train-and-evaluate with the corresponding learning_rate. After all iterations finish, the pick-best-learning-rate step runs using the collected lists.
A screenshot of the Kubeflow Central Dashboard showing a pipeline run titled "Run of parallel (90bee)" with a selected "train-and-evaluate" step displaying an input parameter "learning_rate" set to 0.001. The left sidebar shows navigation items like Home, Notebooks, TensorBoards, and Pipelines.
Example viewable outputs after a run (example random outputs from this dummy training):
Use Collected only on an output of a ParallelFor iteration (for example, run.outputs["accuracy"]). Collected aggregates that output over all iterations into a list that can be passed to downstream components.
This pattern is ideal for hyperparameter sweeps, ensemble experiments, or any workflow that runs the same step with different inputs and then aggregates and selects the best result. Links and references:

Watch Video