Overview of Katib Experiments on Kubernetes for automated hyperparameter tuning, defining search space, objective, search algorithm, and trial templates to run and manage parallel trials
Katib Experiment is the Kubernetes-native resource for automated hyperparameter tuning. A single Experiment manifest declares the complete optimization workflow: which hyperparameters to search, which metric to optimize, which search algorithm to use, and how to run each trial. After you submit the manifest, Katib launches many trials across the cluster, collects metrics, and orchestrates the search automatically so you can find the best configuration without manual trial-and-error.
Core components of a Katib Experiment
Objective metric: the metric Katib optimizes (for example, minimize loss or maximize accuracy).
Search algorithm: the exploration strategy (e.g., Random Search, Grid Search, Bayesian Optimization).
Parameter (search) space: ranges and types for each hyperparameter.
Trial template: the Kubernetes workload template that runs each training trial and how sampled parameters are injected.
Component
Purpose
Example
Objective metric
Defines what to optimize and the optimization direction
maximize accuracy
Search algorithm
Strategy Katib uses to sample hyperparameters
random / bayesian
Parameter (search) space
Allowed values, ranges, and types for each parameter
n-estimators: 50-200
Trial template
Kubernetes manifest template that runs training jobs with injected parameters
Define the search space
The search space constrains which hyperparameter values Katib can try. You typically define numeric ranges, discrete choices, or categorical values. For example, when tuning a decision-tree-based model you might vary the number of estimators and maximum depth; Katib will sample combinations from those feasible ranges.
Trial template and parameter injection
Each trial is a standalone Kubernetes workload. The trial template shows how to run your training job and how sampled parameters are substituted into the container command or environment. Below is an example fragment of a trial template command where Katib injects sampled parameters:
Submitting and running an Experiment
Once your Experiment YAML is ready, submit it to the cluster:
kubectl apply -f experiment.yaml
Katib will create Trials, schedule training jobs, collect metrics, and continue the search until stopping criteria are met (e.g., max trials, max time, or converged objective).
Ensure you use the correct namespace for your Katib installation (commonly kubeflow — lowercase) when running kubectl commands.
Monitoring Katib
Use standard Kubernetes commands to inspect Experiments and watch Trials as they run:
kubectl get experiments -n kubeflowkubectl get trials -n kubeflow --watch
Benefits of using Katib
Katib automates what would otherwise be manual experimentation. Instead of single-threaded trial-and-error, you get:
Parallel trials across the cluster
Continuous metric aggregation and evaluation
Built-in search algorithms and pluggable strategies
Kubernetes-native scheduling and resource management
This Kubernetes-native approach to hyperparameter tuning is a central idea in modern MLOps: automate repetitive experimentation, scale with Kubernetes, and accelerate model development by letting Katib explore the configuration space for you.