Skip to main content
The existing Random Forest training script embeds hyperparameter values directly in the source code:
Hard-coded hyperparameters make ad-hoc experiments simple, but they block automated hyperparameter optimization systems (for example, Katib) from exploring different configurations. Tuners like Katib require the training program to accept hyperparameter values at runtime so each trial can pass different values without changing the code. Making hyperparameters configurable enables seamless integration with hyperparameter tuning frameworks and supports reproducible, automated experiments.
Accepting hyperparameters via command-line arguments keeps defaults for manual runs while enabling external systems (for example, Katib) to override values for automated trials.

Make hyperparameters configurable with argparse

We add argument parsing so the script reads hyperparameter values from command-line options instead of relying on hard-coded constants. The example below shows a minimal train.py that accepts --n-estimators and --max-depth and uses them to configure a scikit-learn RandomForestRegressor.
Notes:
  • The script sets reasonable defaults so it remains convenient for local development.
  • The --max-depth argument accepts -1 to indicate unlimited depth (converted to None in the code).
  • Keep random_state for reproducibility across trials unless you intentionally want stochastic runs.

Quick comparison: hard-coded vs configurable

Example usage

Override defaults on the command line:
A hyperparameter tuning system like Katib launches multiple trials, each supplying different command-line arguments. Conceptually, trials might run:
Because the training program reads hyperparameters from command-line arguments, Katib (or other tuners) can fully automate the tuning process without editing source code between trials.
When enabling external systems to set hyperparameters, be careful to:
  • Validate inputs if unexpected values could break training.
  • Avoid exposing sensitive information via command-line arguments.
  • Keep reproducibility in mind by setting random_state where appropriate.

Integration tips for Katib and other tuners

  • Ensure the training container entrypoint accepts command-line flags (examples above).
  • Map Katib experiment parameters to the same flag names the script expects (e.g., --n-estimators).
  • Log metrics (for example, RMSE) to standard output or the framework-specific metrics endpoint so the tuner can read trial results.
  • Use sensible defaults to allow local debugging without the tuner.

Summary

  • Hard-coded hyperparameters prevent automated tuning systems from exploring the parameter space.
  • Adding argparse enables runtime configuration while preserving defaults for manual execution.
  • Pass parsed arguments directly into the model configuration: RandomForestRegressor(n_estimators=args.n_estimators, max_depth=args.max_depth).
  • This change makes the training program compatible with Katib and other hyperparameter optimization tools, enabling fully automated experiment workflows.

Watch Video