> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Why Update the Random Forest Program

> Explains converting hard-coded Random Forest hyperparameters into command-line configurable arguments using argparse to enable automated tuning tools like Katib.

The existing Random Forest training script embeds hyperparameter values directly in the source code:

```python theme={null}
# original hard-coded hyperparameters
n_estimators = 100
max_depth = 5
```

Hard-coded hyperparameters make ad-hoc experiments simple, but they block automated hyperparameter optimization systems (for example, [Katib](https://www.kubeflow.org/docs/components/katib/)) from exploring different configurations. Tuners like Katib require the training program to accept hyperparameter values at runtime so each trial can pass different values without changing the code.

Making hyperparameters configurable enables seamless integration with hyperparameter tuning frameworks and supports reproducible, automated experiments.

<Callout icon="lightbulb" color="#1CB2FE">
  Accepting hyperparameters via command-line arguments keeps defaults for manual runs while enabling external systems (for example, Katib) to override values for automated trials.
</Callout>

## Make hyperparameters configurable with argparse

We add argument parsing so the script reads hyperparameter values from command-line options instead of relying on hard-coded constants. The example below shows a minimal `train.py` that accepts `--n-estimators` and `--max-depth` and uses them to configure a scikit-learn `RandomForestRegressor`.

```python theme={null}
# train.py
import argparse
from sklearn.datasets import fetch_california_housing
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestRegressor
from sklearn.metrics import mean_squared_error

def parse_args():
    parser = argparse.ArgumentParser(
        description="Train a RandomForest model with configurable hyperparameters."
    )
    parser.add_argument(
        "--n-estimators", type=int, default=100,
        help="Number of trees in the forest."
    )
    parser.add_argument(
        "--max-depth", type=int, default=5,
        help="Maximum depth of each tree (use -1 or None for unlimited depth)."
    )
    return parser.parse_args()

def main():
    args = parse_args()

    # Load example dataset (replace with your dataset)
    data = fetch_california_housing()
    X_train, X_test, y_train, y_test = train_test_split(
        data.data, data.target, test_size=0.2, random_state=42
    )

    # Configure model using parsed arguments
    max_depth_value = None if args.max_depth == -1 else args.max_depth
    model = RandomForestRegressor(
        n_estimators=args.n_estimators,
        max_depth=max_depth_value,
        random_state=42
    )
    model.fit(X_train, y_train)

    preds = model.predict(X_test)
    rmse = mean_squared_error(y_test, preds, squared=False)
    print(
        f"n_estimators={args.n_estimators}, "
        f"max_depth={max_depth_value}, RMSE={rmse:.4f}"
    )

if __name__ == "__main__":
    main()
```

Notes:

* The script sets reasonable defaults so it remains convenient for local development.
* The `--max-depth` argument accepts `-1` to indicate unlimited depth (converted to `None` in the code).
* Keep `random_state` for reproducibility across trials unless you intentionally want stochastic runs.

## Quick comparison: hard-coded vs configurable

| Aspect | Hard-coded | Configurable via `argparse` |
| - | -: | - |
| Flexibility | Low | High — values can be changed per run |
| Automation | Not compatible with tuners | Compatible with Katib and other tuners |
| Reproducibility | Deterministic but inflexible | Deterministic when `random_state` set; tuners can explore parameters |
| Example | `n_estimators = 100` | `python train.py --n-estimators 200 --max-depth 8` |

## Example usage

Override defaults on the command line:

```bash theme={null}
$ python train.py --n-estimators 200 --max-depth 8
```

A hyperparameter tuning system like [Katib](https://www.kubeflow.org/docs/components/katib/) launches multiple trials, each supplying different command-line arguments. Conceptually, trials might run:

```text theme={null}
Trial 1: --n-estimators 50  --max-depth 3
Trial 2: --n-estimators 100 --max-depth 5
Trial 3: --n-estimators 200 --max-depth 8
```

Because the training program reads hyperparameters from command-line arguments, Katib (or other tuners) can fully automate the tuning process without editing source code between trials.

<Callout icon="warning" color="#FF6B6B">
  When enabling external systems to set hyperparameters, be careful to:

  * Validate inputs if unexpected values could break training.
  * Avoid exposing sensitive information via command-line arguments.
  * Keep reproducibility in mind by setting `random_state` where appropriate.
</Callout>

## Integration tips for Katib and other tuners

* Ensure the training container entrypoint accepts command-line flags (examples above).
* Map Katib experiment parameters to the same flag names the script expects (e.g., `--n-estimators`).
* Log metrics (for example, RMSE) to standard output or the framework-specific metrics endpoint so the tuner can read trial results.
* Use sensible defaults to allow local debugging without the tuner.

## Summary

* Hard-coded hyperparameters prevent automated tuning systems from exploring the parameter space.
* Adding `argparse` enables runtime configuration while preserving defaults for manual execution.
* Pass parsed arguments directly into the model configuration: `RandomForestRegressor(n_estimators=args.n_estimators, max_depth=args.max_depth)`.
* This change makes the training program compatible with Katib and other hyperparameter optimization tools, enabling fully automated experiment workflows.

## Links and references

* [Katib — Kubeflow Hyperparameter Tuning](https://www.kubeflow.org/docs/components/katib/)
* [argparse — Python documentation](https://docs.python.org/3/library/argparse.html)
* [scikit-learn RandomForestRegressor](https://scikit-learn.org/stable/modules/generated/sklearn.ensemble.RandomForestRegressor.html)

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/kubeflow/module/d9b1b119-0c6f-494b-b063-8eccd99dbff7/lesson/c857ed8d-7a9b-4972-bbf4-dc20d7c2d0ea" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.