Skip to main content
Most machine learning work starts and often ends inside a notebook: train a model, check accuracy, maybe export a file — then consider the task done. Production machine learning is different. When a model must serve predictions to applications, handle traffic, scale automatically, and continuously improve, it becomes an infrastructure problem. That shift—from a single notebook to a live, reliable service—is the focus of this section. Operationalizing ML introduces new engineering questions:
  • How do we expose models reliably through APIs?
  • How do we scale inference for steady, bursty, or unpredictable traffic patterns?
  • How do we monitor and continuously improve model performance in production?
  • How do we automate experimentation (hyperparameter search) instead of manually tuning parameters?
Kubernetes provides the orchestration layer that addresses many of these questions. On top of Kubernetes, ML-focused systems handle model serving and autoscaling:
  • KServe simplifies deploying models as scalable, production-ready inference services on Kubernetes.
  • Knative provides serverless primitives for autoscaling and event-driven workloads.
An infographic slide titled "Production Machine Learning Requires New Systems" comparing naive model duplication (where copies fail randomly) on the left with Kubernetes-managed scaling on the right, showing multiple pods labeled Pod 1–4 all running. The image illustrates orchestration moving from fragile copies to a Kubernetes control plane managing healthy pods.
This section concentrates on two core, complementary systems within the modern MLOps stack:
  • KServe — a Kubernetes-native model serving platform that exposes models as scalable HTTP/gRPC inference APIs and integrates with autoscaling and Kubernetes networking.
  • Katib — an automated hyperparameter optimization framework that runs experiments to discover better-performing model configurations.
A slide titled "The Two Core Systems in This Section" showing two cards: KServe — a model serving platform for deploying models as scalable APIs with Kubernetes-native inference; and Katib — a hyperparameter optimization platform that automates experimentation to find better-performing models.
Together, KServe and Katib create a production-style ML workflow: automate experiments to find strong models, then serve the selected model reliably at scale. What to expect in this section:
  • A practical end-to-end example using the Iris dataset.
  • Training a Random Forest model and evaluating how hyperparameters affect performance.
  • Using Katib to automate hyperparameter search and run reproducible experiments.
  • Deploying the optimized model with KServe and exposing it as a live inference API on Kubernetes.
Quick comparison
Before you begin: ensure you have access to a Kubernetes cluster and a working kubectl context. Installing KServe and Katib requires cluster admin privileges. For guided installs and permissions, see the linked course pages above.
By the end of this section you will understand:
  • Why ML in production is primarily an infrastructure problem.
  • How Kubernetes enables reliable, scalable model serving.
  • How Katib automates experimentation so teams can find better model configurations faster.
  • How to combine these tools into repeatable, production-ready MLOps workflows.
Links and references

Watch Video