
Roadmap: Core Kubeflow Components
Notebooks
- Kubeflow Notebooks deliver an interactive development environment where data scientists and ML engineers explore data, prototype models, and run experiments.
- Rather than running Jupyter locally, Kubeflow launches notebooks inside Kubernetes for centralized infrastructure, consistent environments, and shared compute.
- Notebooks typically run JupyterLab, but can also include other IDEs such as code-server or RStudio when those tools are present in the container image.
Automation with Pipelines
- Manually running notebooks is not scalable for production workflows. Kubeflow Pipelines enables repeatable, automated workflows that run reliably on Kubernetes.
- We’ll decompose workflows into reusable components, pass parameters and artifacts between steps, and implement control flow so pipelines run on schedules or on demand.
- Kubeflow Pipelines is container-native and portable, designed for building scalable ML workflows that integrate with CI/CD and observability tooling.

Model Optimization with Katib
- Building a model is only part of the task — tuning hyperparameters is essential for peak performance. Katib automates hyperparameter search using strategies like grid search, random search, and Bayesian optimization.
- You’ll learn to define search spaces and configure experiments, and let Katib evaluate many configurations automatically, including early stopping and custom metrics.
- Katib is Kubeflow’s Kubernetes-native AutoML component for systematic hyperparameter tuning and model optimization.
Model Serving with KServe
- A trained model only creates value when it serves predictions. KServe provides Kubernetes-native model serving, exposing inference endpoints for real-time or batch predictions.
- We’ll package trained models, deploy them as scalable services, and cover operational concerns such as autoscaling, GPU-backed inference, canary rollouts, and traffic splitting.
- KServe integrates with monitoring and logging so you can track inference performance and observability.
Multi-user Collaboration and Operations
- Production platforms are shared: data scientists, ML engineers, and platform operators need isolated but integrated workspaces.
- Kubeflow uses Profiles, Kubernetes namespaces, and RBAC to give teams isolated environments while enabling centralized platform management.
- We’ll cover how to set up Profiles, control access, manage quotas, and share common resources so teams can work independently while contributing to a unified platform.
Course goal: gain hands-on experience across the full ML lifecycle—from experimentation and training to tuning, deployment, and multi-user operations—so you can design, operate, and maintain production-ready ML platforms with Kubeflow.
What you’ll be able to do after this course
- Design an end-to-end ML workflow that spans development, training, tuning, deployment, and operations.
- Run reproducible experiments in Kubeflow Notebooks and automate them with Kubeflow Pipelines.
- Optimize models with Katib and deploy production-grade inference with KServe.
- Configure Profiles, namespaces, and RBAC for safe multi-user collaboration and platform governance.
Links and references
- Kubeflow documentation
- Kubernetes documentation
- Jupyter Project
- JupyterLab
- Katib (Kubeflow AutoML)
- KServe (model serving)