Skip to main content
This guide shows how to run KServe locally on a Kind cluster and deploy a pre-trained scikit-learn Iris model. Step-by-step you’ll: confirm prerequisites, create a Kind cluster, install KServe, create an InferenceService, and test the model via kubectl port-forward.

Prerequisites (local)

Install the following CLI tools on your workstation: Quick checks:
If you run into permission errors with Docker on Linux, ensure your user is in the docker group or use sudo for commands that need elevated privileges.

Create a local Kind cluster

Create a cluster with a default configuration:
Example condensed output:
Verify the node(s):
Expected output:

Install KServe

KServe adds a CustomResourceDefinition (CRD) InferenceService and a controller that manages serving resources. Install the KServe release used in this demo:
Verify the CRD is installed:
Expected output:
A CRD extends Kubernetes with new resource types. Here, InferenceService lets KServe translate a high-level serving spec into Deployments, Services, and other Kubernetes resources for you.

Create a namespace for the demo

Use a dedicated namespace to keep demo resources isolated:

Deploy the Iris InferenceService

Create an iris.yaml with this InferenceService spec:
Apply it in the kserve-test namespace:
Verify the InferenceService:
Sample output:

KServe runtime and resources

KServe controller watches InferenceService CRs and creates the required resources (Deployments, Services, etc.). Check pods and services that KServe provisioned:
Example pods output:
Example services output:

Test the model (port-forward + curl)

Because ClusterIP services are only accessible inside the cluster, use kubectl port-forward to expose the predictor service on localhost for local testing. Start port-forwarding (run in one terminal):
You should see:
Prepare the input file iris-input.json with two instances:
From another terminal, call the model:
Expected response:
Interpretation: 1 typically corresponds to the Versicolor class (class mapping depends on the trained model; commonly 0 = Setosa, 1 = Versicolor, 2 = Virginica).
A screenshot of a browser-based Excalidraw diagram inside a code editor. The diagram shows iris-input.json being curl'ed to localhost:8000 which is port-forwarded into a Kind Kubernetes cluster running KServe with an inference service.

Summary of the request flow

  • Your curl targets localhost:8080, where kubectl port-forward exposes the predictor Service.
  • The ClusterIP Service routes the request to the predictor pod(s).
  • The KServe controller created those pods and the Service from the InferenceService spec.
  • The predictor runtime (scikit-learn server) returns the JSON predictions.

Next steps & tips

  • Expose an InferenceService externally for production using an Ingress, LoadBalancer, or NodePort depending on your cluster environment (this demo uses port-forward for local testing).
  • Inspect KServe controller logs if resources fail to come up:
    • kubectl logs -n kserve-system deploy/kserve-controller
  • Customize the InferenceService spec to set resources, replicas, or advanced runtime settings for production workloads.
  • For more details, refer to the KServe docs: https://kserve.github.io
Avoid exposing models directly to the public without proper authentication, rate limiting, and monitoring. Use ingress authentication or API gateways in production.
This completes the basic demo of deploying an Iris-trained scikit-learn model using KServe on a local Kind cluster.

Watch Video