> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# The Iris Dataset

> Overview of the Iris dataset, its four numeric features and three species classification task, with a scikit-learn example for loading, splitting, preprocessing, training, and evaluation.

The Iris dataset is one of the most iconic datasets in machine learning. Introduced by statistician Ronald Fisher in 1936, it remains a standard teaching and benchmarking dataset for classification tasks. Because it is small, clean, and easy to understand, the Iris dataset is often called the "hello world" of machine learning.

Throughout this article we use the Iris dataset to demonstrate core supervised-learning concepts: data inspection, feature understanding, train/test splitting, preprocessing, training, and evaluation.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/MGkgrGfKHDtoCnUb/images/Kubeflow/KServe-and-Katib/The-Iris-Dataset/iris-dataset-hello-world-classification.jpg?fit=max&auto=format&n=MGkgrGfKHDtoCnUb&q=85&s=0c57aac11fcb97ee8da655b0d0148e13" alt="A slide titled &#x22;What is the Iris Dataset?&#x22; with five panels summarizing it as a famous dataset (introduced in 1936) that contains iris flower measurements, is used for classification, and is often called the &#x22;Hello World&#x22; of machine learning." width="1920" height="1080" data-path="images/Kubeflow/KServe-and-Katib/The-Iris-Dataset/iris-dataset-hello-world-classification.jpg" />
</Frame>

## Problem type: classification

The Iris dataset is a multiclass classification problem: given numerical measurements of a flower, predict its species. A model learns patterns from the numeric input features and outputs one of the target species. This is a canonical example for teaching supervised classification pipelines.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/MGkgrGfKHDtoCnUb/images/Kubeflow/KServe-and-Katib/The-Iris-Dataset/iris-dataset-classification-slide.jpg?fit=max&auto=format&n=MGkgrGfKHDtoCnUb&q=85&s=75cc637dac616d250e6576a3e75462f5" alt="A presentation slide titled &#x22;What Problem Does the Iris Dataset Solve?&#x22; with a question-mark graphic and four colored points. It explains that Iris is a classification task to predict flower species using flower measurements and that the model learns patterns from the data." width="1920" height="1080" data-path="images/Kubeflow/KServe-and-Katib/The-Iris-Dataset/iris-dataset-classification-slide.jpg" />
</Frame>

## Target classes: three species

There are three labeled species in the dataset:

* Iris setosa
* Iris versicolor
* Iris virginica

These species serve as the target labels the model must predict. Their measurements differ enough for many algorithms to separate them effectively, but some class pairs are more challenging than others.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/MGkgrGfKHDtoCnUb/images/Kubeflow/KServe-and-Katib/The-Iris-Dataset/iris-three-species-petals-sepals.jpg?fit=max&auto=format&n=MGkgrGfKHDtoCnUb&q=85&s=95e8a2d484a9697833f58855c2f9fdef" alt="A slide titled &#x22;Iris Flower Species&#x22; showing photos of three iris types (Iris setosa, Iris versicolor, Iris virginica) with arrows labeling their petals and sepals." width="1920" height="1080" data-path="images/Kubeflow/KServe-and-Katib/The-Iris-Dataset/iris-three-species-petals-sepals.jpg" />
</Frame>

## Features: four numeric measurements

Each observation contains four numeric features:

* sepal length (cm)
* sepal width (cm)
* petal length (cm)
* petal width (cm)

Features are the inputs the model uses to learn the mapping to species labels. Understanding which features are present and their scales helps decide preprocessing steps like scaling or feature selection.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/MGkgrGfKHDtoCnUb/images/Kubeflow/KServe-and-Katib/The-Iris-Dataset/dataset-features-flower-measurements-table.jpg?fit=max&auto=format&n=MGkgrGfKHDtoCnUb&q=85&s=fb224943ba68ac96c7d41c8823589e7b" alt="A presentation slide titled &#x22;Features in the Dataset&#x22; showing a table of four flower measurements—sepal length, sepal width, petal length, and petal width—with brief descriptions for each." width="1920" height="1080" data-path="images/Kubeflow/KServe-and-Katib/The-Iris-Dataset/dataset-features-flower-measurements-table.jpg" />
</Frame>

## Why the Iris dataset is so popular

The Iris dataset is widely used because it strikes a balance between simplicity and pedagogical value:

* Easy to understand and visualize
* Clean and balanced classes
* Small enough to train quickly
* Great for demonstrating full ML workflows (inspect → split → preprocess → train → evaluate)

These properties make it ideal for tutorials, demos, and exploratory model experiments.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/MGkgrGfKHDtoCnUb/images/Kubeflow/KServe-and-Katib/The-Iris-Dataset/iris-dataset-popular-reasons.jpg?fit=max&auto=format&n=MGkgrGfKHDtoCnUb&q=85&s=b46c92cbeedc0d99bf9fe5b7a940fc0a" alt="A slide titled &#x22;Why is the Iris Dataset So Popular?&#x22; showing five panels that list reasons: Easy to Understand, Clean & Balanced, Trains Quickly, Beginner Friendly, and Demonstrates Workflows. Each panel has an icon and a short explanatory phrase." width="1920" height="1080" data-path="images/Kubeflow/KServe-and-Katib/The-Iris-Dataset/iris-dataset-popular-reasons.jpg" />
</Frame>

## Where to get the dataset

The Iris dataset is bundled with several ML libraries and is also available from classic repositories:

* scikit-learn: load with a single function call
* UCI Machine Learning Repository: canonical dataset source
* Many tutorials and teaching resources include ready-to-use copies

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/MGkgrGfKHDtoCnUb/images/Kubeflow/KServe-and-Katib/The-Iris-Dataset/iris-dataset-sources-scikit-uci-tutorials.jpg?fit=max&auto=format&n=MGkgrGfKHDtoCnUb&q=85&s=8bc1b852a878f07c324aaf1fdea2706b" alt="A presentation slide titled &#x22;Where Can We Find the Iris Dataset?&#x22; showing three numbered colored panels listing sources: Scikit-learn, the UCI Repository, and Tutorials & Demos. Each panel gives a brief note (e.g., &#x22;Available directly in Scikit-learn&#x22;)." width="1920" height="1080" data-path="images/Kubeflow/KServe-and-Katib/The-Iris-Dataset/iris-dataset-sources-scikit-uci-tutorials.jpg" />
</Frame>

## Quick dataset facts

| Property | Value |
| - | - |
| Observations | 150 |
| Features | 4 (numeric) |
| Classes | 3 (Setosa, Versicolor, Virginica) |
| Typical use | Supervised multiclass classification |
| Common sources | [scikit-learn](https://scikit-learn.org/stable/modules/generated/sklearn.datasets.load_iris.html), [UCI Repository](https://archive.ics.uci.edu/ml/datasets/iris) |

<Callout icon="lightbulb" color="#1CB2FE">
  A typical workflow is: load the dataset, inspect features and labels, split into training and test sets (use `stratify` to preserve class balance), apply preprocessing (e.g., scaling), train a model, and evaluate on held-out data.
</Callout>

## Example: load, inspect, split, and train

Below is a compact scikit-learn example that demonstrates the essential steps. Comments explain each step. Run this in a Python environment with scikit-learn and pandas installed.

```python theme={null}
# Load the dataset and inspect it
from sklearn.datasets import load_iris

iris = load_iris(as_frame=True)  # returns a Bunch with a pandas DataFrame
X = iris.data
y = iris.target

print("Feature names:", iris.feature_names)
print("Target names:", list(iris.target_names))
print("Data shape:", X.shape)
```

Expected console output (example):

```text theme={null}
Feature names: ['sepal length (cm)', 'sepal width (cm)', 'petal length (cm)', 'petal width (cm)']
Target names: ['setosa', 'versicolor', 'virginica']
Data shape: (150, 4)
```

Next, split the data into training and test sets. Use `stratify=y` to keep class proportions similar between splits.

```python theme={null}
from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42, stratify=y
)
```

Train a simple pipeline with standard scaling and Logistic Regression:

```python theme={null}
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression

clf = make_pipeline(
    StandardScaler(),
    LogisticRegression(max_iter=200, random_state=42)
)
clf.fit(X_train, y_train)

test_accuracy = clf.score(X_test, y_test)
print("Test accuracy:", round(test_accuracy, 3))
```

On the Iris dataset, a basic pipeline like this often yields high accuracy (commonly > 0.90) because the classes are well-separated by these features.

## Summary

This article covered:

* What the Iris dataset is and why it's important for ML education
* The classification problem and the three species involved
* The four numeric features used by models
* Where to find the dataset and a quick scikit-learn example showing loading, splitting, preprocessing, and training

These steps form the foundation of most supervised-learning workflows and are a great starting point for experimenting with classification algorithms.

## Links and references

* scikit-learn: load\_iris — [https://scikit-learn.org/stable/modules/generated/sklearn.datasets.load\_iris.html](https://scikit-learn.org/stable/modules/generated/sklearn.datasets.load_iris.html)
* UCI Machine Learning Repository: Iris Data Set — [https://archive.ics.uci.edu/ml/datasets/iris](https://archive.ics.uci.edu/ml/datasets/iris)

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/kubeflow/module/d9b1b119-0c6f-494b-b063-8eccd99dbff7/lesson/ac8dd891-4c3a-46ae-9528-9bc9e2be7cd8" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.