> ## Documentation Index
> Fetch the complete documentation index at: https://notes.kodekloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Data Features Labels and Feature Engineering

> Explains data, features, labels, and feature engineering for ML using housing dataset examples, covering preprocessing, encoding, data quality, bias, and a hands-on cleaning lab.

Now that we understand machine learning as learning from examples, the next step is to define the core ingredients of those examples: the data, the features, and the labels.

## What is data in ML?

Data is the raw material a model learns from — emails, images, audio, transaction logs, medical scans, sensor readings, or rows in a spreadsheet. A concrete, widely-used example is house-price prediction.

Kaggle’s "House Prices: Advanced Regression Techniques" competition uses the Ames Housing dataset (79 explanatory variables describing homes in Ames, Iowa) to predict final sale price. See the competition for a real-world dataset and benchmarks:

* [Kaggle: House Prices: Advanced Regression Techniques](https://www.kaggle.com/competitions/house-prices-advanced-regression-techniques)

Imagine a spreadsheet of previously sold houses. Each row is one house (one sample, or training example). Each column describes an attribute: square footage, number of bedrooms, neighborhood, lot size, year built, presence of a garage, etc. That whole spreadsheet is the dataset.

* Features: the columns that describe the house (inputs).
* Label / target: the value the model must predict — in this example, the sale price.

This is the basic structure of supervised learning: examples where the correct answer is known (features → label). If labeled answers are not available, the task may be unsupervised learning (discover patterns or structure), which we’ll cover elsewhere.

## Feature engineering — why it matters

Feature engineering is the process of creating, transforming, or combining features so the model has better information to learn from. The goal is to represent information in a way that makes predictive patterns easier to find, not to “cheat.”

Examples:

* Convert `year_built` into `age` (more meaningful to many models).
* Combine `bedrooms` and `bathrooms` into `total_rooms`.
* Convert a categorical `neighborhood` into numeric form via `one-hot encoding`.

Common preprocessing steps include missing-value handling, deduplication, scaling/normalization, and encoding categorical variables.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/qBe6x55gupKpUs7F/images/Machine-Learning-Fundamentals/Data-and-Training-Basics/Data-Features-Labels-and-Feature-Engineering/house-prices-dirty-data-ml-overfit.jpg?fit=max&auto=format&n=qBe6x55gupKpUs7F&q=85&s=faff3729803a58ff815a72a9c3ba15e3" alt="A hand-drawn diagram of a machine-learning pipeline showing a house-prices dataset (columns like sqft, beds, area, price) with missing/wrong entries. Arrows lead through feature engineering and data preparation into a neural-network model that produces predictions (noting overfit and wrong outcomes)." width="1920" height="1080" data-path="images/Machine-Learning-Fundamentals/Data-and-Training-Basics/Data-Features-Labels-and-Feature-Engineering/house-prices-dirty-data-ml-overfit.jpg" />
</Frame>

## Types of features and how to prepare them

| Feature type | Description | Typical preprocessing |
| - | - | - |
| Numerical | Continuous or integer values (e.g., `sqft`, `age`) | Imputation, scaling, binning |
| Categorical | Discrete categories (e.g., `neighborhood`, `house_style`) | `one-hot encoding`, `label encoding` |
| Ordinal | Categories with an order (e.g., `condition` levels) | Map to integers preserving order |
| Text | Free-form strings (e.g., descriptions) | Tokenize, embed, or extract features |
| Date/time | Timestamps (e.g., sale date) | Extract year/month/age, cyclic transforms |

Converting categorical features to numbers is a crucial step because most models operate on numeric input. Common strategies include `one-hot encoding` for nominal categories and ordinal encoding when category order matters.

## Data quality, distribution, and bias

Data preparation is often the most time-consuming part of a successful ML project. If features are wrong, incomplete, or biased, the model will learn incorrect patterns — “garbage in, garbage out.”

A key concept is data distribution: the pattern of examples in your training set. If the real-world distribution differs from your training distribution (domain shift), model performance can drop. For example, a model trained mainly on suburban homes may underperform on tiny urban studios.

Bias in this context can mean any non-representative or skewed dataset, not just intentional unfairness. Always evaluate whether your dataset reflects the diversity and conditions of the environment where the model will be used.

<Frame>
  <img src="https://mintcdn.com/kodekloud-c4ac6d9a/qBe6x55gupKpUs7F/images/Machine-Learning-Fundamentals/Data-and-Training-Basics/Data-Features-Labels-and-Feature-Engineering/house-prices-neural-network-data-bias.jpg?fit=max&auto=format&n=qBe6x55gupKpUs7F&q=85&s=610ce3032ef5d9f97f0ee6edc7fffb0d" alt="A hand-drawn diagram titled &#x22;Basic Ingredients&#x22; showing a house-prices dataset table (sqft, beds, area, $$) feeding into a neural-network-style model. The sketch highlights problems like &#x22;garbage&#x22; data, bias and non-representative distributions, with a prediction example (New York crossed out, $3000/mo)." width="1920" height="1080" data-path="images/Machine-Learning-Fundamentals/Data-and-Training-Basics/Data-Features-Labels-and-Feature-Engineering/house-prices-neural-network-data-bias.jpg" />
</Frame>

## Hands-on lab: cleaning a small housing dataset

A hands-on lab accompanying these materials walks you through cleaning a raw dataset of 300 house sales that contains realistic issues: missing sale prices, duplicate rows, inconsistent fields, etc. The objective is to produce a cleaned CSV ready for training.

Typical lab tasks:

* Drop rows with missing or duplicate data
* Create a new `age` feature from `year_built`
* Convert `neighborhood` names into numeric columns using `one-hot encoding`
* Validate and save the cleaned output as a new CSV

Example terminal session from the lab:

```bash theme={null}
Welcome to the KodeKloud Hands-On lab
KODEKLOUD
All rights reserved

controlplane ~ on ☁ (us-east-1) -> []
controlplane ~ on ☁ (us-east-1) -> head -3 houses.csv
sqft,bedrooms,year_built,neighborhood,sale_price
1850,3,2005,riverside,412000
2400,4,1998,hilltop,

controlplane ~ on ☁ (us-east-1) -> python3 clean.py
wrote houses_clean.csv (300 rows)

controlplane ~ on ☁ (us-east-1) -> head -1 houses_clean.csv
sqft,bedrooms,year_built,sale_price,age,neighborhood_downtown,neighborhood_hilltop,...
```

The resulting `houses_clean.csv` is ready for model training and evaluation.

<Callout icon="lightbulb" color="#1CB2FE">
  Data quality and feature representation typically have a larger impact on model performance than the choice of algorithm. Allocate ample time to cleaning, validating, and engineering features — those efforts yield the biggest improvements.
</Callout>

## Quick glossary

* Feature: An input variable used by a model (e.g., `sqft`).
* Label / Target: The value to predict (e.g., `sale_price`).
* Feature engineering: Creating or transforming features to improve model learning.
* One-hot encoding: Converting a categorical variable to a set of binary columns.
* Data distribution: How examples are spread across feature space; mismatch causes performance issues.
* Bias: Skew or incompleteness in the data that leads to poor generalization.

## Links and references

* [Kaggle: House Prices: Advanced Regression Techniques](https://www.kaggle.com/competitions/house-prices-advanced-regression-techniques)
* Scikit-learn preprocessing: `OneHotEncoder` and `LabelEncoder` — [https://scikit-learn.org/stable/modules/preprocessing.html](https://scikit-learn.org/stable/modules/preprocessing.html)
* Introduction to feature engineering (blog): [https://www.feature-engineering.org/](https://www.feature-engineering.org/)

<CardGroup>
  <Card title="Watch Video" icon="video" cta="Learn more" href="https://learn.kodekloud.com/user/courses/machine-learning-fundamentals/module/afd8936e-2a0c-4082-a5be-64b0780c9d9d/lesson/74e8df92-d347-4b05-912f-388c7b3496c9" />

  <Card title="Practice Lab" icon="flask-conical" cta="Learn more" href="https://learn.kodekloud.com/user/courses/machine-learning-fundamentals/module/afd8936e-2a0c-4082-a5be-64b0780c9d9d/lesson/d063ca12-8343-4c9a-aaa1-224c32924f69" />
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.