What is data in ML?
Data is the raw material a model learns from — emails, images, audio, transaction logs, medical scans, sensor readings, or rows in a spreadsheet. A concrete, widely-used example is house-price prediction. Kaggle’s “House Prices: Advanced Regression Techniques” competition uses the Ames Housing dataset (79 explanatory variables describing homes in Ames, Iowa) to predict final sale price. See the competition for a real-world dataset and benchmarks: Imagine a spreadsheet of previously sold houses. Each row is one house (one sample, or training example). Each column describes an attribute: square footage, number of bedrooms, neighborhood, lot size, year built, presence of a garage, etc. That whole spreadsheet is the dataset.- Features: the columns that describe the house (inputs).
- Label / target: the value the model must predict — in this example, the sale price.
Feature engineering — why it matters
Feature engineering is the process of creating, transforming, or combining features so the model has better information to learn from. The goal is to represent information in a way that makes predictive patterns easier to find, not to “cheat.” Examples:- Convert
year_builtintoage(more meaningful to many models). - Combine
bedroomsandbathroomsintototal_rooms. - Convert a categorical
neighborhoodinto numeric form viaone-hot encoding.

Types of features and how to prepare them
Converting categorical features to numbers is a crucial step because most models operate on numeric input. Common strategies include
one-hot encoding for nominal categories and ordinal encoding when category order matters.
Data quality, distribution, and bias
Data preparation is often the most time-consuming part of a successful ML project. If features are wrong, incomplete, or biased, the model will learn incorrect patterns — “garbage in, garbage out.” A key concept is data distribution: the pattern of examples in your training set. If the real-world distribution differs from your training distribution (domain shift), model performance can drop. For example, a model trained mainly on suburban homes may underperform on tiny urban studios. Bias in this context can mean any non-representative or skewed dataset, not just intentional unfairness. Always evaluate whether your dataset reflects the diversity and conditions of the environment where the model will be used.
Hands-on lab: cleaning a small housing dataset
A hands-on lab accompanying these materials walks you through cleaning a raw dataset of 300 house sales that contains realistic issues: missing sale prices, duplicate rows, inconsistent fields, etc. The objective is to produce a cleaned CSV ready for training. Typical lab tasks:- Drop rows with missing or duplicate data
- Create a new
agefeature fromyear_built - Convert
neighborhoodnames into numeric columns usingone-hot encoding - Validate and save the cleaned output as a new CSV
houses_clean.csv is ready for model training and evaluation.
Data quality and feature representation typically have a larger impact on model performance than the choice of algorithm. Allocate ample time to cleaning, validating, and engineering features — those efforts yield the biggest improvements.
Quick glossary
- Feature: An input variable used by a model (e.g.,
sqft). - Label / Target: The value to predict (e.g.,
sale_price). - Feature engineering: Creating or transforming features to improve model learning.
- One-hot encoding: Converting a categorical variable to a set of binary columns.
- Data distribution: How examples are spread across feature space; mismatch causes performance issues.
- Bias: Skew or incompleteness in the data that leads to poor generalization.
Links and references
- Kaggle: House Prices: Advanced Regression Techniques
- Scikit-learn preprocessing:
OneHotEncoderandLabelEncoder— https://scikit-learn.org/stable/modules/preprocessing.html - Introduction to feature engineering (blog): https://www.feature-engineering.org/