Skip to main content
Now that we understand machine learning as learning from examples, the next step is to define the core ingredients of those examples: the data, the features, and the labels.

What is data in ML?

Data is the raw material a model learns from — emails, images, audio, transaction logs, medical scans, sensor readings, or rows in a spreadsheet. A concrete, widely-used example is house-price prediction. Kaggle’s “House Prices: Advanced Regression Techniques” competition uses the Ames Housing dataset (79 explanatory variables describing homes in Ames, Iowa) to predict final sale price. See the competition for a real-world dataset and benchmarks: Imagine a spreadsheet of previously sold houses. Each row is one house (one sample, or training example). Each column describes an attribute: square footage, number of bedrooms, neighborhood, lot size, year built, presence of a garage, etc. That whole spreadsheet is the dataset.
  • Features: the columns that describe the house (inputs).
  • Label / target: the value the model must predict — in this example, the sale price.
This is the basic structure of supervised learning: examples where the correct answer is known (features → label). If labeled answers are not available, the task may be unsupervised learning (discover patterns or structure), which we’ll cover elsewhere.

Feature engineering — why it matters

Feature engineering is the process of creating, transforming, or combining features so the model has better information to learn from. The goal is to represent information in a way that makes predictive patterns easier to find, not to “cheat.” Examples:
  • Convert year_built into age (more meaningful to many models).
  • Combine bedrooms and bathrooms into total_rooms.
  • Convert a categorical neighborhood into numeric form via one-hot encoding.
Common preprocessing steps include missing-value handling, deduplication, scaling/normalization, and encoding categorical variables.
A hand-drawn diagram of a machine-learning pipeline showing a house-prices dataset (columns like sqft, beds, area, price) with missing/wrong entries. Arrows lead through feature engineering and data preparation into a neural-network model that produces predictions (noting overfit and wrong outcomes).

Types of features and how to prepare them

Converting categorical features to numbers is a crucial step because most models operate on numeric input. Common strategies include one-hot encoding for nominal categories and ordinal encoding when category order matters.

Data quality, distribution, and bias

Data preparation is often the most time-consuming part of a successful ML project. If features are wrong, incomplete, or biased, the model will learn incorrect patterns — “garbage in, garbage out.” A key concept is data distribution: the pattern of examples in your training set. If the real-world distribution differs from your training distribution (domain shift), model performance can drop. For example, a model trained mainly on suburban homes may underperform on tiny urban studios. Bias in this context can mean any non-representative or skewed dataset, not just intentional unfairness. Always evaluate whether your dataset reflects the diversity and conditions of the environment where the model will be used.
A hand-drawn diagram titled "Basic Ingredients" showing a house-prices dataset table (sqft, beds, area, $$) feeding into a neural-network-style model. The sketch highlights problems like "garbage" data, bias and non-representative distributions, with a prediction example (New York crossed out, $3000/mo).

Hands-on lab: cleaning a small housing dataset

A hands-on lab accompanying these materials walks you through cleaning a raw dataset of 300 house sales that contains realistic issues: missing sale prices, duplicate rows, inconsistent fields, etc. The objective is to produce a cleaned CSV ready for training. Typical lab tasks:
  • Drop rows with missing or duplicate data
  • Create a new age feature from year_built
  • Convert neighborhood names into numeric columns using one-hot encoding
  • Validate and save the cleaned output as a new CSV
Example terminal session from the lab:
The resulting houses_clean.csv is ready for model training and evaluation.
Data quality and feature representation typically have a larger impact on model performance than the choice of algorithm. Allocate ample time to cleaning, validating, and engineering features — those efforts yield the biggest improvements.

Quick glossary

  • Feature: An input variable used by a model (e.g., sqft).
  • Label / Target: The value to predict (e.g., sale_price).
  • Feature engineering: Creating or transforming features to improve model learning.
  • One-hot encoding: Converting a categorical variable to a set of binary columns.
  • Data distribution: How examples are spread across feature space; mismatch causes performance issues.
  • Bias: Skew or incompleteness in the data that leads to poor generalization.

Watch Video

Practice Lab