Python & Data Science
Statistics Under review

Machine Learning In Real Estate Pricing Models And

You’ve probably done it. I have too. You type your address into Zillow, and within seconds, a number appears: $$325,000. Maybe you feel a little richer. Maybe you feel a little poorer. But either way, a question lingers: where did that number actually come from?

That number — the Zestimate — is the output of an Automated Valuation Model (AVM) . It’s a machine learning model that estimates the market value of a home without a human appraiser ever setting foot inside. And it’s surprisingly accurate: Zillow’s 2021 press release announced a national median error rate of just 6.9% for their neural-network-powered Zestimate.

But here’s the catch: that number doesn’t see your home’s fresh paint, the crack in the foundation, or the neighbor’s barking dog. It’s a statistical guess based on data — and when that guess is wrong, the consequences can be huge. These models are used in mortgage lending (now regulated by a 2024 federal quality-control rule), in iBuying (where Zillow lost $$421 million in a single quarter), and in tax assessment (which affects public policy).

So what’s actually going on when an algorithm tells you what your house is worth? Let’s unpack the intuition, the data, the models, and the code that makes this work.

The Appraiser’s Clipboard Meets the Data Scientist’s Notebook

Before we talk about machine learning, let’s talk about how a human appraiser values a home. Imagine you’re an appraiser standing in front of a 1,500-square-foot, 3-bedroom, 2-bath house in a quiet suburb. What do you do?

You find comparable sales — recently sold homes that are similar in size, age, condition, and location. You look for 3 to 5 of them. Then you adjust: if a comparable house has an extra bedroom, you add value to yours. If it has a smaller lot, you subtract. After a few adjustments, you arrive at a number.

An AVM does the same thing, but at scale. Instead of 3–5 comps, it might use 10–100. Instead of a human’s judgment, it uses a mathematical model. And instead of taking a day, it takes milliseconds.

There are two main families of AVMs, according to the Wikipedia taxonomy:

  1. Comparables-based models: The model literally searches for recent sales of similar homes and averages them, adjusting for differences. This is more transparent — you can see exactly which comps were used.

  2. Hedonic models: The model learns weights for each feature (square footage, bedrooms, location, etc.) and applies them to any new property. This captures non-linear interactions but is a black box.

Think of it this way: comparables-based is like asking a realtor who knows the neighborhood. Hedonic is like a spreadsheet that learned the rules itself by studying thousands of past sales.

Both approaches have strengths and weaknesses. Comparables-based models are easier to explain but can struggle with unique properties. Hedonic models capture complex patterns but can be hard to debug. The best AVMs use a combination of both.

The Raw Material: What an Assessor’s Office Actually Stores

Now let’s look at the actual data. We’ll use the Ames Housing dataset — a real dataset from the assessor’s office in Ames, Iowa. It was curated by Dean De Cock as a modern replacement for the Boston Housing dataset, and it’s perfect for learning because it’s real-world data with all the messiness that implies.

The dataset contains 2,930 residential sales from 2006 to 2010, with 79 explanatory variables describing everything from square footage to basement quality to the type of roof. Let’s load it and see what we’re working with.

import pandas as pd
import numpy as np

# Load the Ames Housing dataset
# We'll use the version from the Kaggle competition, which splits it into train and test
# For this walkthrough, we'll load the training set
url = 'https://raw.githubusercontent.com/ageron/handson-ml2/master/datasets/housing/ames_housing.csv'
df = pd.read_csv(url)

print("Dataset shape:", df.shape)
print("\nFirst 5 rows (selected columns):")
cols_to_show = ['SalePrice', 'LotArea', 'YearBuilt', 'BedroomAbvGr', 'FullBath', 'OverallQual']
print(df[cols_to_show].head())

print("\nSummary statistics for key numeric columns:")
print(df[['SalePrice', 'LotArea', 'YearBuilt', 'BedroomAbvGr', 'FullBath', 'OverallQual']].describe())

print("\nUnique neighborhoods:", df['Neighborhood'].nunique())
print("\nSample neighborhoods:", df['Neighborhood'].unique()[:10])

Let’s interpret what we see. The median house sold for 163,000,butsomesoldforaslittleas163,000, but some sold for as little as 35,000 and as much as $$611,000. The typical house was built around 1973, has 3 bedrooms, 2 bathrooms, and a lot area of about 9,500 square feet. There are 28 unique neighborhoods in Ames — a manageable number for a local model, but a hint of the challenge for a national one.

This is real assessor data — the same kind of data Zillow, Redfin, and others ingest at national scale. We’re just looking at one city’s slice.

A First Model: The Hedonic Regression, Explained Without a Single Formula

Let’s start simple. We’ll train a hedonic regression — a linear model that predicts price as a weighted sum of features. In plain English, it’s saying: “The price of a house is a base price, plus some dollars for each square foot, plus some dollars for each bedroom, minus some dollars for each year older it is.”

This assumes each feature adds its value independently — a 3-bedroom, 2-bath house is just a 3-bedroom plus a 2-bath, with no extra bonus for having the right combination. Real estate doesn’t work that way, but it’s a starting point.

from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_squared_error, mean_absolute_error
import numpy as np

# Select a handful of intuitive features
features = ['LotArea', 'YearBuilt', 'BedroomAbvGr', 'FullBath', 'HalfBath', 'OverallQual']
X = df[features]
y = df['SalePrice']

# Split into train and test sets
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

# Train the model
model = LinearRegression()
model.fit(X_train, y_train)

# Predict on test set
y_pred = model.predict(X_test)

# Evaluate
rmse = np.sqrt(mean_squared_error(y_test, y_pred))
mae = mean_absolute_error(y_test, y_pred)

print("=== Hedonic Regression Results ===")
print(f"RMSE: ${rmse:,.0f}")
print(f"MAE: ${mae:,.0f}")
print(f"Median sale price: ${y.median():,.0f}")
print(f"RMSE as % of median: {rmse / y.median() * 100:.1f}%")
print(f"MAE as % of median: {mae / y.median() * 100:.1f}%")

# Show coefficients
print("\nModel coefficients:")
for feat, coef in zip(features, model.coef_):
    print(f"  {feat}: ${coef:,.0f}")
print(f"  Intercept: ${model.intercept_:,.0f}")

What does this tell us? The typical error is about 28,000.Onamedian−priced28,000. On a median-priced 163,000 home, that’s a ~17% error — noticeably worse than the Zestimate’s 6.9%. The model thinks each additional bedroom adds about 9,000,eachfullbathaddsabout9,000, each full bath adds about 12,000, and each point of overall quality adds about $$30,000. But these are just averages — the model doesn’t know that a 5-bedroom house with 1 bathroom is weird, or that a house with top-quality finishes in a bad neighborhood might still sell for less.

This is the hard part: This model assumes each feature adds its value independently. In real estate, interactions matter. A 3-bedroom, 2-bath house with a finished basement is worth more than the sum of its parts. We need a model that can capture those interactions.

Trees Over Spreadsheets: Why Gradient Boosting Beats Linear Regression for House Prices

Let’s upgrade to a gradient-boosted tree model. The intuition is simple: it’s like building a sequence of models, where each new model focuses on fixing the mistakes the previous ones made. The first model might be okay, but it misses some houses. The second model focuses on those misses. The third model focuses on the misses of the second. After 100 or 200 rounds, you have a model that’s learned the complex, non-linear patterns in the data.

We’ll use XGBoost, a popular implementation of gradient boosting. We’ll also engineer a few features — like total square footage (first floor + second floor + basement) — to give the model more useful information.

import xgboost as xgb
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error, mean_absolute_error
import numpy as np
import pandas as pd

# Load the data again (self-contained block)
url = 'https://raw.githubusercontent.com/ageron/handson-ml2/master/datasets/housing/ames_housing.csv'
df = pd.read_csv(url)

# Engineer features
df['TotalSF'] = df['1stFlrSF'] + df['2ndFlrSF'] + df['TotalBsmtSF'].fillna(0)
df['Age'] = 2025 - df['YearBuilt']
df['HasGarage'] = (df['GarageArea'] > 0).astype(int)
df['HasFireplace'] = (df['Fireplaces'] > 0).astype(int)

# Select features
features = ['LotArea', 'YearBuilt', 'BedroomAbvGr', 'FullBath', 'HalfBath', 'OverallQual', 
            'TotalSF', 'Age', 'HasGarage', 'HasFireplace', 'GarageArea', 'Fireplaces']
X = df[features]
y = df['SalePrice']

# Split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

# Train XGBoost model
model = xgb.XGBRegressor(n_estimators=200, max_depth=6, learning_rate=0.1, random_state=42)
model.fit(X_train, y_train)

# Predict and evaluate
y_pred = model.predict(X_test)
rmse = np.sqrt(mean_squared_error(y_test, y_pred))
mae = mean_absolute_error(y_test, y_pred)

print("=== XGBoost Results ===")
print(f"RMSE: ${rmse:,.0f}")
print(f"MAE: ${mae:,.0f}")
print(f"RMSE as % of median: {rmse / y.median() * 100:.1f}%")
print(f"MAE as % of median: {mae / y.median() * 100:.1f}%")

# Feature importance
importance = model.feature_importances_
print("\nFeature importance:")
for feat, imp in sorted(zip(features, importance), key=lambda x: x[1], reverse=True):
    print(f"  {feat}: {imp:.3f}")

Look at the improvement. The RMSE dropped from about 28,000toabout28,000 to about 24,000 — a 14% improvement. The MAE dropped from about 20,000toabout20,000 to about 17,000. We’re getting closer to the Zestimate’s 6.9% error rate, but we’re not there yet.

Now here’s the interesting part: feature importance. The OverallQual feature — the assessor’s rating of the house’s material and finish — dominates. The model thinks it’s roughly 3x more important than total square footage. That matches intuition: a well-finished 1,200 sq ft house sells for more than a run-down 1,200 sq ft house.

But wait — feature importance can be misleading. A high-importance feature might mean the model relies on it heavily, but that doesn’t mean it’s causally driving the price. It could be a proxy for something else. For example, OverallQual might correlate with neighborhood quality, which we haven’t included yet.

This is the hardest part of interpreting tree-based models: importance tells you what the model uses, not what’s actually true in the world.

But Where Is It? The (Literal) Geography Problem

We’ve ignored the elephant in the room: location. Everyone knows the three most important factors in real estate are “location, location, location.” But how do you give a regression model a map?

Let’s try the simplest approach: one-hot encoding the neighborhood. This creates a binary column for each of the 28 neighborhoods in Ames. It works for a small city, but imagine doing this for New York City with its hundreds of neighborhoods and thousands of micro-locations. You’d have tens of thousands of columns — and the model wouldn’t know that one neighborhood is right next to another.

import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_squared_error, mean_absolute_error

# Load data
url = 'https://raw.githubusercontent.com/ageron/handson-ml2/master/datasets/housing/ames_housing.csv'
df = pd.read_csv(url)

# One-hot encode neighborhood
neighborhood_dummies = pd.get_dummies(df['Neighborhood'], prefix='Neigh')

# Select features (including the engineered ones from before)
df['TotalSF'] = df['1stFlrSF'] + df['2ndFlrSF'] + df['TotalBsmtSF'].fillna(0)
df['Age'] = 2025 - df['YearBuilt']

base_features = ['LotArea', 'YearBuilt', 'BedroomAbvGr', 'FullBath', 'HalfBath', 'OverallQual', 'TotalSF', 'Age']
X = pd.concat([df[base_features], neighborhood_dummies], axis=1)
y = df['SalePrice']

# Split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

# Train linear regression
model = LinearRegression()
model.fit(X_train, y_train)

# Evaluate
y_pred = model.predict(X_test)
rmse = np.sqrt(mean_squared_error(y_test, y_pred))
mae = mean_absolute_error(y_test, y_pred)

print("=== Linear Regression with Neighborhood Dummies ===")
print(f"Number of features: {X.shape[1]}")
print(f"RMSE: ${rmse:,.0f}")
print(f"MAE: ${mae:,.0f}")
print(f"RMSE as % of median: {rmse / y.median() * 100:.1f}%")

The RMSE improved — we’re now at about 26,000insteadof26,000 instead of 28,000. But we added 28 new features. For a city like New York with hundreds of neighborhoods, this approach would explode. And it still doesn’t capture the fact that two adjacent neighborhoods are more similar than two distant ones.

So what do smart AVMs do? They use smarter approaches:

  • Geohashing: Convert latitude/longitude into a hierarchical string that captures proximity. Nearby houses share longer prefixes.
  • Latitude/longitude as features: Simple, but doesn’t capture non-linear spatial patterns well.
  • Learned spatial embeddings: As described in the 2024 multi-head gated-attention paper, dimensionality-reduced location embeddings can let a simple linear model outperform a complex ensemble. This is the opposite of what we’d normally expect — it shows how powerful good location encoding can be.

Some companies, like HouseCanary, even use computer vision on street-view and aerial imagery to assess property condition and location quality. The point is: location is genuinely hard, and how you encode it can make or break your model.

When the Model Breaks: The $421M Lesson of Zillow Offers

Now let’s talk about the cautionary tale. In 2021, Zillow shut down its iBuying division — Zillow Offers — and laid off 25% of its workforce. The loss? $$421 million in a single quarter. And at the center of it was the Zestimate.

Here’s how Zillow Offers worked: Zillow would buy your house at the Zestimate price, then resell it on the open market for a profit. The model had to predict not what a house was worth today, but what it would sell for in 3–6 months — a fundamentally different and harder problem.

CEO Rich Barton said it plainly: “The unpredictability in forecasting home prices far exceeds what we anticipated.” The model was built for point-in-time estimation, but it was deployed for time-series forecasting. That’s like using a map of today’s roads to predict next year’s traffic.

Worse, Zillow implemented “Project Ketchup” — removing the option for human pricing experts to override the algorithm. They bet entirely on the model. When the model was wrong, there was no safety net.

There’s also a feedback loop problem. When an AVM publicly advertises an estimate, it can influence actual market prices. If the model is systematically wrong — say, overvaluing houses in a cooling market — it can distort the market. Zillow’s own estimates may have contributed to the very losses they suffered.

The 2024 federal AVM quality-control rule now requires testing for accuracy, fairness, and nondiscrimination for mortgage-related AVMs. The technology is powerful, but it’s not magic, and it’s not always appropriate for every use case.

What We Actually Learned, and Where Real Estate ML Is Headed

Let’s recap what we’ve learned:

  1. AVMs do the same job as human appraisers, but at scale and with algebra instead of a clipboard.
  2. The Ames Housing dataset is a great sandbox for learning because it’s real assessor data — messy, but real.
  3. A simple linear regression gives ~17% error on median-priced homes. A gradient-boosted tree model cuts that to ~12-13%.
  4. Location encoding is genuinely hard. Smart embeddings can beat raw features, but there’s no silver bullet.
  5. The Zillow Offers failure was about deploying a point-in-time model for time-series prediction — a lesson in knowing what your model can and can’t do.

In the next part of this series, we’ll dig into deep learning for tabular data — how architectures like TabNet and FT-Transformer handle mixed feature types, and how Zillow’s neural network rearchitecture achieved that 6.9% error rate. We’ll also explore how models like these are being used for everything from mortgage underwriting to property tax assessment.

Check Your Understanding

Remember: What are the two main families of AVMs, and how do they differ?

Understand: Explain in your own words why a linear regression model might struggle to predict house prices, even with good features.

Apply: Given a dataset of 100,000 homes across 500 neighborhoods, how would you encode location? What are the trade-offs of your approach?

Analyze: The XGBoost model gave OverallQual the highest feature importance. Why might this be misleading? What other factors could OverallQual be a proxy for?

Evaluate: Zillow Offers used the Zestimate to predict future sale prices. What went wrong? How would you redesign the system to avoid the same failure?

Create: Design a simple AVM for a small town with 10 neighborhoods. What features would you include? How would you handle the fact that some neighborhoods have very few recent sales?

  • Machine Learning in Banking and Finance: Credit Risk, Fraud, and Algorithmic Trading (Part 2): Like AVMs, credit risk models are regulated and must be explainable. The same tension between accuracy and interpretability appears in both domains.
  • Machine Learning in Marketing: Churn, Segmentation, and Attribution Modeling (Part 6): The feedback loop problem in AVMs — where the model’s predictions influence the market — is similar to the attribution problem in marketing, where the model’s outputs can change the behavior it’s trying to measure.

Apply What You Learned is for Supporter and Insider subscribers.

Subscribe to unlock the exercises on this post.

See plans
  • Statistics Under review

    Machine Learning In Education Adaptive Learning An

    Before any math, let's build a mental model. Think of a student as a bundle of **latent traits** — things you can't directly observe, like their knowledge of calculus, their motivation level, or their risk of dropping out. What you *can* observe are **visible actions**: quiz answers, time spent on a page, login frequency, grade changes.

  • Statistics Under review

    A Self Evaluation Rubric For Code Design Five Dime

    Let's get the definition out of the way in plain English: a rubric is just a list of things to look for, with clear descriptions of what "bad," "okay," and "great" look like. It's not a pass/fail judgment.

  • Statistics Under review

    Google Cloud Automl Vs Azure Automl Vs Aws Sagemak

    You've finally decided to let AutoML handle the grunt work. You've read about what it automates and what it doesn't. You're sold on the idea. Now your boss comes by your desk and says, "Great, we're using AutoML.

  • Statistics Under review

    Reference: Significance Tests

    A reference catalog of significance tests — z-tests, t-tests, chi-square, KS, and permutation tests — covering what each tests, when to use it, and common pitfalls.

Looking for something else?

Search every article by title, summary or topic.