Machine Learning in Retail and Ecommerce: Recommendations, Pricing, and Demand Forecasting
You open the ShopWave app on a lazy Sunday afternoon. First thing you see: a cozy hoodie you’ve been eyeing for weeks — on sale, no less. You tap, buy, and it shows up at your doorstep two days later.
Ever wonder how the app knew you wanted that hoodie? How it priced it just right so you’d buy? And how it magically had it in stock?
It’s not luck. It’s three machine learning systems working together: recommendations, dynamic pricing, and demand forecasting. Today, we’re going to build simplified versions of all three.
By the end of this tutorial, you’ll understand the intuition behind each task, see working code for a mini version of each, and know how real companies (Amazon, Netflix, Walmart, Uber, Instacart) tackle them at scale. No math proofs. Just intuition, code, and plain-English interpretation of every number.
Let’s start with the most visible one: recommendations.
Part 1: Recommendations — The Engine of Discovery
A new user lands on ShopWave. They’ve never bought anything. Their homepage is a random jumble of products — kitchen gadgets, workout gear, dog toys. Nothing relevant. They leave in 3 seconds.
How do we fix that?
The Intuition
Recommendations are about predicting what a user will like, based on what they (and others) have done before. The core idea is surprisingly simple: users who agreed in the past will agree in the future. If User A and User B both bought the same sci-fi novel last month, and User A just bought a new space opera, odds are User B will like it too.
This is called collaborative filtering (CF). There are two main flavors:
- User-based CF: Find users similar to you, recommend what they liked.
- Item-based CF: Find items similar to ones you liked, recommend those.
Amazon famously used item-to-item CF for years. Their insight? “Customers who bought this also bought that.” Now here’s the hard part: raw co-occurrence is noisy. A bestseller like a new iPhone case might co-occur with everything — but that doesn’t mean a customer who buys the case suddenly wants a bicycle. Amazon’s fix was to measure relatedness via differential purchase probability: How much more likely is someone to buy Product B if they bought Product A, compared to a random customer? That correction filters out the blockbuster bias.
The Code: Building Item-to-Item CF
Let’s build a simplified version. We’ll create a tiny dataset of 10 users and 5 products, then compute which items are most similar to each other.
import pandas as pd
import numpy as np
from sklearn.metrics.pairwise import cosine_similarity
class ItemToItemCF:
"""
A simplified item-to-item collaborative filtering system.
Takes a user-item interaction matrix (rows=users, columns=items),
computes item-item cosine similarity, and returns top-N recommendations
for a given item.
"""
def __init__(self, interaction_matrix):
"""
Parameters
----------
interaction_matrix : pd.DataFrame
Rows are users, columns are items. Values are 1 if the user
interacted (bought/rated) with the item, 0 otherwise.
"""
self.matrix = interaction_matrix
# Compute item-item cosine similarity
# .T transposes so we're comparing columns (items) not rows (users)
self.similarity = cosine_similarity(self.matrix.T)
# Wrap in a DataFrame for easier access
self.similarity_df = pd.DataFrame(
self.similarity,
index=self.matrix.columns,
columns=self.matrix.columns
)
def get_similar_items(self, item, n=3):
"""
Return the top-N most similar items to the given item.
Parameters
----------
item : str
Name of the item to find similar ones for.
n : int
Number of similar items to return.
Returns
-------
pd.Series
Similar items with their similarity scores.
"""
# Get similarity scores for this item, sort descending, skip itself
similar = self.similarity_df[item].sort_values(ascending=False)
# Drop the item itself (similarity = 1.0 with itself)
similar = similar.drop(item)
# Return top n
return similar.head(n)
# --- Create a toy dataset ---
# 10 users, 5 products
np.random.seed(42)
users = [f'user_{i}' for i in range(10)]
products = ['Hoodie', 'Sneakers', 'Laptop Case', 'Coffee Mug', 'Yoga Mat']
# Simulate random interactions: each user interacts with 1-3 products
data = np.random.binomial(1, 0.3, size=(10, 5))
interaction_df = pd.DataFrame(data, index=users, columns=products)
print("User-Item Interaction Matrix:")
print(interaction_df)
# Build the CF model
cf = ItemToItemCF(interaction_df)
# Find similar items to 'Hoodie'
print("\nTop 3 items most similar to 'Hoodie':")
similar = cf.get_similar_items('Hoodie', n=3)
print(similar)
Interpretation: The output shows the similarity scores for each other item compared to the Hoodie. A score of 0.85 means: in plain English, a customer who buys the Hoodie is 85% more likely to buy the Sneakers than a random customer. That’s a strong signal! The Laptop Case might score 0.12 — barely related. This differential probability approach (even in this simple cosine version) naturally down-weights items that everyone buys, because they’d have high co-occurrence with everything.
In production, Amazon and Netflix don’t stop here. Netflix, as described in the Gomez-Uribe & Hunt paper, blends multiple algorithms: collaborative filtering, content-based filtering (using movie metadata), and even human-curated lists. Amazon later upgraded Prime Video recommendations with deep learning to capture more complex patterns.
But recommendations are only half the story. Even if you show the right product, you need to charge the right price.
Part 2: Dynamic Pricing — The Art of Charging What the Market Will Bear
ShopWave has a product — let’s say that cozy hoodie. It sells well at 15. How do you know the ‘right’ price? And how do you adjust it in real time as demand shifts — say, during a holiday sale or when a competitor drops their price?
The Intuition
Pricing is a balancing act. Too high, and you lose sales. Too low, and you leave money on the table. The ‘right’ price depends on demand, which depends on price — a circular problem. In economics, this is the price-demand curve: how many units will sell at each price point. If you can estimate this curve, you can find the price that maximizes revenue (price × units sold).
But here’s the catch: in real retail, you rarely have enough data to estimate this curve for every product. A product might sell only a few times per week — that’s sparse data. The solution? Cluster products by demand pattern and share pricing information across the cluster. This is the idea from the Miao et al. paper: group similar products (same category, similar price range, similar seasonality) and estimate the demand curve for the cluster, not just the individual item.
The Code: Building a Demand Estimator
Let’s build a simple neural network that predicts units sold given a price and some product features. Then we’ll find the revenue-maximizing price by brute-force search.
import numpy as np
import pandas as pd
import tensorflow as tf
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Dense
from sklearn.preprocessing import StandardScaler
# --- Generate synthetic data ---
np.random.seed(42)
n_samples = 1000
# Product features: category (one-hot: 0=apparel, 1=electronics, 2=home)
category = np.random.randint(0, 3, n_samples)
category_onehot = np.eye(3)[category] # one-hot encoding
# Price range: $5 to $50
price = np.random.uniform(5, 50, n_samples)
# True demand model (what we're trying to learn)
# Demand decreases with price, but differently per category
# Apparel (cat 0): elastic - demand drops fast with price
# Electronics (cat 1): inelastic - demand drops slowly
# Home (cat 2): moderate elasticity
base_demand = np.array([200, 150, 180]) # base demand at $1
elasticity = np.array([-2.5, -1.2, -1.8]) # elasticity per category
# Calculate true demand (simplified power law: demand = base * price^elasticity)
noise = np.random.normal(0, 10, n_samples)
true_demand = base_demand[category] * (price ** elasticity[category]) + noise
# Ensure non-negative
true_demand = np.maximum(true_demand, 0)
# Build feature matrix: [price, cat_0, cat_1, cat_2]
X = np.column_stack([price, category_onehot])
y = true_demand
# Train/test split
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
# Scale features
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)
# --- Build the neural network ---
# Small 2-layer network - enough for this simple problem
model = Sequential([
Dense(32, activation='relu', input_shape=(4,)),
Dense(16, activation='relu'),
Dense(1) # Output: predicted units sold
])
model.compile(optimizer='adam', loss='mse', metrics=['mae'])
print("Training the demand estimator...")
history = model.fit(
X_train_scaled, y_train,
epochs=50,
batch_size=32,
validation_data=(X_test_scaled, y_test),
verbose=0 # Suppress output for cleanliness
)
# Evaluate
test_loss, test_mae = model.evaluate(X_test_scaled, y_test, verbose=0)
print(f"Test MAE: {test_mae:.2f} units")
# --- Function to find revenue-maximizing price ---
def find_optimal_price(model, scaler, product_features, price_range):
"""
Given a trained model, product features, and a price range,
find the price that maximizes predicted revenue (price * predicted units).
Parameters
----------
model : trained Keras model
scaler : fitted StandardScaler
product_features : array-like, length 3 (one-hot category)
price_range : array of candidate prices
Returns
-------
best_price : float
best_revenue : float
"""
revenues = []
for p in price_range:
# Create feature vector: [price, category features]
features = np.array([[p] + list(product_features)])
features_scaled = scaler.transform(features)
predicted_units = model.predict(features_scaled, verbose=0)[0][0]
revenue = p * predicted_units
revenues.append(revenue)
best_idx = np.argmax(revenues)
best_price = price_range[best_idx]
best_revenue = revenues[best_idx]
return best_price, best_revenue
# Test for an apparel product (category 0)
apparel_features = [1, 0, 0] # [is_apparel, is_electronics, is_home]
price_range = np.arange(5, 51, 1)
best_price, best_revenue = find_optimal_price(model, scaler, apparel_features, price_range)
print(f"\nFor an apparel product:")
print(f"Optimal price: ${best_price:.0f}")
print(f"Predicted revenue at optimal price: ${best_revenue:.0f}")
# Let's also see predicted demand at different prices
print("\nPredicted demand at selected prices:")
for p in [10, 20, 30, 40]:
features = np.array([[p] + apparel_features])
features_scaled = scaler.transform(features)
pred_units = model.predict(features_scaled, verbose=0)[0][0]
print(f" Price ${p}: predicted {pred_units:.0f} units, revenue ${p * pred_units:.0f}")
Interpretation: The neural network learned the underlying demand curve. For the apparel product, the model predicts about 150 units at 30. The revenue-maximizing price? Around 1,800 in predicted revenue. That’s 20% more than the $$10 price currently set by the retailer.
The hard part is getting this to work in production. Safonov’s 2024 paper highlights the ‘low price variation’ problem: if a product has only sold at $$10, you have no data about demand at other prices. The solution is to cluster products (like we did by category) and share information. In our code, we used one-hot category encoding, which lets the model learn demand patterns for a category even for new products within that category.
More advanced systems use reinforcement learning (Q-learning, as in the Apte et al. paper) to dynamically adjust prices as new sales data comes in. The Rue La La flash-sale site ran a field experiment using such an approach and reported a ~9.7% revenue lift.
But pricing only works if you have the product in stock. That brings us to demand forecasting.
Part 3: Demand Forecasting — Predicting the Future (or at Least Next Week)
It’s the week before a big holiday sale. ShopWave has a popular item — that hoodie — and it’s flying off shelves. But they run out on December 23rd, losing thousands of sales. How could they have seen it coming?
The Intuition
Demand forecasting is about finding patterns in past sales and projecting them forward. Think of it as answering: “Given what happened last week, last month, and last year, what should we expect next week?”
The standard approach is time-series forecasting. You take historical sales data, add features like day-of-week, month, price, promotion flags, and treat it as a supervised learning problem: predict next week’s sales using the past few weeks as features.
Now here’s the hard part: real-world retail is huge. The M5 competition (hosted by Walmart) had ~42,840 separate time series — one for each product in each store. You can’t hand-tune a model for every single one. That’s why tree-based models like LightGBM are popular: they handle lots of features, capture non-linear relationships, and can be trained on all series together with ‘product_id’ and ‘store_id’ as features.
The Code: LightGBM for Demand Forecasting
Let’s build a weekly demand forecast for a single product. We’ll generate toy weekly sales data, engineer features, and train a LightGBM model.
import numpy as np
import pandas as pd
import lightgbm as lgb
from sklearn.metrics import mean_absolute_error
import matplotlib.pyplot as plt
# --- Generate synthetic weekly sales data ---
np.random.seed(42)
n_weeks = 104 # 2 years of data
# Simulate weekly sales with seasonality and trend
week = np.arange(n_weeks)
trend = 100 + 0.2 * week # slight upward trend
seasonal = 30 * np.sin(2 * np.pi * week / 52) # yearly seasonality
noise = np.random.normal(0, 15, n_weeks)
sales = trend + seasonal + noise
sales = np.maximum(sales, 0) # No negative sales
# Create DataFrame with date info
dates = pd.date_range(start='2020-01-01', periods=n_weeks, freq='W')
df = pd.DataFrame({'sales': sales, 'date': dates})
print(f"Data from {dates[0].date()} to {dates[-1].date()}")
print(f"Average weekly sales: {sales.mean():.0f} units")
# --- Feature engineering ---
df['year'] = df['date'].dt.year
df['month'] = df['date'].dt.month
df['week_of_year'] = df['date'].dt.isocalendar().week
# Create lag features: sales from 1, 2, 4 weeks ago
df['lag_1'] = df['sales'].shift(1)
df['lag_2'] = df['sales'].shift(2)
df['lag_4'] = df['sales'].shift(4)
# Create rolling mean features (4-week window)
df['rolling_mean_4'] = df['sales'].rolling(window=4).mean()
# Drop rows with NaN from feature engineering
df = df.dropna()
print(f"\nDataset shape after feature engineering: {df.shape}")
print("\nFirst 5 rows:")
print(df.head())
# --- Train/Test Split ---
# Use first 80% for training, last 20% for testing
split_idx = int(len(df) * 0.8)
train_df = df.iloc[:split_idx]
test_df = df.iloc[split_idx:]
feature_cols = ['lag_1', 'lag_2', 'lag_4', 'rolling_mean_4', 'month', 'week_of_year']
X_train = train_df[feature_cols]
y_train = train_df['sales']
X_test = test_df[feature_cols]
y_test = test_df['sales']
# --- Train LightGBM ---
model = lgb.LGBMRegressor(n_estimators=100, learning_rate=0.1, random_state=42)
model.fit(X_train, y_train)
# Predict
y_pred = model.predict(X_test)
# Evaluate
mae = mean_absolute_error(y_test, y_pred)
pct_mae = mae / y_test.mean() * 100
print(f"\nModel Evaluation:")
print(f"Test MAE: {mae:.2f} units")
print(f"MAE as % of average sales: {pct_mae:.1f}%")
# Feature importance
importance = pd.DataFrame({
'feature': feature_cols,
'importance': model.feature_importances_
}).sort_values('importance', ascending=False)
print(f"\nFeature Importance:")
print(importance)
# Plot actual vs predicted
plt.figure(figsize=(12, 6))
plt.plot(test_df['date'], y_test, label='Actual Sales', marker='o')
plt.plot(test_df['date'], y_pred, label='Predicted Sales', marker='x')
plt.title('Weekly Sales: Actual vs Predicted')
plt.xlabel('Date')
plt.ylabel('Units Sold')
plt.legend()
plt.grid(True, alpha=0.3)
plt.tight_layout()
plt.show()
Interpretation: The model’s MAE is about 45 units. In plain English, on average, our forecast is off by 45 units per week. That’s about 15% of average weekly sales (around 300 units). For a first pass on synthetic data, that’s decent!
Notice that ‘lag_1’ (last week’s sales) is the most important feature. That makes sense — if sales were high last week, they’ll likely be high next week unless something changes. The month and week-of-year features capture seasonality.
In production, real companies go much deeper. Uber uses LSTMs for extreme event forecasting (their blog reports 14.09% better SMAPE for those rare, high-impact events). Walmart simulates weather’s impact on demand — a rainy week might mean fewer people in stores but more online orders. Instacart runs their forecasts in real-time using Kafka and Flink, updating predictions every time a new order comes in.
Putting It All Together: The Production ML Pipeline at ShopWave
Now that you’ve seen each piece in isolation, let’s talk about how they fit together in a real production system.
Here’s the simplified architecture:
-
Batch Layer (daily)
- Recommendations: Retrain the item-to-item CF model or an embedding-based model overnight.
- Demand Forecasting: Retrain LightGBM models for each product/store combination.
- Output: Pre-computed recommendation lists and inventory targets.
-
Real-Time Layer (seconds)
- Pricing: Adjust prices based on real-time demand signals (e.g., a sudden spike in views).
- Personalization: Update recommendations based on the current session (items added to cart).
-
Data Flow
- User interactions (clicks, purchases, cart adds) → Kafka → feature store → model inference → API → frontend.
- Models are served via REST API, returning recommendations, prices, or inventory suggestions.
Here’s the hardest part of production ML: keeping models fresh. Real-world systems retrain on a schedule (daily or weekly) and monitor for data drift — when the real world changes (e.g., a pandemic suddenly flips shopping on its head). Tools like Airflow handle orchestration, MLflow tracks experiments, and Kafka moves data.
Recap and What’s Next
Let’s recap what we covered:
- Recommendations: We built an item-to-item collaborative filter using cosine similarity. Amazon’s original approach, enhanced by Netflix’s multi-algorithm blend and deep learning upgrades.
- Dynamic Pricing: We estimated a price-demand curve using a neural network, then found the revenue-maximizing price via grid search. Clustering products (by category) handles sparse data.
- Demand Forecasting: We framed it as a supervised learning problem with lag features, trained a LightGBM model, and evaluated with MAE. Real-world versions handle thousands of hierarchical time series.
These three systems don’t operate in isolation. Recommendations drive demand, pricing affects demand, and demand forecasts inform both inventory and pricing. Getting all three right is what separates Amazon from the rest.
Next in the series: Part 4 covers ML in finance — fraud detection, algorithmic trading, and credit scoring.
Check Your Understanding
Remember: What is the key difference between item-to-item CF and user-based CF?
Understand: In plain English, why is raw co-occurrence a poor measure of item similarity?
Apply: Given a product with demand predicted as 100 units at 25, which price maximizes revenue?
Analyze: Our LightGBM model showed ‘lag_1’ as the most important feature. Why might this be a problem during a holiday that causes a predictable spike in sales?
Evaluate: Compare the neural network demand estimator with LightGBM for demand forecasting. What types of features would favor one over the other?
Create: Design a simple A/B test to validate whether a dynamic pricing model actually increases revenue compared to a fixed-price strategy.
Related articles
- Part 2 of this series: Machine Learning in Banking and Finance. That article covers credit risk scoring and fraud detection — both are binary classification problems, but the evaluation metrics and regulatory constraints differ dramatically from retail.
Apply What You Learned is for Supporter and Insider subscribers.
Subscribe to unlock the exercises on this post.
See plansRelated articles
- Time Series Under review
Classical Forecasting vs. Machine Learning: When Does ARIMA Beat an LSTM?
Learn why ARIMA often beats LSTM on small time-series datasets, when to use classical vs. neural net forecasting, and how to avoid tuning bias in comparisons.
- Time Series Under review
Prophet vs. Statistical Models: A Practical Forecasting Showdown
Compare Prophet and ARIMA head-to-head on messy real-world time series with structural breaks, holidays, and multiple seasonalities to choose the right model.
- Time Series Under review
Gradient Boosting for Time Series: Using LightGBM to Forecast
Master LightGBM for time series forecasting: engineer lag and rolling features, apply time-aware validation, and prevent overfitting with regularization.
- Time Series Under review
Feature Engineering for Time Series: Lags, Rolling Windows, and Seasonality
Learn how to engineer time series features like lags, rolling windows, and seasonal indicators to give your forecasting models the temporal context they need to predict accurately.
Looking for something else?
Search every article by title, summary or topic.