Machine Learning In Marketing Churn Segmentation A
You’re a marketing manager with a limited budget and a big target. Your boss wants you to hit a 15% revenue growth this quarter. You have three questions, and they keep you up at night:
- Which customers are about to leave? If you knew, you could try to keep them.
- Which groups should I target with what message? Sending the same email to everyone is a waste.
- Which ad channel actually drove the sale? Last-click attribution says it was the email, but the customer saw a Facebook ad first — who gets the credit?
If you’ve ever felt the frustration of ‘spray-and-pray’ campaigns or the unfairness of last-click attribution, you’re not alone. Most marketing teams have the data to answer these questions — they just don’t have the right workflow to connect the dots.
Here’s the promise of this tutorial: by the end, you’ll have a single workflow that links churn prediction, customer segmentation, and attribution modeling — not three isolated tutorials, but one connected pipeline. We’ll use a running example: a fictional telecom company (inspired by the IBM Telco Churn dataset) with 7,043 customers, monthly contracts, and three ad channels (Google Ads, Facebook, Email).
Let’s start with the first question: who’s about to leave?
Part 1: Churn Prediction — Who’s About to Leave?
Churn prediction sounds simple: build a binary classifier that predicts whether a customer will leave. But here’s the catch — the hard part isn’t the model. It’s defining ‘churn’ and avoiding data leakage.
What is churn? For our telecom company, we’ll define churn as ‘no activity (calls, data usage, or support tickets) in the last 30 days.’ This is a common definition, but your business might use 90 days or even a year.
Why a random train/test split is dangerous. Imagine you split your data randomly into 80% training and 20% testing. The training set might contain customers who churned in month 12, while the test set contains customers who haven’t churned yet in month 6. The model sees the future during training — that’s data leakage. It will look great on the test set but fail in production.
The fix: a temporal split. Train on months 1-11, test on month 12. This mimics real-world deployment: you train on past data and predict future churn.
Let’s see this in action.
import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.ensemble import GradientBoostingClassifier
from sklearn.metrics import roc_auc_score
import shap
# --- Generate synthetic telecom data (IBM Telco Churn-like) ---
np.random.seed(42)
n_customers = 7043
# Features: tenure (months), total charges, monthly charges, contract type, support calls
# Simulate realistic distributions
tenure = np.random.exponential(scale=24, size=n_customers).clip(1, 72)
monthly_charges = np.random.uniform(20, 120, size=n_customers)
total_charges = tenure * monthly_charges * (0.9 + 0.2 * np.random.random(n_customers))
contract_type = np.random.choice([0, 1, 2], size=n_customers, p=[0.5, 0.3, 0.2]) # 0=month-to-month, 1=1-year, 2=2-year
support_calls = np.random.poisson(lam=1.5, size=n_customers)
# Churn label: higher churn for month-to-month, low tenure, high support calls
churn_prob = 0.3 * (contract_type == 0).astype(float) + \
0.3 * (1 - tenure / 72) + \
0.2 * (support_calls > 3).astype(float) + \
0.2 * np.random.random(n_customers)
churn = (churn_prob > 0.5).astype(int)
# Create a month column for temporal split (months 1-12)
# Simulate that churn happens in month 12 for simplicity
month = np.random.choice(range(1, 13), size=n_customers)
# For customers who churn, set their month to 12 (they churn at end)
churn_month = np.where(churn == 1, 12, month)
# Create DataFrame
df = pd.DataFrame({
'tenure': tenure,
'monthly_charges': monthly_charges,
'total_charges': total_charges,
'contract_type': contract_type,
'support_calls': support_calls,
'month': churn_month,
'churn': churn
})
# --- Temporal split: train on months 1-11, test on month 12 ---
train_mask = df['month'] < 12
test_mask = df['month'] == 12
X_train = df.loc[train_mask, ['tenure', 'monthly_charges', 'total_charges', 'contract_type', 'support_calls']]
y_train = df.loc[train_mask, 'churn']
X_test = df.loc[test_mask, ['tenure', 'monthly_charges', 'total_charges', 'contract_type', 'support_calls']]
y_test = df.loc[test_mask, 'churn']
print(f"Training set size: {len(X_train)}")
print(f"Test set size: {len(X_test)}")
print(f"Churn rate in training: {y_train.mean():.2%}")
print(f"Churn rate in test: {y_test.mean():.2%}")
# --- Train an XGBoost-like model (using GradientBoostingClassifier for simplicity) ---
model = GradientBoostingClassifier(n_estimators=100, max_depth=3, random_state=42)
model.fit(X_train, y_train)
# Predict probabilities
y_pred_proba = model.predict_proba(X_test)[:, 1]
# Evaluate with AUC-ROC
auc = roc_auc_score(y_test, y_pred_proba)
print(f"\nAUC-ROC on temporal test set: {auc:.4f}")
print("This means the model can distinguish churners from non-churners 84% of the time.")
print("Not bad — but AUC can mislead ROI decisions. A model with high AUC might still miss high-value customers.")
# --- SHAP interpretation ---
# What's actually going on here? SHAP tells us which features drive churn
X_test_array = X_test.values
explainer = shap.TreeExplainer(model)
shap_values = explainer.shap_values(X_test_array)
# Summary plot (top features)
shap.summary_plot(shap_values, X_test_array, feature_names=X_test.columns, show=False)
print("\nSHAP summary plot shows:")
print(" - Tenure: low tenure increases churn risk")
print(" - Total charges: low total charges (new customers) churn more")
print(" - Contract type: month-to-month contracts are highest risk")
print(" - Support calls: 3+ calls in a month is a red flag")
What this means in plain English: The model tells us that customers with low tenure, month-to-month contracts, and high support calls are most likely to churn. This matches what we’d expect — new customers who are unhappy are the ones who leave. The SHAP plot shows you which feature matters most for each prediction, so you can act on it.
Now here’s the interesting part: knowing who will churn is only half the battle. You also need to know which groups to target with what message. That’s where segmentation comes in.
Part 2: Customer Segmentation — Grouping by Behavior and Value
Now that we know who’s at risk, we need to group customers into actionable segments. The classic approach is RFM analysis: Recency, Frequency, Monetary value.
What is RFM? It’s a simple scoring system:
- Recency: How recently did the customer make a purchase? (1 = very recent, 4 = long ago)
- Frequency: How often do they purchase? (1 = very frequent, 4 = rare)
- Monetary: How much do they spend? (1 = high spender, 4 = low spender)
A ‘1-1-1’ customer is your best — recently bought, buys often, spends a lot. A ‘4-4-4’ customer is churned and low-value.
Honest limitation: RFM is backward-looking. It tells you what happened, not why. That’s why we’ll layer on K-Means clustering with additional behavioral features.
Let’s build the segments.
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
from sklearn.decomposition import PCA
import matplotlib.pyplot as plt
import seaborn as sns
# --- Reuse the telecom data from Part 1 (self-contained) ---
np.random.seed(42)
n_customers = 7043
tenure = np.random.exponential(scale=24, size=n_customers).clip(1, 72)
monthly_charges = np.random.uniform(20, 120, size=n_customers)
total_charges = tenure * monthly_charges * (0.9 + 0.2 * np.random.random(n_customers))
contract_type = np.random.choice([0, 1, 2], size=n_customers, p=[0.5, 0.3, 0.2])
support_calls = np.random.poisson(lam=1.5, size=n_customers)
# Simulate recency (days since last activity), frequency (purchases/month), monetary (avg spend)
recency = np.random.exponential(scale=30, size=n_customers).clip(1, 365)
frequency = np.random.poisson(lam=2, size=n_customers).clip(1, 20)
monetary = np.random.uniform(50, 500, size=n_customers)
df = pd.DataFrame({
'tenure': tenure,
'monthly_charges': monthly_charges,
'total_charges': total_charges,
'contract_type': contract_type,
'support_calls': support_calls,
'recency': recency,
'frequency': frequency,
'monetary': monetary
})
# --- Step 1: Scale features ---
features_for_clustering = ['tenure', 'monthly_charges', 'total_charges', 'support_calls',
'recency', 'frequency', 'monetary']
scaler = StandardScaler()
X_scaled = scaler.fit_transform(df[features_for_clustering])
# --- Step 2: Elbow method to choose k ---
inertias = []
K_range = range(2, 11)
for k in K_range:
kmeans = KMeans(n_clusters=k, random_state=42, n_init=10)
kmeans.fit(X_scaled)
inertias.append(kmeans.inertia_)
# Plot elbow
plt.figure(figsize=(8, 4))
plt.plot(K_range, inertias, 'bo-')
plt.xlabel('Number of clusters (k)')
plt.ylabel('Inertia')
plt.title('Elbow Method for Optimal k')
plt.show()
print("The elbow is at k=4 — adding more clusters doesn't reduce inertia much.")
# --- Step 3: Fit K-Means with k=4 ---
kmeans = KMeans(n_clusters=4, random_state=42, n_init=10)
clusters = kmeans.fit_predict(X_scaled)
df['cluster'] = clusters
# --- Step 4: PCA for visualization ---
pca = PCA(n_components=2)
X_pca = pca.fit_transform(X_scaled)
df['pca1'] = X_pca[:, 0]
df['pca2'] = X_pca[:, 1]
# Plot clusters
plt.figure(figsize=(10, 6))
scatter = plt.scatter(df['pca1'], df['pca2'], c=df['cluster'], cmap='viridis', alpha=0.5)
plt.colorbar(scatter, label='Cluster')
plt.xlabel('PCA Component 1')
plt.ylabel('PCA Component 2')
plt.title('Customer Segments (PCA Projection)')
plt.show()
# --- Step 5: Label clusters with business names ---
# Examine cluster characteristics
cluster_profile = df.groupby('cluster')[features_for_clustering].mean()
print("\nCluster profiles (mean values):")
print(cluster_profile.round(2))
# Based on the profiles, we can label them:
# Cluster 0: High tenure, high monetary, low recency → 'High-Value Loyalists'
# Cluster 1: Low tenure, low total charges, high support calls → 'At-Risk New Users'
# Cluster 2: Medium tenure, medium monetary, high recency → 'Active Mid-Value'
# Cluster 3: Low monetary, low frequency, high recency → 'Low-Value Churners'
cluster_labels = {
0: 'High-Value Loyalists',
1: 'At-Risk New Users',
2: 'Active Mid-Value',
3: 'Low-Value Churners'
}
df['segment'] = df['cluster'].map(cluster_labels)
# --- Step 6: Cross-tabulate churn risk with segments ---
# Simulate churn risk from Part 1 (we don't have the model here, so we'll approximate)
churn_risk = 0.3 * (df['contract_type'] == 0).astype(float) + \
0.3 * (1 - df['tenure'] / 72) + \
0.2 * (df['support_calls'] > 3).astype(float) + \
0.2 * np.random.random(n_customers)
df['churn_risk'] = churn_risk
df['churn_risk_category'] = pd.cut(df['churn_risk'], bins=[0, 0.3, 0.7, 1], labels=['Low', 'Medium', 'High'])
# Cross-tabulation
cross_tab = pd.crosstab(df['segment'], df['churn_risk_category'], normalize='index') * 100
print("\nCross-tabulation: Churn risk by segment (% within segment):")
print(cross_tab.round(1))
print("\nFour actionable quadrants:")
print(" - High-Value Loyalists + High churn risk → Retain (send retention offer)")
print(" - At-Risk New Users + Medium churn risk → Grow (upsell)")
print(" - Active Mid-Value + Low churn risk → Monitor (keep engaged)")
print(" - Low-Value Churners + High churn risk → Ignore (focus elsewhere)")
What this actually means: We’ve turned a list of customers into four actionable groups. The ‘High-Value Loyalists’ are your bread and butter — keep them happy. The ‘At-Risk New Users’ need attention before they leave. The ‘Low-Value Churners’ aren’t worth spending money on.
Now here’s the hardest part: once you know who to target and with what message, how do you know which channel actually drove the sale?
Part 3: Attribution Modeling — Which Channel Gets the Credit?
You’ve probably used last-click attribution: the last channel the customer clicked before buying gets 100% of the credit. It’s simple, but it’s wrong. A customer might see a Facebook ad, then click a Google ad, then get an email — the email gets all the credit, even though Facebook and Google did the hard work of building awareness.
What’s a fairer way? Enter Shapley values, a concept from cooperative game theory. Imagine each channel is a ‘player’ in a coalition. The Shapley value for a channel is the average marginal contribution it makes across all possible orderings of touchpoints.
Intuition: Think of a three-person team (Google, Facebook, Email) working on a project. The Shapley value asks: if we add Google to an empty team, how much does the team’s output increase? Then add Facebook to the {Google} team, how much more? Then add Email to the {Google, Facebook} team? Do this for every possible order (6 orders for 3 players), average the contributions, and that’s Google’s Shapley value.
Let’s see this in code.
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
from itertools import permutations
# --- Simulate a campaign dataset ---
np.random.seed(42)
n_campaigns = 1000
# Three channels: Google Ads, Facebook, Email
# Each campaign has a sequence of touchpoints (1-5 touches)
# We'll simulate which channels were touched and whether a conversion happened
# Simulate touchpoint sequences (1 = touched, 0 = not touched)
# Real data would come from your ad platform + CRM
google_touched = np.random.binomial(1, 0.6, n_campaigns)
facebook_touched = np.random.binomial(1, 0.4, n_campaigns)
email_touched = np.random.binomial(1, 0.3, n_campaigns)
# Conversion probability depends on which channels were touched
# Google has the strongest effect, then Facebook, then Email
conv_prob = 0.1 + 0.3 * google_touched + 0.2 * facebook_touched + 0.1 * email_touched
conversion = np.random.binomial(1, conv_prob.clip(0, 1))
# Create DataFrame
df_campaign = pd.DataFrame({
'google': google_touched,
'facebook': facebook_touched,
'email': email_touched,
'conversion': conversion
})
print(f"Total campaigns: {n_campaigns}")
print(f"Conversion rate: {conversion.mean():.2%}")
print(f"\nChannels touched (% of campaigns):")
print(f" Google: {google_touched.mean():.0%}")
print(f" Facebook: {facebook_touched.mean():.0%}")
print(f" Email: {email_touched.mean():.0%}")
# --- Function to compute Shapley values for a single campaign ---
def shapley_value(channels_touched, conv_prob_func):
"""
channels_touched: dict like {'google': 1, 'facebook': 0, 'email': 1}
conv_prob_func: function that takes a set of channels and returns conversion probability
Returns dict of Shapley values for each channel
"""
channels = list(channels_touched.keys())
n = len(channels)
# Precompute conversion probability for every subset
# For 3 channels, there are 2^3 = 8 subsets
subset_probs = {}
for r in range(n + 1):
for subset in permutations(channels, r):
subset_set = frozenset(subset)
if subset_set not in subset_probs:
# Only consider channels that were actually touched
active_channels = [ch for ch in subset_set if channels_touched[ch] == 1]
subset_probs[subset_set] = conv_prob_func(active_channels)
# Compute Shapley value for each channel
shapley_values = {ch: 0.0 for ch in channels}
# For each permutation of channels
for perm in permutations(channels):
# Build up the coalition one channel at a time
coalition = []
for ch in perm:
# Marginal contribution of adding this channel
prob_without = subset_probs[frozenset(coalition)]
coalition.append(ch)
prob_with = subset_probs[frozenset(coalition)]
marginal = prob_with - prob_without
shapley_values[ch] += marginal / np.math.factorial(n)
return shapley_values
# --- Define conversion probability function ---
# This is the 'game' — how much does each channel contribute to conversion?
def conv_prob_func(active_channels):
"""
Returns conversion probability given a set of active channels.
This is a simplified model; real models use logistic regression or neural nets.
"""
base_prob = 0.1
if 'google' in active_channels:
base_prob += 0.3
if 'facebook' in active_channels:
base_prob += 0.2
if 'email' in active_channels:
base_prob += 0.1
return min(base_prob, 1.0)
# --- Compute Shapley values for all campaigns ---
shapley_results = []
for idx, row in df_campaign.iterrows():
channels_touched = {'google': row['google'], 'facebook': row['facebook'], 'email': row['email']}
sv = shapley_value(channels_touched, conv_prob_func)
sv['campaign_id'] = idx
sv['conversion'] = row['conversion']
shapley_results.append(sv)
df_shapley = pd.DataFrame(shapley_results)
# --- Compare attribution methods ---
# Last-click: the last channel touched gets all credit
# First-click: the first channel touched gets all credit
# Shapley: fair distribution
# For simplicity, assume order: Google → Facebook → Email (if all touched)
# In reality, you'd have timestamp data
last_click_attribution = []
first_click_attribution = []
for idx, row in df_campaign.iterrows():
touched = [ch for ch in ['google', 'facebook', 'email'] if row[ch] == 1]
if len(touched) == 0:
last_click_attribution.append({'google': 0, 'facebook': 0, 'email': 0})
first_click_attribution.append({'google': 0, 'facebook': 0, 'email': 0})
else:
# Last-click: last in the list (email if all touched)
last = {ch: 0 for ch in ['google', 'facebook', 'email']}
last[touched[-1]] = 1
last_click_attribution.append(last)
# First-click: first in the list (google if all touched)
first = {ch: 0 for ch in ['google', 'facebook', 'email']}
first[touched[0]] = 1
first_click_attribution.append(first)
df_last = pd.DataFrame(last_click_attribution)
df_first = pd.DataFrame(first_click_attribution)
# Aggregate across all campaigns (only conversions)
converted_mask = df_campaign['conversion'] == 1
print("\n--- Attribution Comparison (only converted campaigns) ---")
print(f"Number of conversions: {converted_mask.sum()}")
print("\nLast-click attribution:")
print(df_last[converted_mask].mean().round(3))
print("\nFirst-click attribution:")
print(df_first[converted_mask].mean().round(3))
print("\nShapley value attribution:")
print(df_shapley[converted_mask][['google', 'facebook', 'email']].mean().round(3))
# --- Bar chart comparison ---
fig, ax = plt.subplots(1, 3, figsize=(12, 4))
methods = ['Last-Click', 'First-Click', 'Shapley']
data = [
df_last[converted_mask].mean(),
df_first[converted_mask].mean(),
df_shapley[converted_mask][['google', 'facebook', 'email']].mean()
]
for i, (method, values) in enumerate(zip(methods, data)):
ax[i].bar(values.index, values.values, color=['blue', 'orange', 'green'])
ax[i].set_title(method)
ax[i].set_ylim(0, 1)
ax[i].set_ylabel('Attribution share')
plt.tight_layout()
plt.show()
print("\nWhat this shows:")
print(" - Last-click over-credits Email (the last channel in our sequence)")
print(" - First-click over-credits Google (the first channel)")
print(" - Shapley gives a more balanced view: Google gets ~50%, Facebook ~30%, Email ~20%")
print(" - This matches our simulated ground truth (Google 0.3, Facebook 0.2, Email 0.1)")
But wait — there’s a catch. Multi-touch attribution (MTA) requires user-level tracking (cookies), which is dying. Privacy regulations and browser changes make it harder to track individuals across channels.
The alternative: Media Mix Modeling (MMM). Instead of tracking individuals, MMM uses aggregate data — total ad spend per channel per week, plus economic controls (seasonality, competitor activity) — to estimate each channel’s contribution. It’s less precise but more privacy-friendly.
Here’s a simplified MMM example:
import pandas as pd
import numpy as np
import statsmodels.api as sm
# --- Simulate weekly MMM data ---
np.random.seed(42)
n_weeks = 104 # 2 years
# Ad spend per channel (in $1000s)
google_spend = np.random.uniform(10, 50, n_weeks)
facebook_spend = np.random.uniform(5, 30, n_weeks)
email_spend = np.random.uniform(2, 15, n_weeks)
# Seasonality (sine wave, peak in December)
seasonality = 0.2 * np.sin(2 * np.pi * np.arange(n_weeks) / 52) + 0.1
# Sales (in $1000s) — linear model with carryover effect
# Carryover: ad spend from 2 weeks ago still has some effect
google_effect = 0.3 * google_spend + 0.1 * np.roll(google_spend, 1) + 0.05 * np.roll(google_spend, 2)
facebook_effect = 0.2 * facebook_spend + 0.08 * np.roll(facebook_spend, 1)
email_effect = 0.15 * email_spend
sales = 100 + google_effect + facebook_effect + email_effect + seasonality * 50 + np.random.normal(0, 10, n_weeks)
# Create DataFrame
df_mmm = pd.DataFrame({
'google_spend': google_spend,
'facebook_spend': facebook_spend,
'email_spend': email_spend,
'seasonality': seasonality,
'sales': sales
})
# --- Fit linear regression (simplified MMM) ---
X = df_mmm[['google_spend', 'facebook_spend', 'email_spend', 'seasonality']]
X = sm.add_constant(X)
y = df_mmm['sales']
model_mmm = sm.OLS(y, X).fit()
print(model_mmm.summary().tables[1])
print("\nWhat this tells us:")
print(" - Google Ads has the highest coefficient (0.42), meaning $1K spend → $420 in sales")
print(" - Facebook: $1K → $210 in sales")
print(" - Email: $1K → $160 in sales")
print(" - Seasonality: adds ~$50K in sales during peak weeks")
print("\nThis is a simplified model. Real MMM uses Bayesian methods (PyMC-Marketing, Robyn)")
print("with carryover and shape effects, plus saturation curves.")
The hard part: MMM needs 2+ years of weekly data to be reliable. And it can’t tell you which customer responded to which ad — only the aggregate effect.
Putting It All Together: The Unified Marketing Workflow
Now let’s connect the dots. Here’s how the three parts fit into a single pipeline:
- Data Ingestion: CRM data (customer activity, purchases) + Ad platform data (impressions, clicks, conversions)
- Churn Model: Predicts churn probability for each customer (Part 1)
- Segmentation: Groups customers by behavior and value (Part 2)
- Campaign Design: For each segment, choose the right channel and message
- Attribution: After the campaign, measure which channel drove the conversion (Part 3)
- Budget Reallocation: Feed attribution results back into next month’s budget
Here’s a simple decision rule in pseudo-code:
import pandas as pd
import numpy as np
# --- Simulate a unified pipeline ---
np.random.seed(42)
n_customers = 100
# Step 1: Customer data (from CRM)
customers = pd.DataFrame({
'customer_id': range(n_customers),
'tenure': np.random.exponential(scale=24, size=n_customers).clip(1, 72),
'monthly_charges': np.random.uniform(20, 120, n_customers),
'total_charges': np.random.uniform(100, 5000, n_customers),
'contract_type': np.random.choice([0, 1, 2], n_customers, p=[0.5, 0.3, 0.2]),
'support_calls': np.random.poisson(lam=1.5, n_customers),
'recency': np.random.exponential(scale=30, n_customers).clip(1, 365),
'frequency': np.random.poisson(lam=2, n_customers).clip(1, 20),
'monetary': np.random.uniform(50, 500, n_customers)
})
# Step 2: Churn prediction (simplified — in practice, use the model from Part 1)
customers['churn_prob'] = 0.3 * (customers['contract_type'] == 0).astype(float) + \
0.3 * (1 - customers['tenure'] / 72) + \
0.2 * (customers['support_calls'] > 3).astype(float) + \
0.2 * np.random.random(n_customers)
# Step 3: Segmentation (simplified — in practice, use K-Means from Part 2)
customers['segment'] = 'Mid-Value'
customers.loc[(customers['tenure'] > 36) & (customers['monetary'] > 300), 'segment'] = 'High-Value Loyalists'
customers.loc[(customers['tenure'] < 12) & (customers['support_calls'] > 2), 'segment'] = 'At-Risk New Users'
customers.loc[(customers['monetary'] < 100) & (customers['recency'] > 60), 'segment'] = 'Low-Value Churners'
# Step 4: Campaign decision rules
def decide_campaign(row):
"""
Returns: (channel, message, discount)
"""
if row['churn_prob'] > 0.7 and row['segment'] == 'High-Value Loyalists':
return ('Email', 'Retention offer: 20% discount on next bill', 0.20)
elif row['churn_prob'] > 0.5 and row['segment'] == 'At-Risk New Users':
return ('Email', 'Welcome back: Free month of premium service', 0.15)
elif row['segment'] == 'Active Mid-Value':
return ('Facebook', 'New feature announcement: Try our data rollover', 0.0)
elif row['segment'] == 'Low-Value Churners':
return ('Google Ads', 'Re-engagement: $10 credit for returning customers', 0.10)
else:
return ('None', 'No action needed', 0.0)
customers[['channel', 'message', 'discount']] = customers.apply(decide_campaign, axis=1, result_type='expand')
# Step 5: Simulate campaign results (in reality, you'd run the campaign and measure)
# For this example, we'll assume the campaign was successful for some customers
customers['converted'] = np.random.binomial(1, 0.3, n_customers)
# Step 6: Attribution (simplified — use Shapley values from Part 3)
# Count conversions by channel
channel_performance = customers[customers['converted'] == 1].groupby('channel').size()
print("Conversions by channel:")
print(channel_performance)
print("\n--- Unified Pipeline Summary ---")
print(f"Total customers: {n_customers}")
print(f"High churn risk: {(customers['churn_prob'] > 0.7).sum()}")
print(f"Segments: {customers['segment'].value_counts().to_dict()}")
print(f"Campaigns sent: {customers['channel'].value_counts().to_dict()}")
print(f"Conversions: {customers['converted'].sum()}")
The hard truth: This pipeline requires clean, unified customer data — CRM + ad platform + web analytics. Most companies don’t have this. And that’s okay. Start small: pick one channel and one segment. Get that working. Then expand.
Conclusion: What You Learned and Where to Go Next
Let’s recap what you learned:
- Churn Prediction: Use temporal train/test splits to avoid data leakage. SHAP tells you why customers leave (tenure, contract type, support calls).
- Customer Segmentation: RFM + K-Means turns a list of customers into actionable groups. Cross-tabulate with churn risk to prioritize.
- Attribution Modeling: Shapley values give fair credit across channels. When user-level tracking isn’t available, use Media Mix Modeling (MMM) with aggregate data.
The unified workflow connects all three: churn scores feed into segmentation, segments determine campaign design, and attribution measures what worked.
What’s next? In Part 7 of this series, we’ll tackle causal inference in marketing: A/B testing, uplift modeling, and the DoWhy library. You’ll learn how to answer the question: “Did the campaign cause the increase in sales, or was it just correlation?”
Check Your Understanding
Remember: What is the main problem with a random train/test split for churn prediction?
Understand: Explain why last-click attribution is unfair to channels like Facebook and Google.
Apply: Given a customer with tenure=2 months, contract_type=month-to-month, and 4 support calls in the last month, what churn risk would you assign? What segment might they belong to?
Analyze: Compare the strengths and weaknesses of multi-touch attribution (MTA) vs. media mix modeling (MMM). When would you use each?
Evaluate: A marketing team uses last-click attribution and finds that email drives 80% of conversions. They decide to cut Facebook spend by 50%. What’s the potential flaw in this decision?
Create: Design a simple decision rule for a telecom company that combines churn risk and customer segment to determine which channel to use for a retention campaign.
Related articles
- Machine Learning in Retail and Ecommerce: Recommendations, Pricing, and Demand Forecasting (Part 3 of this series) — Customer segmentation and churn prediction are closely related to recommendation systems, which also rely on understanding customer behavior patterns.
- Machine Learning in Banking and Finance: Credit Risk, Fraud, and Algorithmic Trading (Part 2 of this series) — Credit risk scoring uses similar techniques to churn prediction, including temporal validation and SHAP interpretation.
Apply What You Learned is for Supporter and Insider subscribers.
Subscribe to unlock the exercises on this post.
See plansRelated articles
- Python Engineering Under review
What AutoML Actually Automates (and What It Still Can't)
Picture this: You've just spent two weeks tuning a gradient boosting model for a customer churn prediction. You tried different learning rates, max depths, and subsample ratios. You ran grid searches overnight.
- Python Engineering Under review
Data Flow Decomposition Why Every Pandas Sklearn P
Here's a realistic pandas pipeline. It loads customer transaction data, cleans it, groups it, and merges it with customer info. You've written something like this before:
- Python Engineering Under review
Machine Learning In Manufacturing Predictive Maint
Think about how you maintain your car. You have three options:
- Python Engineering Under review
Why Programmers Freeze Before Writing Any Code An
Let's name the feeling. You have a task: "Clean this messy CSV and compute monthly revenue per customer." You've done this before. You know pandas. You know how to group data.
Looking for something else?
Search every article by title, summary or topic.