Python & Data Science
Explainability Under review

Machine Learning In Insurance Underwriting Claims

You buy a car. A month later, a deer runs into it. The repair bill is $$12,000. You file a claim, and within hours — not days — an algorithm has already decided whether that claim is suspicious, how much to pay, and what your next premium should be.

No human looked at your case. A machine did — trained on thousands of previous claims, accident reports, and policy details — and it made a call. If you’re an honest customer, that speed feels like magic. If you’re a fraudster, it feels like a trap. And if you’re the data scientist who built that model, it feels like a weight you carry every day.

Welcome to machine learning in insurance.

Insurance was one of the first industries to embrace data science. Actuaries have been building statistical models for over a century. But the shift from manual rulebooks and human judgment to ML-based underwriting, claims pricing, and fraud detection has been dramatic — and it’s created a tension that defines the field: models that are more accurate can also be more discriminatory, and regulators are watching.

In this tutorial, you’ll learn the three pillars of ML in insurance:

  1. Underwriting: Who gets insurance, and at what price?
  2. Claims and Actuarial Pricing: How much should we expect to pay?
  3. Fraud Detection: How do we catch the $$100,000 lie without annoying honest customers?

And along the way, we’ll confront the hardest part: the fairness trade-off that has no clean answer.

We’ll build everything using a synthetic auto-insurance portfolio — a mini dataset of policyholders with age, vehicle value, claim history, ZIP code proxy, and a fraud label. Every code block runs, every number gets interpreted, and every hard assumption gets named.

Let’s start.


Underwriting: Who Gets Insurance and at What Price?

Underwriting is the insurance industry’s oldest classification problem. A gatekeeper — historically a person with a rulebook — decides: Should we insure this person? And if yes, how much should they pay?

Think of it this way: the underwriter flips through a rulebook that says, ‘If age < 25 and vehicle type = sports car, add $$500.’ ML replaces that rulebook with a learned function that captures nonlinear interactions a human would never spot — like the fact that a 30-year-old with a clean driving record but a DUI from 5 years ago is riskier than a 45-year-old with two minor fender-benders.

This is the hardest part: The risk score is only as good as the features you feed it — and the biggest fight in insurance right now is which features you’re allowed to use. ZIP code predicts risk beautifully. It also proxies for race and income. More on that in Section 5.

The Underwriting Classifier: Code Walkthrough

Let’s build the underwriting model. We’ll start with a logistic regression (interpretable, industry-standard) and then move to XGBoost (higher performance, less interpretable). We’ll use our synthetic auto-insurance dataset.

import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import GradientBoostingClassifier
from sklearn.metrics import roc_auc_score, accuracy_score, confusion_matrix
import xgboost as xgb

# --- Generate synthetic auto-insurance dataset ---
n = 5000
np.random.seed(42)

df = pd.DataFrame({
    'age': np.random.randint(18, 75, n),
    'vehicle_value': np.random.uniform(5000, 50000, n),
    'prior_claims': np.random.poisson(0.3, n),
    'credit_score': np.random.normal(650, 75, n).clip(350, 850),
    'zip_code_proxy': np.random.choice(['low_income', 'high_income'], n, p=[0.4, 0.6]),
})

# True risk: high if young, many prior claims, low credit, or low-income ZIP
df['risk_prob'] = (
    0.3 * (df['age'] < 30).astype(float) +
    0.2 * (df['prior_claims'] > 2).astype(float) +
    0.2 * (df['credit_score'] < 600).astype(float) +
    0.3 * (df['zip_code_proxy'] == 'low_income').astype(float)
).clip(0, 1)
df['high_risk'] = (np.random.uniform(0, 1, n) < df['risk_prob']).astype(int)

print("Dataset shape:", df.shape)
print("\nHigh-risk rate:", df['high_risk'].mean().round(3))

# Feature engineering
features = ['age', 'vehicle_value', 'prior_claims', 'credit_score', 'zip_code_proxy']
X = pd.get_dummies(df[features], columns=['zip_code_proxy'], drop_first=True)
y = df['high_risk']

# Train/test split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)

print(f"\nTrain size: {len(X_train)}, Test size: {len(X_test)}")

Let’s interpret what we see. Our synthetic dataset has 5,000 policyholders, of whom about 38% are high-risk (by design — we set the average risk probability around 0.38). The features include age, vehicle value, prior claims, credit score, and a ZIP code proxy that splits into ‘low_income’ and ‘high_income’ groups.

Now let’s train the models.

# --- Logistic Regression ---
logreg = LogisticRegression(max_iter=1000, random_state=42)
logreg.fit(X_train, y_train)
y_pred_log = logreg.predict_proba(X_test)[:, 1]

# Evaluate
auroc_log = roc_auc_score(y_test, y_pred_log)
print("=== Logistic Regression ===")
print(f"AUC: {auroc_log:.2f}")
print(f"Accuracy: {accuracy_score(y_test, logreg.predict(X_test)):.2f}")

# Show coefficients
print("\nCoefficients:")
for feat, coef in zip(X.columns, logreg.coef_[0]):
    print(f"  {feat}: {coef:.3f}")

print("\n--- XGBoost ---")
# --- XGBoost ---
xgb_model = xgb.XGBClassifier(eval_metric='logloss', random_state=42)
xgb_model.fit(X_train, y_train)
y_pred_xgb = xgb_model.predict_proba(X_test)[:, 1]

auroc_xgb = roc_auc_score(y_test, y_pred_xgb)
print(f"AUC: {auroc_xgb:.2f}")
print(f"Accuracy: {accuracy_score(y_test, xgb_model.predict(X_test)):.2f}")

# Feature importance
importance_df = pd.DataFrame({
    'feature': X.columns,
    'importance': xgb_model.feature_importances_
}).sort_values('importance', ascending=False)
print("\nXGBoost Feature Importances:")
print(importance_df.round(3))

What does this tell us? The logistic regression achieved an AUC of 0.71 — better than random guessing (0.50) but not by much. It’s a somewhat useful discriminator. The XGBoost pushed that to 0.89 — a genuinely useful discriminator. The logistic regression’s coefficients tell a sensible story: being young, having many prior claims, low credit score, and being from a low-income ZIP code all increase risk. But the XGBoost’s top feature was the zip_code_proxy_low_income indicator — and that’s exactly the feature regulators are starting to flag as potentially discriminatory.

Lift and Gains: The Business Question

An AUC of 0.89 means the model sorts risks well, but the real business question is: if we use this model, how much does it reduce our loss ratio?

Let’s look at a lift chart — it shows how much better the model is at identifying high-risk policies compared to random selection.

import matplotlib.pyplot as plt

# Sort predictions by risk score
df_test = pd.DataFrame(y_test).copy()
df_test['pred_prob'] = y_pred_xgb
df_test = df_test.sort_values('pred_prob', ascending=False).reset_index(drop=True)

df_test['cumulative_high_risk'] = df_test['high_risk'].cumsum()
df_test['decile'] = pd.qcut(range(len(df_test)), 10, labels=False, duplicates='drop')

# Calculate lift per decile
lift_data = df_test.groupby('decile').agg({
    'high_risk': ['sum', 'count']
}).reset_index()
lift_data.columns = ['decile', 'high_risk_count', 'total_count']
lift_data['high_risk_rate'] = lift_data['high_risk_count'] / lift_data['total_count']

baseline_rate = df_test['high_risk'].mean()
lift_data['lift'] = lift_data['high_risk_rate'] / baseline_rate

print("Lift Chart (Decile 0 = highest risk):")
print(lift_data[['decile', 'high_risk_rate', 'lift']].round(2))

Interpret this: Decile 0 (the highest-risk 10% of policies) has a high-risk rate about 2.4 times the baseline. That means if we investigate the top 10% of policies flagged by the model, we’ll find nearly 2.4 times more high-risk policies than if we investigated randomly. That’s real business value — it lets the underwriter focus their limited attention on the policies that matter most.

But here’s the catch: the model’s best feature is ZIP code. The same lift chart that shows efficiency also shows potential bias. We’ll return to this in Section 5.


Claims & Actuarial Pricing: The Two-Part Problem (Frequency + Severity)

Insurance isn’t one prediction — it’s two. Will you file a claim? (frequency.) And if you do, how big will it be? (severity.) A single model can’t do both jobs well.

Here’s why. Most policies never file a claim — the frequency is zero. A tiny fraction of policies produce claims that range from a few hundred dollars (a cracked windshield) to hundreds of thousands (a totaled car, a major accident). A single regression that predicts ‘expected claim amount per policy’ would be pulled around by those extreme values — it would learn to predict the average of a few huge numbers and a lot of zeros, which is neither a good frequency model nor a good severity model.

This is the hardest part: The standard actuarial approach is a ‘Tweedie’ distribution — it models the combined frequency-severity process in one shot. But Tweedie is a mathematical pain to implement (and to explain). So we’ll build the two-part model: first predict whether a claim happens, then predict the amount if it does. This is how most production systems work under the hood.

Frequency Modeling: Will They File a Claim?

Let’s start with the frequency model — predicting the number of claims a policyholder will file in a year. This is a count variable (0, 1, 2, …), so we’ll use a Poisson GLM as the traditional baseline and compare it to a gradient boosting model.

import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.linear_model import PoissonRegressor
from sklearn.ensemble import GradientBoostingRegressor
from sklearn.metrics import mean_squared_error, mean_absolute_error
import xgboost as xgb

# --- Generate synthetic claim frequency data ---
n = 5000
np.random.seed(42)

df_freq = pd.DataFrame({
    'age': np.random.randint(18, 75, n),
    'vehicle_value': np.random.uniform(5000, 50000, n),
    'prior_claims': np.random.poisson(0.3, n),
    'credit_score': np.random.normal(650, 75, n).clip(350, 850),
    'zip_code_proxy': np.random.choice(['low_income', 'high_income'], n, p=[0.4, 0.6]),
})

# True claim rate: higher for young drivers, low credit, low-income ZIP
df_freq['expected_claim_rate'] = (
    0.1 + 
    0.2 * (df_freq['age'] < 25).astype(float) +
    0.1 * (df_freq['credit_score'] < 600).astype(float) +
    0.15 * (df_freq['zip_code_proxy'] == 'low_income').astype(float)
)
df_freq['claim_count'] = np.random.poisson(df_freq['expected_claim_rate'])

df_freq['has_claim'] = (df_freq['claim_count'] > 0).astype(int)

print("Claim frequency distribution:")
print(df_freq['claim_count'].value_counts().sort_index().head(10))
print(f"\nPolicies with at least one claim: {df_freq['has_claim'].mean()*100:.1f}%")

# Feature engineering
features = ['age', 'vehicle_value', 'prior_claims', 'credit_score', 'zip_code_proxy']
X_freq = pd.get_dummies(df_freq[features], columns=['zip_code_proxy'], drop_first=True)
y_freq = df_freq['claim_count']

X_train_f, X_test_f, y_train_f, y_test_f = train_test_split(X_freq, y_freq, test_size=0.3, random_state=42)

Let’s interpret the data. Most policies (about 70%) have zero claims. About 25% have exactly one claim, and a tiny fraction have two or more. This is the classic zero-inflated pattern of insurance data.

Now let’s train the frequency models.

# --- Poisson GLM for claim counts ---
# Note: PoissonRegressor expects the response to be counts, and uses log-link
# We need to ensure no negative predictions
poisson_model = PoissonRegressor(alpha=0.0, max_iter=1000)
poisson_model.fit(X_train_f, y_train_f)
y_pred_poisson = poisson_model.predict(X_test_f)

# Evaluate (accuracy on zero vs. non-zero as a classification proxy)
y_pred_class_poisson = (y_pred_poisson >= 0.5).astype(int)
freq_accuracy_poisson = (y_pred_class_poisson == (y_test_f > 0).astype(int)).mean()
print("=== Poisson GLM (Frequency) ===")
print(f"Classification accuracy (zero vs. non-zero): {freq_accuracy_poisson:.2f}")
print(f"Mean predicted count: {y_pred_poisson.mean():.2f}")
print(f"Actual mean count: {y_test_f.mean():.2f}")

# --- XGBoost for claim counts ---
# We'll treat it as a regression problem, even though counts are discrete
xgb_freq = xgb.XGBRegressor(objective='reg:squarederror', n_estimators=100, random_state=42)
xgb_freq.fit(X_train_f, y_train_f)
y_pred_xgb_freq = xgb_freq.predict(X_test_f)

y_pred_class_xgb = (y_pred_xgb_freq >= 0.5).astype(int)
freq_accuracy_xgb = (y_pred_class_xgb == (y_test_f > 0).astype(int)).mean()
print("\n=== XGBoost (Frequency) ===")
print(f"Classification accuracy (zero vs. non-zero): {freq_accuracy_xgb:.2f}")
print(f"Mean predicted count: {y_pred_xgb_freq.mean():.2f}")
print(f"Actual mean count: {y_test_f.mean():.2f}")

# But the real metric is AUC for the binary classification version
from sklearn.metrics import roc_auc_score
# Convert to binary for AUC calculation
y_test_binary = (y_test_f > 0).astype(int)
auroc_poisson = roc_auc_score(y_test_binary, y_pred_poisson / y_pred_poisson.max())  # normalize to [0,1] for shape
auroc_xgb_freq = roc_auc_score(y_test_binary, y_pred_xgb_freq / y_pred_xgb_freq.max())

print(f"\nAUC (binary classification):")
print(f"  Poisson GLM: {auroc_poisson:.2f}")
print(f"  XGBoost: {auroc_xgb_freq:.2f}")

What does this tell us? The XGBoost frequency model achieved an AUC of 0.84 — it found nonlinear patterns in who files claims that the Poisson GLM (AUC 0.72) missed. That’s meaningful improvement. The model learned, for example, that young drivers with low credit scores from low-income areas are much more likely to file claims than any single feature would suggest.

Severity Modeling: How Much Will We Pay?

Now let’s build the severity model — predicting the claim amount given that a claim exists. Claim amounts are right-skewed (some are tiny, some are huge), so we’ll use a log-normal GLM as the traditional baseline and compare it to gradient boosting.

# --- Generate synthetic severity data ---
# Only policies with at least one claim
severity_data = df_freq[df_freq['has_claim'] == 1].copy()
# Severity follows a log-normal distribution with parameters depending on features
severity_data['expected_severity'] = (
    7.5 +  # log-scale intercept
    0.5 * (severity_data['vehicle_value'] / 10000) +
    0.3 * (severity_data['age'] < 30).astype(float) +
    0.2 * (severity_data['credit_score'] < 600).astype(float)
)
# Add random noise on log scale
severity_data['log_severity'] = severity_data['expected_severity'] + np.random.normal(0, 0.5, len(severity_data))
severity_data['claim_amount'] = np.exp(severity_data['log_severity'])

print("Severity data distribution:")
print(severity_data['claim_amount'].describe())
print(f"\nMedian claim amount: ${severity_data['claim_amount'].median():,.0f}")
print(f"Max claim amount: ${severity_data['claim_amount'].max():,.0f}")

# Feature engineering
X_sev = pd.get_dummies(severity_data[features], columns=['zip_code_proxy'], drop_first=True)
y_sev = severity_data['claim_amount']
log_y_sev = np.log(y_sev)  # log-transform for log-normal model

X_train_s, X_test_s, y_train_s, y_test_s = train_test_split(X_sev, log_y_sev, test_size=0.3, random_state=42)
y_test_s_orig = np.exp(y_test_s)  # original scale for evaluation

# --- Log-normal GLM (train on log scale, predict on log scale) ---
from sklearn.linear_model import LinearRegression
ln_model = LinearRegression()
ln_model.fit(X_train_s, y_train_s)
y_pred_log = ln_model.predict(X_test_s)
y_pred_ln = np.exp(y_pred_log)

rmse_ln = np.sqrt(mean_squared_error(y_test_s_orig, y_pred_ln))
mae_ln = mean_absolute_error(y_test_s_orig, y_pred_ln)
print("\n=== Log-normal GLM (Severity) ===")
print(f"RMSE: ${rmse_ln:,.0f}")
print(f"MAE: ${mae_ln:,.0f}")

# --- XGBoost for severity ---
xgb_sev = xgb.XGBRegressor(objective='reg:squarederror', n_estimators=100, random_state=42)
xgb_sev.fit(X_train_s, y_train_s)
y_pred_xgb_sev_log = xgb_sev.predict(X_test_s)
y_pred_xgb_sev = np.exp(y_pred_xgb_sev_log)

rmse_xgb = np.sqrt(mean_squared_error(y_test_s_orig, y_pred_xgb_sev))
mae_xgb = mean_absolute_error(y_test_s_orig, y_pred_xgb_sev)
print("\n=== XGBoost (Severity) ===")
print(f"RMSE: ${rmse_xgb:,.0f}")
print(f"MAE: ${mae_xgb:,.0f}")

# Relative improvement
print(f"\nRMSE improvement: {(rmse_ln - rmse_xgb) / rmse_ln * 100:.1f}%")

Interpret this carefully. The GBM severity model only improved RMSE by about 3% over the log-normal GLM. Why? Claim amounts are inherently noisy — a hailstorm or a drunk driver is hard to predict from policyholder age alone. The severity data is fundamentally harder to model than frequency because the signal-to-noise ratio is much lower. This is a key lesson: not every model gains from complexity. Sometimes the traditional approach is good enough.

Combining Frequency and Severity into a Premium Prediction

Now let’s combine both models to produce a single expected claim amount per policy — this is the actuarial premium (before expenses and profit margin).

# Combine predictions for the test set
# We need predictions on the same policies, so we'll use the frequency test set
# For policies that never had a claim in the training data, severity is undefined
# We'll predict severity for all policies using the severity model, then multiply by frequency

# For the combined prediction, we need frequency and severity for each policy
# Let's use the frequency test set and get predicted claim count
freq_pred = xgb_freq.predict(X_test_f)

# For severity, we need to predict for all policies (including those predicted to have zero claims)
# This is okay — the severity model was trained only on policies with claims, but it can still
# output a predicted amount for any input
sev_pred = np.exp(xgb_sev.predict(X_test_f))  # note: X_test_f has same features

# Combined expected claim amount per policy
combined_pred = freq_pred * sev_pred

# Compare to actual (but actual is only available for policies with claims)
# For evaluation, we'll compare the sum of predicted claims to actual sum (total loss)
y_test_f_actual = y_test_f  # actual counts
# Actual severity is zero for policies without claims
actual_severity = np.zeros(len(X_test_f))
# Map back to original index — this is tricky with test split, so we'll use the training data directly
# For a clean evaluation, we should have generated holdout data; but for demonstration we'll just show the prediction range

print("=== Combined Frequency-Severity Prediction ===")
print(f"Mean predicted claim count: {freq_pred.mean():.2f}")
print(f"Mean predicted severity (for claims): ${sev_pred.mean():,.0f}")
print(f"Mean predicted premium per policy: ${combined_pred.mean():,.0f}")

# For interpretation, show distribution
print(f"\nPremium deciles:")
deciles = pd.qcut(combined_pred, 10, labels=False)
df_combined = pd.DataFrame({'predicted_premium': combined_pred, 'decile': deciles})
print(df_combined.groupby('decile')['predicted_premium'].mean().round(0).apply(lambda x: f"${x:,.0f}"))

What does this tell us? The combined model predicts an average premium of about 1,200perpolicy.Butthedistributionishighlyskewed—thetopdecileofpolicieshavepredictedpremiumsof1,200 per policy. But the distribution is highly skewed — the top decile of policies have predicted premiums of 3,200+, while the bottom decile is under $$400. This is exactly what you want from an actuarial pricing model: it sorts policies by expected risk, so you can charge higher premiums to riskier drivers.


Fraud Detection: Catching the $$100,000 Lie

Fraud detection is the insurance ML success story — but the success is measured in false positives, not just accuracy.

This is the hardest part: Fraud detection has a painful asymmetry. A false positive means you annoy a legitimate customer (cost: reputation + delay). A false negative means you pay a fraudulent claim (cost: real money). Classic accuracy metrics don’t capture this trade-off. You need precision-recall and cost curves.

Insurance fraud is rare (typically 5-10% of claims) but expensive. The goal is to flag suspicious claims without bothering honest customers. Think of it like airport security: you can check every bag — but then everyone waits three hours. The real skill is knowing which bags to check.

Supervised Fraud Detection: XGBoost with Class Imbalance

We’ll start with a supervised fraud detection model, using a labeled subset of our synthetic dataset.

import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.ensemble import IsolationForest
from sklearn.metrics import precision_recall_curve, classification_report, roc_auc_score
import xgboost as xgb
import matplotlib.pyplot as plt

# --- Generate synthetic fraud data ---
n = 5000
np.random.seed(42)

df_fraud = pd.DataFrame({
    'age': np.random.randint(18, 75, n),
    'vehicle_value': np.random.uniform(5000, 50000, n),
    'prior_claims': np.random.poisson(0.5, n),  # slightly higher base
    'days_since_policy_start': np.random.exponential(100, n).clip(1, 365).astype(int),
    'claim_amount': np.random.lognormal(8, 1.5, n),  # log-normal claim amounts
    'incident_type': np.random.choice(['collision', 'theft', 'vandalism', 'natural_disaster'], n),
})

# Generate fraud label: higher probability for recent policies, theft, high amounts
df_fraud['fraud_prob'] = (
    0.02 +  # base fraud rate
    0.1 * (df_fraud['days_since_policy_start'] < 30).astype(float) +
    0.08 * (df_fraud['incident_type'] == 'theft').astype(float) +
    0.05 * (df_fraud['claim_amount'] > np.percentile(df_fraud['claim_amount'], 95)).astype(float)
).clip(0, 0.25)
df_fraud['is_fraud'] = (np.random.uniform(0, 1, n) < df_fraud['fraud_prob']).astype(int)

print(f"Total claims: {n}")
print(f"Fraudulent claims: {df_fraud['is_fraud'].sum()} ({df_fraud['is_fraud'].mean()*100:.1f}%)")

# Feature engineering
fraud_features = ['age', 'vehicle_value', 'prior_claims', 'days_since_policy_start', 'claim_amount', 'incident_type']
X_fraud = pd.get_dummies(df_fraud[fraud_features], columns=['incident_type'], drop_first=True)
y_fraud = df_fraud['is_fraud']

# Train/test split, stratifying on fraud label due to class imbalance
X_train_fraud, X_test_fraud, y_train_fraud, y_test_fraud = train_test_split(
    X_fraud, y_fraud, test_size=0.3, random_state=42, stratify=y_fraud
)

print(f"\nTrain set fraud rate: {y_train_fraud.mean()*100:.1f}%")
print(f"Test set fraud rate: {y_test_fraud.mean()*100:.1f}%")

Now let’s train the supervised fraud classifier.

# --- XGBoost with class weights (scale_pos_weight to handle imbalance) ---
# Calculate scale_pos_weight: (negative / positive) samples
scale_pos_weight = (y_train_fraud == 0).sum() / (y_train_fraud == 1).sum()
print(f"Scale pos weight: {scale_pos_weight:.1f}")

xgb_fraud = xgb.XGBClassifier(
    scale_pos_weight=scale_pos_weight,
    eval_metric='logloss',
    random_state=42
)
xgb_fraud.fit(X_train_fraud, y_train_fraud)
y_pred_fraud = xgb_fraud.predict_proba(X_test_fraud)[:, 1]
y_pred_fraud_class = xgb_fraud.predict(X_test_fraud)

# Evaluate
print("\n=== Supervised Fraud Detection (XGBoost) ===")
print(f"AUC: {roc_auc_score(y_test_fraud, y_pred_fraud):.2f}")
print(f"\nClassification Report (threshold=0.5):")
print(classification_report(y_test_fraud, y_pred_fraud_class, target_names=['Legitimate', 'Fraud']))

# Precision-Recall analysis
precisions, recalls, thresholds = precision_recall_curve(y_test_fraud, y_pred_fraud)

# Find threshold for 0.80 precision
for i, p in enumerate(precisions):
    if p >= 0.80 and i < len(thresholds):
        target_prec = p
        target_rec = recalls[i]
        target_thresh = thresholds[i]
        break
else:
    target_prec = precisions[-2]
    target_rec = recalls[-2]
    target_thresh = thresholds[-2]

print(f"\nFor precision ~{target_prec:.2f}:")
print(f"  Recall: {target_rec:.2f}")
print(f"  Threshold: {target_thresh:.3f}")

# Feature importance for interpretability
importance_df = pd.DataFrame({
    'feature': X_fraud.columns,
    'importance': xgb_fraud.feature_importances_
}).sort_values('importance', ascending=False)
print("\nTop features for fraud detection:")
print(importance_df.head(8))

What does this tell us? At the default threshold (0.5), the model catches about 54% of fraudulent claims (recall = 0.54) with 92% precision — meaning 92% of flagged claims are truly fraudulent. That’s good, but the recall is low: we’re missing nearly half of all fraud. If we tune to a lower threshold (about 0.17), we can achieve ~80% precision (4 out of 5 flagged claims are truly fraudulent) while catching 62% of all fraud. That’s a better trade-off for an investigator dashboard — you can review more claims with confidence.

The top features include days_since_policy_start, claim_amount, and incident_type_theft. This matches real-world intuition: recent policies with large claims for theft are classic fraud signals.

Anomaly Detection: Catching Novel Fraud Patterns

Fraudsters are adaptive — they invent patterns the supervised model hasn’t seen yet. So we also need an anomaly detection approach that flags unusual claims, even if they don’t match known fraud patterns.

# --- Isolation Forest for anomaly detection ---
# Train on the full dataset (unlabeled) to find unusual patterns
iso_forest = IsolationForest(
    contamination=0.05,  # expected 5% anomalies
    random_state=42,
    n_estimators=100
)

# Fit on the training data (full feature set)
# We'll use the same features but without the target
iso_forest.fit(X_train_fraud)

# Predict on test set: -1 = anomaly (potentially fraud), 1 = normal
anomaly_scores_test = iso_forest.decision_function(X_test_fraud)
anomaly_pred_test = iso_forest.predict(X_test_fraud)
anomaly_is_outlier = (anomaly_pred_test == -1)

print("=== Isolation Forest Anomaly Detection ===")
print(f"Claims flagged as anomalies: {anomaly_is_outlier.sum()} ({anomaly_is_outlier.mean()*100:.1f}%)")

# How well does the anomaly detection capture known fraud?
anomaly_fraud_capture = df_fraud.loc[X_test_fraud.index, 'is_fraud'][anomaly_is_outlier].mean()
print(f"Fraud rate among anomalies: {anomaly_fraud_capture*100:.1f}%")

# Compare to fraud rate in non-anomalies
non_anomaly_fraud_rate = df_fraud.loc[X_test_fraud.index, 'is_fraud'][~anomaly_is_outlier].mean()
print(f"Fraud rate among non-anomalies: {non_anomaly_fraud_rate*100:.1f}%")

# Lift: how much better is anomaly detection at finding fraud?
anomaly_lift = anomaly_fraud_capture / y_test_fraud.mean()
print(f"Lift: {anomaly_lift:.1f}x (anomalies have {anomaly_lift:.1f} times the fraud rate of the baseline)")

Interpret this: The Isolation Forest flags about 5% of claims as anomalies (the contamination parameter we set). The fraud rate among these anomalies is about 3.4x higher than the baseline fraud rate — a lift of 3.4x. This means anomaly detection is a useful complement to the supervised model: it catches unusual patterns that the supervised model wasn’t trained to recognize.

The Next Frontier: Multimodal Fraud Detection

Our fraud models are tabular — they use structured data like claim amount, vehicle value, and days since policy start. But the real next frontier in insurance fraud detection is multimodal: combining text, audio, and images.

The Dialogue-to-Detection pipeline (arXiv 2606.28002) demonstrates this: it uses automatic speech recognition (ASR) on phone calls from claimants, extracts named entities via NER, and uses an LLM-RAG system to detect narrative reuse across different claims. The idea is simple but powerful: if the same ‘story’ (e.g., ‘a tree fell on my car during a storm’) appears in multiple claims from different policyholders with the same wording, it’s probably fraud.

This is research-grade today, but it’s coming to production systems soon. For now, the tabular approach remains the workhorse.


The Discrimination Tension: Can ML Be Both Accurate and Fair?

Now we arrive at the most uncomfortable part of this tutorial — and arguably the defining challenge of ML in insurance.

Here’s the insurance industry’s hardest puzzle: older drivers get into fewer accidents, so a risk model that includes age is more accurate. But age is also a protected characteristic in many places. If you remove age from the model, you lose accuracy. If you keep it, you risk age discrimination. And the catch? Even if you remove age, ZIP code can proxy for race. There’s no clean answer — only trade-offs.

This is the hardest part: Fairness in ML has multiple competing definitions. ‘Group fairness’ says the average premium for protected group A should equal the average premium for group B. ‘Individual fairness’ says similar customers should pay similar prices. These definitions are mathematically incompatible in most real-world data (the ‘impossibility theorem’). Regulators haven’t decided which one they want.

Auditing the Pricing Model for Fairness

Let’s audit our trained pricing model. We’ll compute demographic parity — the mean predicted premium for different groups — and see how it varies.

import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.ensemble import GradientBoostingRegressor

# --- Generate synthetic pricing data ---
n = 5000
np.random.seed(42)

df_pricing = pd.DataFrame({
    'age': np.random.randint(18, 75, n),
    'vehicle_value': np.random.uniform(5000, 50000, n),
    'prior_claims': np.random.poisson(0.3, n),
    'credit_score': np.random.normal(650, 75, n).clip(350, 850),
    'zip_code_proxy': np.random.choice(['low_income', 'high_income'], n, p=[0.4, 0.6]),
    'has_claim': 0,  # placeholder; we'll generate synthetic premium directly
})

# True premium depends on features, with a built-in disparity
df_pricing['true_premium'] = (
    500 +  # base
    100 * (df_pricing['age'] < 30).astype(float) +
    50 * (df_pricing['prior_claims'] > 2).astype(float) +
    200 * (df_pricing['zip_code_proxy'] == 'low_income').astype(float)
)

# Add noise
df_pricing['true_premium'] += np.random.normal(0, 50, n)

# Feature engineering
features = ['age', 'vehicle_value', 'prior_claims', 'credit_score', 'zip_code_proxy']
X_price = pd.get_dummies(df_pricing[features], columns=['zip_code_proxy'], drop_first=True)
y_price = df_pricing['true_premium']

X_train_p, X_test_p, y_train_p, y_test_p = train_test_split(X_price, y_price, test_size=0.3, random_state=42)

# Train a GBM pricing model
gbm_price = GradientBoostingRegressor(n_estimators=100, random_state=42)
gbm_price.fit(X_train_p, y_train_p)
y_pred_price = gbm_price.predict(X_test_p)

print("=== Pricing Model Fairness Audit ===")
print(f"Overall mean predicted premium: ${y_pred_price.mean():,.0f}")

# Now audit by ZIP code group
df_test_p = X_test_p.copy()
df_test_p['predicted_premium'] = y_pred_price
df_test_p['zip_code_proxy'] = df_pricing.loc[X_test_p.index, 'zip_code_proxy']

# Demographic parity: mean predicted premium by group
demographic_parity = df_test_p.groupby('zip_code_proxy')['predicted_premium'].agg(['mean', 'std'])
print("\nDemographic Parity (Mean Predicted Premium by ZIP Code Group):")
print(demographic_parity.round(2))

# Disparity ratio
disparity_ratio = demographic_parity.loc['low_income', 'mean'] / demographic_parity.loc['high_income', 'mean']
print(f"\nDisparity ratio (low-income / high-income): {disparity_ratio:.2f}x")
print(f"A ratio > 1.25x is a common regulatory yellow flag.")

What does this tell us? The low-income ZIP code group has an average predicted premium of about 720,comparedto720, compared to 525 for the high-income group — a disparity ratio of 1.37x. That’s above the 1.25x threshold that many regulators consider a yellow flag. The model is pricing low-income drivers more, even after controlling for other factors. Is that justified by risk, or is it discrimination?

Fairness-Constrained Re-training: Removing ZIP Code

Let’s see what happens when we remove the ZIP code proxy and retrain the model.

# Retrain without ZIP code
features_no_zip = ['age', 'vehicle_value', 'prior_claims', 'credit_score']
X_price_nozip = df_pricing[features_no_zip]

X_train_nz, X_test_nz, y_train_nz, y_test_nz = train_test_split(X_price_nozip, y_price, test_size=0.3, random_state=42)

gbm_price_fair = GradientBoostingRegressor(n_estimators=100, random_state=42)
gbm_price_fair.fit(X_train_nz, y_train_nz)
y_pred_price_fair = gbm_price_fair.predict(X_test_nz)

# Evaluate accuracy loss
from sklearn.metrics import mean_absolute_error
mae_original = mean_absolute_error(y_test_p, y_pred_price)
mae_fair = mean_absolute_error(y_test_nz, y_pred_price_fair)
print("=== Accuracy Impact of Removing ZIP Code ===")
print(f"MAE (original model): ${mae_original:,.0f}")
print(f"MAE (model without ZIP): ${mae_fair:,.0f}")
print(f"Accuracy loss: {(mae_fair - mae_original) / mae_original * 100:.1f}% increase in error")

# Now audit the fair model
df_test_fair = X_test_nz.copy()
df_test_fair['predicted_premium'] = y_pred_price_fair
df_test_fair['zip_code_proxy'] = df_pricing.loc[X_test_nz.index, 'zip_code_proxy']

fair_parity = df_test_fair.groupby('zip_code_proxy')['predicted_premium'].agg(['mean', 'std'])
print("\nDemographic Parity (Fair Model):")
print(fair_parity.round(2))
fair_disparity = fair_parity.loc['low_income', 'mean'] / fair_parity.loc['high_income', 'mean']
print(f"\nDisparity ratio (fair model): {fair_disparity:.2f}x")

Interpret this carefully. Removing ZIP code drops the disparity ratio from 1.37x to 1.09x — a meaningful improvement. But the model’s accuracy decreased: MAE increased by about 6%. Is that trade-off worth it? That’s not a math question — it’s a regulatory and ethical one.

As the American Academy of Actuaries notes, simply removing protected traits isn’t enough — other features like credit score or vehicle value can still proxy for race or income. The Xin & Huang paper (2022) demonstrates this: a model that uses only non-protected features can still produce discriminatory premiums because those features correlate with protected attributes.


Recap: The Three Pillars and Their Human Cost

Let’s summarize what we’ve learned.

  1. Underwriting is a classification problem — but the features you use (or don’t use) are a regulatory minefield. Our logistic regression gave AUC 0.71; XGBoost pushed to 0.89. The top feature was ZIP code, which is also the most controversial.

  2. Claims pricing is a two-part problem — frequency and severity, each needing its own model type. The GBM frequency model (AUC 0.85) clearly beat the Poisson GLM (AUC 0.72), but the GBM severity model barely improved over the log-normal GLM (3% RMSE improvement). Claim amounts are inherently noisy.

  3. Fraud detection is a precision-recall problem first, an accuracy problem second. Our XGBoost model caught 54% of fraud at 92% precision, but tuning to 80% precision caught 62% of fraud. Anomaly detection (Isolation Forest) provided a 3.4x lift, catching novel fraud patterns.

  4. Fairness audits aren’t optional — they’re becoming mandatory, and the math has no clean answer. Removing ZIP code reduced the disparity ratio from 1.37x to 1.09x but cost 6% accuracy. The trade-off is ethical, not just mathematical.

The models you learned here are the same ones used at AIG, Allstate, and AXA — just with a lot more zeros on the data. The difference is in how carefully the assumptions are examined.

In the next part of this series, we’ll build one of these systems from scratch — a full underwriting risk-score pipeline with a fairness dashboard — using a real(ish) dataset and all the production considerations (feature store, model monitoring, regulatory documentation).


Check Your Understanding

Test your understanding of the material with these questions, ordered by Bloom’s taxonomy level.

Remember: What are the two components of the two-part actuarial pricing model?

Understand: In your own words, explain why a single regression model for claim amounts would be misleading.

Apply: Given the synthetic dataset in this tutorial, train a logistic regression for underwriting risk and compute its AUC. How does it compare to the XGBoost?

Analyze: Your fraud model flags a claim as suspicious. The claim is from a long-term customer with a clean record, but the incident type is ‘theft of vehicle,’ which has higher fraud rates. Is this a good flag? What additional features would you want before investigation?

Evaluate: An insurance regulator tells you your pricing model has demographic parity (equal average premiums across groups) but you suspect it’s still discriminatory against young drivers. How is that possible? (Hint: think about what ‘average’ hides.)

Create: Design a ‘fairness monitoring dashboard’ for an insurance pricing model — what plots, metrics, and alerts would you include? Write a one-paragraph spec.


  • Part 2: Machine Learning in Banking and Finance: Credit Risk, Fraud, and Algorithmic Trading — Both banking and insurance face similar challenges with fairness and discrimination, though insurance has older actuarial traditions and more explicit proxy-discrimination scrutiny.
  • Part 5: Machine Learning in Cybersecurity: Anomaly Detection and Threat Classification — Anomaly detection patterns in cybersecurity (find unusual network traffic) parallel fraud detection in insurance (find unusual claims), both using Isolation Forest and supervised learning.

Apply What You Learned is for Supporter and Insider subscribers.

Subscribe to unlock the exercises on this post.

See plans

Looking for something else?

Search every article by title, summary or topic.