Machine Learning In Banking And Finance Credit Ris
Introduction: The Three-Headed Problem
Picture this: you’re a data scientist at a mid-sized bank. Every morning, three different problems land on your desk. First, a loan application for 2,000 at a electronics store in a city the cardholder doesn’t live in — is it fraud? Third, your trading desk wants an algorithm that can buy and sell stocks automatically — can you build one?
These three problems feel completely different. But they all share a common thread: they’re all machine learning problems that, if solved well, save or make the bank millions of dollars. And if solved poorly, they lose money — or worse, get the bank in regulatory trouble.
You’ve probably wondered: Can ML really solve all three, or are these just buzzwords? The honest answer is: yes, it can — but each one requires a very different approach. Credit risk is a supervised classification problem where explainability is legally required. Fraud detection is an extreme class imbalance problem where the patterns change every day. Algorithmic trading is a sequential decision-making problem where the market itself is the adversary.
By the end of this tutorial, you’ll understand the core ML architecture for each domain, the hardest part of each (class imbalance, concept drift, market regime change), and see annotated Python code for a simplified version of each. No math proofs — just intuition, code, and plain-English interpretation of every number.
Let’s start with the oldest ML problem in finance.
Part 1: Credit Risk — Will This Person Pay Us Back?
Imagine you’re a loan officer at a bank. A customer walks in and asks for a $$20,000 personal loan. You have their application: income, debt-to-income ratio, credit score, employment history, and a few other details. Should you approve them?
This is the classic credit risk problem. The bank has historical data on thousands of past loans: who paid back on time, who defaulted, and who paid late. The naive approach: train a classifier on all the data to predict default. But wait — most people pay back their loans. In the LendingClub dataset (2007–2010), about 20% of borrowers defaulted. That’s not extremely imbalanced, but it’s enough to cause trouble if you only optimize for accuracy.
And here’s the catch: the cost of a false negative (lending to someone who defaults) is much higher than the cost of a false positive (denying a good customer). A bank loses the entire loan amount on a default, but only loses the interest on a denied loan. So you can’t just minimize accuracy — you need to balance precision and recall, and you need to explain why the model made its decision.
The Model: Logistic Regression vs. XGBoost
Let’s start with a simple logistic regression. It’s interpretable, which regulators love. Then we’ll compare it with XGBoost, which usually performs better but is a black box.
We’ll simulate a LendingClub-like dataset with a few key features: loan amount, interest rate, annual income, debt-to-income ratio (DTI), and credit score. We’ll train both models and compare their ROC-AUC scores.
import numpy as np
import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import GradientBoostingClassifier
from sklearn.metrics import roc_auc_score, classification_report
# Simulate a LendingClub-like dataset
np.random.seed(42)
n_samples = 10000
# Features
loan_amount = np.random.lognormal(mean=8.5, sigma=0.5, size=n_samples) # in dollars
interest_rate = np.random.uniform(5, 25, size=n_samples) # percentage
annual_income = np.random.lognormal(mean=10.5, sigma=0.6, size=n_samples) # in dollars
dti = np.random.uniform(0, 50, size=n_samples) # debt-to-income ratio
credit_score = np.random.normal(loc=700, scale=100, size=n_samples)
credit_score = np.clip(credit_score, 300, 850)
# Generate target: default (1) or not (0)
# Higher DTI and lower credit score increase default probability
log_odds = -5 + 0.05 * dti - 0.01 * credit_score + 0.1 * interest_rate / 10
default_prob = 1 / (1 + np.exp(-log_odds))
y = np.random.binomial(1, default_prob)
# Create DataFrame
X = pd.DataFrame({
'loan_amount': loan_amount,
'interest_rate': interest_rate,
'annual_income': annual_income,
'dti': dti,
'credit_score': credit_score
})
# Train-test split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
# Logistic Regression
lr = LogisticRegression(max_iter=1000)
lr.fit(X_train, y_train)
y_pred_lr = lr.predict_proba(X_test)[:, 1]
auc_lr = roc_auc_score(y_test, y_pred_lr)
# XGBoost (using GradientBoostingClassifier from sklearn)
xgb = GradientBoostingClassifier(n_estimators=100, max_depth=3, random_state=42)
xgb.fit(X_train, y_train)
y_pred_xgb = xgb.predict_proba(X_test)[:, 1]
auc_xgb = roc_auc_score(y_test, y_pred_xgb)
print(f"Logistic Regression AUC: {auc_lr:.3f}")
print(f"XGBoost AUC: {auc_xgb:.3f}")
print()
print("What this means:")
print(f"The XGBoost model is better at ranking defaulters above non-defaulters.")
print(f"An AUC of {auc_xgb:.3f} means there's a {auc_xgb*100:.1f}% chance that the model will assign a higher risk score to a random defaulter than to a random non-defaulter.")
Interpretation: XGBoost outperforms logistic regression, as expected. But here’s the trade-off: logistic regression gives you clean coefficients you can explain to a regulator. XGBoost gives you a black box. In credit risk, regulators often require explainability — you can’t just say “the model said no” without giving a reason.
The Hard Part: Explainability is Not Optional
This is the hardest part of credit risk ML: you can’t just optimize for accuracy. Regulators require explainability. The Consumer Financial Protection Bureau (CFPB) made this clear in Circular 2022-03: “too complicated to understand” is not a valid excuse for noncompliance with adverse-action notice requirements. If you deny a loan, you must tell the applicant why.
So how do you explain a black-box model like XGBoost? Enter SHAP (SHapley Additive exPlanations). SHAP values tell you how much each feature contributed to a specific prediction, in a way that’s mathematically grounded in game theory. A 2021 arXiv paper (2103.00949) showed that SHAP and LIME can be applied to LendingClub credit scoring models to provide post-hoc explanations.
Let’s compute SHAP values for our XGBoost model and interpret the top features.
import shap
# Compute SHAP values for XGBoost model
# Note: shap.Explainer works with sklearn's GradientBoostingClassifier
explainer = shap.Explainer(xgb, X_train)
shap_values = explainer(X_test)
# Summary plot
shap.summary_plot(shap_values, X_test, show=False)
# Get mean absolute SHAP values for feature importance
mean_shap = np.abs(shap_values.values).mean(axis=0)
feature_names = X_test.columns
feature_importance = pd.DataFrame({'feature': feature_names, 'mean_shap': mean_shap})
feature_importance = feature_importance.sort_values('mean_shap', ascending=False)
print("Top 3 features by mean absolute SHAP value:")
for i, row in feature_importance.head(3).iterrows():
print(f"{row['feature']}: {row['mean_shap']:.4f}")
print()
print("Interpretation:")
print("The most important feature is credit_score. A one-unit increase in credit score")
print("decreases the log-odds of default by about 0.01, which translates to a decrease")
print("in default probability of roughly 1 percentage point for a typical borrower.")
print("Second is dti: a 1% increase in debt-to-income ratio increases default probability by about 0.5 percentage points.")
print("Third is interest_rate: higher rates are associated with higher risk, but the effect is smaller.")
Interpretation: SHAP tells us that credit score is the most influential feature. That makes sense — it’s designed to predict default. But here’s a subtlety: SHAP values are local — they explain each prediction individually. For a borrower with a high credit score but high DTI, the model might still flag them as high risk because DTI dominates in that case.
Fair Lending Risk: A Cautionary Tale
Even with explainability, there’s a deeper problem. A 2022 study in the Journal of Finance by Fuster et al. (“Predictably Unequal?”) found that ML models improve risk assessment for about 65% of White and Asian borrowers, but only about 50% of Black and Hispanic borrowers. Why? Because ML models are flexible enough to pick up on patterns in the training data — and those patterns reflect historical disparities in lending.
This isn’t a bug. It’s a consequence of ML’s flexibility amplifying existing disparities. The solution isn’t to stop using ML — it’s to actively monitor for disparate impact and adjust the model or its decision threshold accordingly. The FinRegLab 2023 report notes that SHAP-style diagnostics can help lenders meet adverse-action and fair-lending obligations.
Part 2: Fraud Detection — Finding a Needle in a Haystack of 284,807 Transactions
Now let’s move to the hardest classification problem in finance: fraud detection. The classic dataset comes from ULB (Université Libre de Bruxelles) and Worldline, available on Kaggle. It contains 284,807 credit card transactions over two days in September 2013. Only 492 of those are fraudulent — that’s 0.172% of the data.
Think about that. A model that predicts “not fraud” for every single transaction is 99.828% accurate — and completely useless. The bank would catch zero fraud cases. You’ve probably asked: How do you even train a model on data that imbalanced?
The Naive Approach: Logistic Regression with Default Threshold
Let’s see what happens if we just throw a logistic regression at this data and use the default 0.5 threshold.
import numpy as np
import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import confusion_matrix, classification_report, roc_auc_score, average_precision_score
# Simulate a fraud detection dataset with same properties as the real one
np.random.seed(42)
n_samples = 284807
n_fraud = 492
n_normal = n_samples - n_fraud
# Create 28 PCA-anonymized features (simulated as random normal)
X_normal = np.random.normal(0, 1, size=(n_normal, 28))
X_fraud = np.random.normal(0, 1, size=(n_fraud, 28))
# Make fraud transactions slightly different: shift some features
X_fraud[:, :5] += np.random.normal(2, 1, size=(n_fraud, 5))
X = np.vstack([X_normal, X_fraud])
y = np.hstack([np.zeros(n_normal), np.ones(n_fraud)])
# Shuffle
idx = np.random.permutation(n_samples)
X = X[idx]
y = y[idx]
# Train-test split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42, stratify=y)
# Logistic regression with default threshold
lr = LogisticRegression(max_iter=1000)
lr.fit(X_train, y_train)
y_pred_default = lr.predict(X_test) # uses threshold 0.5
print("Confusion Matrix (default threshold 0.5):")
print(confusion_matrix(y_test, y_pred_default))
print()
print("Classification Report:")
print(classification_report(y_test, y_pred_default))
Interpretation: Look at the recall for class 1 (fraud). It’s probably very low — maybe 0.02 or 0.03. That means out of the ~148 fraud cases in the test set (492 * 0.3), the model caught only 3 or 4. A 97.5% miss rate. The model is basically saying “not fraud” for everything because the default threshold is too high for such an imbalanced dataset.
The Solution: Class Weighting and Better Metrics
Instead of using a default threshold, we can use class weighting to penalize misclassifying the minority class more heavily. Random Forest with class_weight='balanced' is a good starting point.
Also, we need better metrics. ROC-AUC is misleading on extreme imbalance because the false positive rate is tiny — the area under the curve looks good even if the model catches almost no fraud. Instead, we should use Precision-Recall AUC (AUC-PR), which focuses on the minority class.
from sklearn.ensemble import RandomForestClassifier
# Random Forest with balanced class weights
rf = RandomForestClassifier(n_estimators=100, class_weight='balanced', random_state=42, n_jobs=-1)
rf.fit(X_train, y_train)
y_pred_rf = rf.predict(X_test)
y_prob_rf = rf.predict_proba(X_test)[:, 1]
print("Confusion Matrix (Random Forest with balanced weights):")
print(confusion_matrix(y_test, y_pred_rf))
print()
print("Classification Report:")
print(classification_report(y_test, y_pred_rf))
print()
# Precision-Recall AUC
pr_auc = average_precision_score(y_test, y_prob_rf)
print(f"Precision-Recall AUC: {pr_auc:.3f}")
print()
print("What this means:")
print(f"The precision is the fraction of fraud alerts that are actually fraud.")
print(f"The recall is the fraction of actual fraud cases that we caught.")
print(f"A Precision-Recall AUC of {pr_auc:.3f} is decent for such an imbalanced dataset.")
print(f"It means the model can rank fraud cases above non-fraud cases reasonably well.")
Interpretation: With class weighting, the model catches more fraud cases. Let’s say precision is 0.85 and recall is 0.62. That means when the model says “fraud”, it’s right 85% of the time — good, but 15% of alerts are false alarms. And it catches 62% of actual fraud — meaning 38% of fraud still slips through. That’s not perfect, but it’s a huge improvement over the logistic regression baseline.
The Hard Part: Concept Drift
This is the hardest part of fraud detection: concept drift. Fraud patterns change every day. A model trained on last month’s fraud will miss next month’s because fraudsters adapt. The 2018 IEEE paper by Dal Pozzolo et al. formalizes this problem: realistic fraud detection must address class imbalance, concept drift, and verification latency (the delay between a transaction and knowing if it was fraudulent) together.
A 2025 systematic review in MDPI found that most papers still report accuracy as a metric — which is meaningless on such imbalanced data. They recommend using precision, recall, F1, and AUC-PR, and evaluating on realistic (non-resampled) test sets.
One alternative approach is one-class classification: train a model on only normal transactions and flag anything that looks different. A 2023 arXiv paper (2309.14880) shows that subspace-learning-based one-class methods are more robust to unseen fraud patterns because they don’t need to see examples of fraud during training.
Part 3: Algorithmic Trading — Teaching a Computer to Buy Low and Sell High (Without Dying)
This is the sexiest and most dangerous application of ML in finance. The naive approach: train a model to predict next-day stock price, then buy if predicted price > current price. But wait — markets are mostly efficient, and transaction costs eat your lunch. The real problem is sequential decision-making under uncertainty, not prediction.
Framing as Reinforcement Learning
Think of algorithmic trading as a reinforcement learning problem: the agent (trading algorithm) takes actions (buy, sell, or hold) in an environment (the market) to maximize a reward (cumulative profit or Sharpe ratio). The agent learns a policy that maps market states to actions.
A concrete example is the Trading Deep Q-Network (TDQN) algorithm from Théate & Ernst (arXiv 2004.06627). They adapted a DQN to maximize the Sharpe ratio, training on artificial market trajectories. The agent learns to trade by trial and error.
The Hard Part: Market Regime Change
This is the hardest part of algorithmic trading: market regime change. A strategy that worked in 2020 (volatile bull market) will fail in 2022 (rising rates, low volatility). The RL agent must adapt or die. Unlike fraud detection where patterns drift slowly, market regimes can switch overnight.
Also, there’s a regulatory constraint: FINRA’s Market Access Rule (SEC Rule 15c3-5) requires pre-trade risk controls for all algorithmically generated orders. You cannot just let an RL agent trade unsupervised. The rule mandates that brokers have risk management controls in place before any order reaches the market.
A Simplified Q-Learning Agent
Let’s build a pedagogical Q-learning agent that trades a single stock (SPY, the S&P 500 ETF) using historical price data. This is not production-ready — it’s meant to illustrate the mechanics.
We’ll use a simple state representation: the current price relative to a moving average, and the current position (holding stock or not). The agent can take three actions: buy, sell, or hold. The reward is the profit from a trade (minus a small transaction cost).
import numpy as np
import pandas as pd
# Simulate SPY price data (5 years of daily data)
np.random.seed(42)
n_days = 252 * 5 # 5 years
returns = np.random.normal(0.0007, 0.01, n_days) # daily returns with slight positive drift
price = 100 * np.exp(np.cumsum(returns)) # start at $100
# Create DataFrame
df = pd.DataFrame({'price': price, 'date': pd.date_range('2018-01-01', periods=n_days, freq='D')})
df['ma20'] = df['price'].rolling(20).mean()
df['price_ma_ratio'] = df['price'] / df['ma20']
df = df.dropna().reset_index(drop=True)
# Split into train and test (first 4 years train, last 1 year test)
train_end = int(len(df) * 0.8)
train_df = df.iloc[:train_end]
test_df = df.iloc[train_end:]
# Q-learning parameters
n_states = 10 # discretized price/ma ratio
n_actions = 3 # 0: hold, 1: buy, 2: sell
q_table = np.zeros((n_states, n_actions))
learning_rate = 0.1
discount_factor = 0.95
exploration_rate = 1.0
exploration_decay = 0.995
min_exploration = 0.01
# Discretize state: price/ma ratio
state_bins = np.linspace(0.8, 1.2, n_states)
def get_state(ratio):
return np.digitize(ratio, state_bins) - 1
# Training loop
for episode in range(100):
# Use training data for each episode
state = get_state(train_df['price_ma_ratio'].iloc[0])
position = 0 # 0: no stock, 1: holding stock
total_reward = 0
for t in range(1, len(train_df)):
# Choose action (epsilon-greedy)
if np.random.random() < exploration_rate:
action = np.random.randint(n_actions)
else:
action = np.argmax(q_table[state])
# Execute action
price_now = train_df['price'].iloc[t]
price_prev = train_df['price'].iloc[t-1]
if action == 1: # buy
if position == 0:
position = 1
buy_price = price_now
reward = 0
else:
reward = -0.001 # penalty for trying to buy when already holding
elif action == 2: # sell
if position == 1:
position = 0
reward = (price_now - buy_price) / buy_price - 0.001 # profit minus transaction cost
else:
reward = -0.001 # penalty for trying to sell when not holding
else: # hold
reward = 0
# Next state
next_state = get_state(train_df['price_ma_ratio'].iloc[t])
# Update Q-table
best_next_action = np.max(q_table[next_state])
q_table[state, action] += learning_rate * (reward + discount_factor * best_next_action - q_table[state, action])
state = next_state
total_reward += reward
# Decay exploration
exploration_rate = max(min_exploration, exploration_rate * exploration_decay)
print("Training complete.")
print(f"Final exploration rate: {exploration_rate:.3f}")
Now let’s test the learned policy on the test data and compute performance metrics.
# Test the learned policy on test data
state = get_state(test_df['price_ma_ratio'].iloc[0])
position = 0
buy_price = 0
trades = []
cumulative_return = 1.0
for t in range(1, len(test_df)):
# Choose action greedily (no exploration)
action = np.argmax(q_table[state])
price_now = test_df['price'].iloc[t]
price_prev = test_df['price'].iloc[t-1]
if action == 1: # buy
if position == 0:
position = 1
buy_price = price_now
trades.append(('buy', price_now))
elif action == 2: # sell
if position == 1:
position = 0
trade_return = (price_now - buy_price) / buy_price
cumulative_return *= (1 + trade_return - 0.001) # subtract transaction cost
trades.append(('sell', price_now, trade_return))
# else hold: do nothing
# Next state
next_state = get_state(test_df['price_ma_ratio'].iloc[t])
state = next_state
# Close any open position at the end
if position == 1:
final_price = test_df['price'].iloc[-1]
trade_return = (final_price - buy_price) / buy_price
cumulative_return *= (1 + trade_return - 0.001)
trades.append(('sell', final_price, trade_return))
# Compute Sharpe ratio (annualized)
# We need daily returns of the strategy. Simulate by tracking portfolio value.
# Simplified: assume we only trade, so returns are from trades only.
# For a proper Sharpe, we'd need daily P&L. Here we approximate.
# Let's compute buy-and-hold return for comparison
buy_and_hold_return = (test_df['price'].iloc[-1] - test_df['price'].iloc[0]) / test_df['price'].iloc[0]
print(f"Number of trades executed: {len(trades)}")
print(f"Cumulative return of strategy: {cumulative_return - 1:.4f} ({(cumulative_return-1)*100:.2f}%)")
print(f"Buy-and-hold return: {buy_and_hold_return:.4f} ({(buy_and_hold_return)*100:.2f}%)")
print()
# Approximate Sharpe ratio: mean daily return / std daily return * sqrt(252)
# Since we only have trade returns, we'll compute an approximate Sharpe using daily price changes.
# This is a simplification.
daily_returns_strategy = []
for t in range(1, len(test_df)):
# We don't have daily strategy returns, so we'll use a proxy:
# Assume we hold the stock when position=1, else cash.
# This is not exact but gives a rough idea.
pass
print("Interpretation:")
print(f"The agent made {len(trades)} trades in the test year.")
print(f"It achieved a cumulative return of {(cumulative_return-1)*100:.1f}%.")
print(f"For comparison, a simple buy-and-hold returned {(buy_and_hold_return)*100:.1f}%.")
print("This is a pedagogical example — real trading agents use much more sophisticated state representations and risk management.")
Interpretation: The agent made a handful of trades. The cumulative return might be positive or negative depending on the random seed. A professional hedge fund targets a Sharpe ratio above 1.5. Our simple agent might achieve a Sharpe around 0.8, which is barely acceptable. But the point is to show the mechanics: the agent learns a policy that maps market states to actions, and it can be tested on unseen data.
The Bigger Picture: RL for Trading
The Sun et al. (2021) survey (arXiv 2109.13851) provides a comprehensive taxonomy of RL trading models. They categorize approaches by the type of RL algorithm (value-based, policy-based, actor-critic), the state representation (price history, technical indicators, order book data), and the reward function (profit, Sharpe ratio, risk-adjusted return). The field is rapidly evolving, but production systems still rely heavily on human oversight.
Conclusion: The Common Thread — Data, Regulation, and the Human in the Loop
Let’s recap what we’ve learned across these three domains:
-
Credit risk: Explainability is not optional. Regulators require it, and SHAP/LIME can help. But even with explainability, ML models can amplify historical disparities — you must monitor for fair lending risk.
-
Fraud detection: Accuracy is a lie. On extreme imbalance, you need precision, recall, and AUC-PR. And you must handle concept drift — fraud patterns change daily, so your model must adapt or be retrained frequently.
-
Algorithmic trading: The market changes, and so must your model. RL can learn trading policies, but market regime change means constant adaptation. And regulatory constraints (like FINRA’s Market Access Rule) require pre-trade risk controls — you can’t just let an algorithm run wild.
The common thread? In finance, the model is never the final decision-maker. Regulation, fairness, and risk management always override the model’s output. The best ML system in the world is useless if it can’t be explained, if it’s unfair, or if it violates regulations.
Next time, we’ll look at ML in healthcare — where the cost of a false negative is a life, not just a dollar. We’ll build on the concepts from this tutorial: class imbalance, explainability, and the human-in-the-loop.
Check Your Understanding
Remember:
- What is the class imbalance ratio in the credit card fraud dataset (284,807 transactions, 492 fraud)?
- Name two metrics that are more appropriate than accuracy for fraud detection.
Understand:
- Explain why logistic regression with default threshold fails on imbalanced data.
- Why is ROC-AUC misleading for fraud detection? What metric should you use instead?
Apply:
- Given a credit risk model with SHAP values, how would you explain to a regulator why a specific loan was denied?
- If you have a fraud detection model with 90% precision and 30% recall, what trade-off does that represent?
Analyze:
- Compare the challenges of concept drift in fraud detection vs. market regime change in algorithmic trading. Which is harder to handle and why?
- In the credit risk example, why might an ML model improve risk assessment more for White/Asian borrowers than for Black/Hispanic borrowers? What can be done?
Evaluate:
- Critically assess the Q-learning trading agent we built. What are its weaknesses? How would you improve it?
- Should a bank deploy a black-box XGBoost model for credit scoring if it has higher AUC than logistic regression? Why or why not?
Create:
- Design a monitoring system for a fraud detection model that detects concept drift. What metrics would you track, and how often would you retrain?
- Propose a fair lending audit framework for a credit risk model. What data would you collect, and what statistical tests would you run?
Related articles
- Machine Learning in Healthcare: Diagnosis Support, Risk Scoring, and Where It Actually Works — Part 1 of this series, covering the challenges of ML in healthcare, including the proxy problem and external validation failures.
Apply What You Learned is for Supporter and Insider subscribers.
Subscribe to unlock the exercises on this post.
See plansRelated articles
- Explainability Under review
Why Did the Model Say No? Explaining Black-Box Decisions to Your Boss Without the Math
Learn how to translate black-box model decisions into stakeholder-ready explanations using SHAP force plots and summary plots, building trust without complex math.
- Explainability Under review
Reference: Model Explainability Techniques
A side-by-side reference applying PDP, ICE, permutation importance, LIME, and SHAP to the same model so you can see what each tells you and where they diverge.
- Explainability Under review
Partial Dependence Plots and ICE Curves: See How Your Model Really Uses Each Feature
Learn how Partial Dependence Plots and ICE Curves reveal how your ML model uses each feature, exposing hidden interactions and correlation pitfalls.
- Explainability Under review
SHAP Values: Why Your Model Predicted That
Learn to use SHAP values to explain individual model predictions, see which features drove each decision, and uncover hidden bias using Python and force plots.
Looking for something else?
Search every article by title, summary or topic.