Python & Data Science

Machine Learning in Cybersecurity: Anomaly Detection and Threat Classification

You’re a security operations center (SOC) analyst at 3 AM. Your phone buzzes — the machine learning-based intrusion detection system just fired an alert. You rub your eyes and pull up the dashboard. The alert says: “High-confidence malicious traffic detected — 99.7% probability.”

You investigate. Ten minutes later, you find it: the flagged traffic is a legitimate software update download. A false positive. The model was so sure it was an attack, but it was dead wrong.

Worse: you later discover that the model did miss a real attack the same night — a SQL injection that the attacker masked by adding small, human-imperceptible delays between each request. The model never blinked.

Why did the model fail?

This is the hardest part of cybersecurity ML. Benchmarks are clean. Real networks are not. Attackers are intelligent adversaries — they’re actively trying to fool your model. And the traffic patterns you trained on? They change the moment you deploy to a different network.

Here’s the promise of this tutorial: you’ll learn to detect anomalies, classify threats, and — more importantly — understand when your model will break in the wild. By the end, you’ll know why a model that scores 99% on a benchmark can fail catastrophically on real traffic, and you’ll have the tools to evaluate your models honestly.

Let’s start with a foundational question: what’s the difference between finding an anomaly and naming the threat?

Anomaly Detection vs Threat Classification: What’s the Difference?

Think of a security guard who knows a building’s normal rhythm. They know employees arrive between 8 and 9 AM, the cleaning crew comes at 7 PM, and the only person who works weekends is the night watchman. If someone enters at 2 AM, the guard’s internal anomaly detection fires: “Something is different.”

Now imagine a different guard — one who has a detailed list of every employee, contractor, and visitor, along with their photo and clearance level. This guard doesn’t just know something is off — they can name the person and their access level. That’s threat classification.

Anomaly detection = finding things that don’t look like the normal baseline. Threat classification = putting a name on the anomaly (e.g., “Brute Force” vs “DDoS” vs “Benign”).

In cybersecurity, you often need both. The anomaly detector catches the unusual behavior. The classifier tells you what kind of attack it is (or that it’s actually benign).

Let’s see this in action with two classic models: IsolationForest (unsupervised anomaly detection) and RandomForest (supervised classification). We’ll use the CIC-IDS2017 dataset, a widely-used benchmark that contains benign traffic and eight attack types.

import pandas as pd
import numpy as np
from sklearn.ensemble import IsolationForest, RandomForestClassifier
from sklearn.decomposition import PCA
from sklearn.preprocessing import StandardScaler
import matplotlib.pyplot as plt

# --- Load and prepare the data ---
# We'll simulate a small sample of CIC-IDS2017-like data
np.random.seed(42)
n_samples = 1000

# Features: packet length, duration, bytes per second, ports
# Normally these would come from a real CSV
features = np.random.randn(n_samples, 4) * 2 + 10

# Create some anomalies (scattered far from normal)
anomaly_indices = np.random.choice(n_samples, size=50, replace=False)
features[anomaly_indices] += 20 * np.random.randn(50, 4)

# Labels: 0 = benign, 1-3 = attack types
labels = np.zeros(n_samples, dtype=int)
labels[anomaly_indices] = np.random.choice([1, 2, 3], size=50)

# Create DataFrame for inspection
df = pd.DataFrame(features, columns=['packet_len', 'duration', 'bytes_per_sec', 'port'])
df['label'] = labels

# --- Train IsolationForest (unsupervised) ---
# It doesn't use labels — it just learns what's "normal"
iso_forest = IsolationForest(contamination=0.05, random_state=42)
iso_preds = iso_forest.fit_predict(features)  # 1 = normal, -1 = anomaly

# --- Train RandomForest (supervised) ---
# It uses labels to learn how to distinguish attack types
rf = RandomForestClassifier(n_estimators=100, random_state=42)
rf.fit(features, labels)
rf_preds = rf.predict(features)

# --- Compare on a single record ---
# Take one malicious record (index 0 of anomalies)
anomaly_idx = anomaly_indices[0]
print(f"Record index: {anomaly_idx}")
print(f"True label: {labels[anomaly_idx]} (1=DoS, 2=DDoS, 3=BruteForce)")
print(f"IsolationForest says: {'ANOMALY' if iso_preds[anomaly_idx] == -1 else 'NORMAL'}")
print(f"RandomForest says: Attack type {rf_preds[anomaly_idx]}")

The output shows the difference clearly. The IsolationForest flags it as unusual — but it can’t tell you why or what kind. The RandomForest names it: “This is a Brute Force attack.”

The Data Pipeline: Cleaning a Noisy Security Log

Real network logs are messy — much messier than our synthetic example. CIC-IDS2017 has over 80 flow features (things like ‘Fwd Packet Length Mean’, ‘Bwd IAT Total’, ‘Flow Bytes/s’). Many of them are missing, infinite, or just noise.

Let’s walk through a realistic pipeline: loading the data, handling missing values, dealing with class imbalance, and selecting useful features.

import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.feature_selection import mutual_info_classif
from sklearn.preprocessing import LabelEncoder

# --- Simulate loading CIC-IDS2017 ---
# Real CSV would be: df = pd.read_csv('cic_ids2017.csv')
np.random.seed(42)
n_samples = 5000

# Simulate 15 features (mimicking real ones)
feature_names = [
    'fwd_pkt_len_mean', 'bwd_pkt_len_mean', 'flow_duration',
    'fwd_iat_tot', 'bwd_iat_tot', 'flow_bytes_per_sec',
    'fwd_pkts_per_sec', 'bwd_pkts_per_sec', 'pkt_len_min',
    'pkt_len_max', 'pkt_len_mean', 'pkt_len_std',
    'fwd_psh_flags', 'bwd_psh_flags', 'init_win_bytes_fwd'
]
X = np.random.randn(n_samples, 15) * 10 + 100

# Introduce realistic missing values and infinities
for i in [0, 3, 5]:  # mess up a few features
    missing_mask = np.random.random(n_samples) < 0.05
    X[missing_mask, i] = np.nan
    inf_mask = np.random.random(n_samples) < 0.01
    X[inf_mask, i] = np.inf

# Create labels: benign (0) and 4 attack types (1-4)
y = np.zeros(n_samples, dtype=int)
attack_indices = np.random.choice(n_samples, size=800, replace=False)
y[attack_indices] = np.random.choice([1, 2, 3, 4], size=800)

# Label names for reference
attack_names = {0: 'Benign', 1: 'DoS', 2: 'DDoS', 3: 'Brute Force', 4: 'Web Attack'}

# --- Step 1: Remove features with >50% missing values ---
df = pd.DataFrame(X, columns=feature_names)
missing_percent = df.isnull().sum() / len(df)
features_to_drop = missing_percent[missing_percent > 0.5].index.tolist()
if features_to_drop:
    print(f"Dropping features with >50% missing: {features_to_drop}")
    df = df.drop(columns=features_to_drop)

# Fill remaining NaNs with median and replace inf with large value
df = df.fillna(df.median())
df = df.replace([np.inf, -np.inf], 1e10)

# --- Step 2: Check class imbalance ---
class_counts = pd.Series(y).value_counts().sort_index()
print("\nClass distribution (before undersampling):")
for label, count in class_counts.items():
    print(f"  {attack_names[label]}: {count} samples")
print(f"\nBenign accounts for {class_counts[0]/n_samples*100:.1f}% of data.")
print("A model predicting 'Benign' every time would be that accurate. Useless.")

# --- Step 3: Feature selection with mutual information ---
mi_scores = mutual_info_classif(df.fillna(0), y, random_state=42)
top_features_idx = np.argsort(mi_scores)[-5:][::-1]  # top 5
selected_features = df.columns[top_features_idx].tolist()
print(f"\nTop 5 features by mutual information: {selected_features}")

X_selected = df[selected_features].values

# --- Step 4: Undersample benign to get ~10:1 ratio ---
benign_idx = np.where(y == 0)[0]
attack_idx = np.where(y != 0)[0]

# Take a subset of benign samples
n_benign_keep = len(attack_idx) * 10  # 10:1 benign:attack ratio
if len(benign_idx) > n_benign_keep:
    benign_idx_keep = np.random.choice(benign_idx, size=n_benign_keep, replace=False)
else:
    benign_idx_keep = benign_idx

# Combine
balanced_idx = np.concatenate([benign_idx_keep, attack_idx])
X_balanced = X_selected[balanced_idx]
y_balanced = y[balanced_idx]

print(f"\nAfter undersampling: {len(balanced_idx)} total samples")
print(f"  Benign: {len(benign_idx_keep)} ({len(benign_idx_keep)/len(balanced_idx)*100:.1f}%)")
print(f"  Attack: {len(attack_idx)} ({len(attack_idx)/len(balanced_idx)*100:.1f}%)")

# --- Step 5: Train/test split ---
X_train, X_test, y_train, y_test = train_test_split(
    X_balanced, y_balanced, test_size=0.3, random_state=42, stratify=y_balanced
)
print(f"\nTraining set: {len(X_train)} samples")
print(f"Test set: {len(X_test)} samples")

Interpretation: The dataset is severely imbalanced — benign traffic dominates. A model that always predicts “Benign” would be 85% accurate, but completely useless for detecting attacks. Undersampling helps, but it’s a trade-off: you lose information from the benign class. In production, you’d also consider oversampling attacks (SMOTE) or using anomaly detection on the raw imbalanced data.

The Benchmark Trap: How Test-Set Accuracy Lulls You Into a False Sense of Security

Now here’s where things get dangerous. You train a Random Forest on the CIC-IDS2017 training set. You evaluate on the test set. Score: 99.4% accuracy. You feel great. You deploy the model.

Then you test it on traffic from a different network — say, a slice of UNSW-NB15 (a dataset with different attack types and different normal traffic patterns).

Let’s see what happens.

from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import classification_report, accuracy_score
import numpy as np

# --- Train on what we prepared earlier ---
# (Repeating key steps for self-contained code block)
np.random.seed(42)
n_train = 4000

# Simulate CIC-IDS2017-like data
X_cic = np.random.randn(n_train, 5) * 10 + 100
y_cic = np.zeros(n_train, dtype=int)
attack_idx_cic = np.random.choice(n_train, size=600, replace=False)
y_cic[attack_idx_cic] = np.random.choice([1, 2, 3, 4], size=600)

# Train Random Forest
rf = RandomForestClassifier(n_estimators=100, random_state=42)
rf.fit(X_cic, y_cic)

# Evaluate on CIC-IDS2017 test set
n_test = 1000
X_cic_test = np.random.randn(n_test, 5) * 10 + 100
y_cic_test = np.zeros(n_test, dtype=int)
attack_idx_test = np.random.choice(n_test, size=150, replace=False)
y_cic_test[attack_idx_test] = np.random.choice([1, 2, 3, 4], size=150)

cic_preds = rf.predict(X_cic_test)
print("In-dataset performance (CIC-IDS2017 test set):")
print(f"Accuracy: {accuracy_score(y_cic_test, cic_preds):.4f}")
print(classification_report(y_cic_test, cic_preds, zero_division=0))

# --- Simulate cross-dataset test (UNSW-NB15-like) ---
# Different distribution: different centers, different attack patterns
X_unsw = np.random.randn(n_test, 5) * 15 + 90  # shifted distribution
y_unsw = np.zeros(n_test, dtype=int)
attack_idx_unsw = np.random.choice(n_test, size=150, replace=False)
y_unsw[attack_idx_unsw] = np.random.choice([1, 2, 3, 4], size=150)

unsw_preds = rf.predict(X_unsw)
print("\nCross-dataset performance (UNSW-NB15-like data):")
print(f"Accuracy: {accuracy_score(y_unsw, unsw_preds):.4f}")
print(classification_report(y_unsw, unsw_preds, zero_division=0))

Interpretation: The in-dataset accuracy might be 99%+. The cross-dataset accuracy drops to near random chance — maybe 50% for benign vs attack, worse for specific attack types. This isn’t a flaw in Random Forest. It’s that network traffic distributions vary wildly between organizations, times of day, and attack tools.

Research confirms this: four classifiers achieved near-perfect accuracy (>99%) within one dataset but dropped to near-random-chance accuracy when tested on another. The benchmark tells you the model learned the dataset, not the problem.

When the Attacker Fights Back: Adversarial Evasion

You train a model. You deploy it. It works… until the attacker adapts.

Attackers don’t just launch the same old attacks. They study your model — or at least, they study how security tools work. They make small changes to their malware or network traffic that are invisible to a human analyst but flip the model’s prediction.

Here’s a concrete example from research: a gradient-based attack modified individual bytes of malware binaries while preserving the executable functionality. The modified malware was classified as benign by deep learning models that had previously detected the original samples with 98% accuracy.

Let’s simulate a simplified version of this.

import numpy as np
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score

# --- Set up a simple scenario ---
np.random.seed(42)

# Simulate 5 numeric features from a malware detection model
# Features: API call count, entropy, file size, import count, section alignment
n_samples = 500
X = np.random.randn(n_samples, 5) * 2 + np.array([50, 7, 10000, 200, 1024])
y = np.random.choice([0, 1], size=n_samples, p=[0.7, 0.3])  # 0=benign, 1=malicious

# Train a simple model
rf = RandomForestClassifier(n_estimators=100, random_state=42)
rf.fit(X, y)

# Find a correctly classified malicious sample
malicious_indices = np.where(y == 1)[0]
correct_malicious = [i for i in malicious_indices if rf.predict([X[i]])[0] == 1]

if len(correct_malicious) > 0:
    sample_idx = correct_malicious[0]
    print(f"Original sample (malicious): Label {y[sample_idx]}")
    print(f"Model prediction: {rf.predict([X[sample_idx]])[0]} (1 = malicious)")
    print(f"Original features: {X[sample_idx]}")
    
    # --- Simulate adversarial perturbation ---
    # Attacker adds small noise to low-importance features
    # (In reality, this would be a carefully crafted gradient-based perturbation)
    perturbed = X[sample_idx].copy()
    
    # Which features are 'least important'?
    # Attacker modifies those because they're less likely to be monitored
    feature_importances = rf.feature_importances_
    least_important_idx = np.argmin(feature_importances)
    
    # Add a perturbation that's 0.5 standard deviations of that feature
    feature_std = np.std(X[:, least_important_idx])
    perturbed[least_important_idx] += 3 * feature_std  # 3 std shift
    
    # Also add a tiny perturbation to another feature
    second_least_idx = np.argsort(feature_importances)[1]
    perturbed[second_least_idx] += 2 * np.std(X[:, second_least_idx])
    
    print(f"\nAfter adversarial perturbation:")
    print(f"Perturbed features: {perturbed}")
    print(f"Model's new prediction: {rf.predict([perturbed])[0]}")
    
    if rf.predict([perturbed])[0] == 0:
        print("\nThe attacker successfully evaded detection!")
        print("The sample is still malicious, but the model now says it's benign.")
    else:
        print("\nThe model held firm this time.")
        print("Real attacks would use gradient-based optimization to find evasion.")
    
    # --- Quantify robustness ---
    # How many malicious samples would be flipped by this perturbation?
    malicious_X = X[malicious_indices].copy()
    for i in range(len(malicious_X)):
        # Perturb the same features for each malicious sample
        malicious_X[i, least_important_idx] += 3 * feature_std
        malicious_X[i, second_least_idx] += 2 * np.std(X[:, second_least_idx])
    
    original_preds = rf.predict(X[malicious_indices])
    perturbed_preds = rf.predict(malicious_X)
    flip_rate = np.mean(original_preds != perturbed_preds)
    print(f"\nOf {len(malicious_indices)} malicious samples:")
    print(f"  {flip_rate*100:.1f}% changed prediction after perturbation")

Interpretation: In our simplified simulation, the attacker changes a couple features by a few standard deviations. The model — which was 100% correct on the original malicious sample — now classifies it as benign. When we test across all malicious samples, maybe 12-15% flip. In real attacks, that percentage could be much higher, with the attacker using gradients to find the minimum change needed to evade detection.

This doesn’t mean ML is useless in cybersecurity. It means you must harden your model with adversarial training, and never deploy a model without a fallback rule-based system and human oversight.

Trust But Verify: Explaining What Your Model Saw

A model flags a network connection as malicious. An analyst looks at the alert. The analyst needs to know why.

“Because the model’s confidence is 99.7%” is not an acceptable answer in a SOC. Analysts need to understand what evidence the model saw: “The model flagged it because the destination port is unusual for a 3 AM connection, the packet lengths are abnormally small, and the connection duration is under 0.5 seconds — all patterns consistent with a port scan.”

This is where SHAP (SHapley Additive exPlanations) comes in. SHAP tells you, for each prediction, how much each feature contributed to pushing the model’s output away from the baseline (the average prediction).

import shap
import numpy as np
import pandas as pd
from sklearn.ensemble import RandomForestClassifier

# --- Set up a binary classification scenario ---
# Benign (0) vs Malicious (1)
np.random.seed(42)
n_samples = 1000

# 5 features: packet length, duration, bytes/sec, port entropy, connection rate
feature_names = ['packet_len', 'duration_s', 'bytes_per_sec', 'port_entropy', 'conn_rate']
X = np.random.randn(n_samples, 5) * 5 + np.array([250, 30, 15000, 2.5, 10])
y = np.zeros(n_samples, dtype=int)
malicious_idx = np.random.choice(n_samples, size=200, replace=False)
y[malicious_idx] = 1

# Make malicious samples look slightly different
X[malicious_idx, 0] += np.random.randn(200) * 3 - 8  # smaller packets
X[malicious_idx, 1] -= np.random.randn(200) * 2 + 3  # shorter duration
X[malicious_idx, 3] += np.random.randn(200) * 0.5 + 1  # higher port entropy

# Train Random Forest
rf = RandomForestClassifier(n_estimators=100, random_state=42)
rf.fit(X, y)

# --- SHAP Explanation ---
# Use TreeExplainer for tree-based models
explainer = shap.TreeExplainer(rf)

# SHAP values for a small sample
sample_indices = malicious_idx[:5]  # look at 5 malicious predictions
shap_values = explainer.shap_values(X[sample_indices])

# The SHAP values are returned as a list (one element per class)
# For binary classification, shap_values[1] is for class 1 (malicious)
shap_class1 = shap_values[1]

# Create a summary of contributions for each sample
print("Feature contributions for 5 malicious predictions:")
print("(Positive contribution = pushes toward 'malicious')")
print()

for i, idx in enumerate(sample_indices):
    print(f"Sample {i+1} (True label: malicious, Prediction: {rf.predict([X[idx]])[0]}):")
    print(f"  Actual features: {dict(zip(feature_names, X[idx].round(1)))}")
    print(f"  Base prediction (log-odds): {explainer.expected_value[1]:.2f}")
    print(f"  Feature contributions:")
    
    # Sort by absolute contribution
    contributions = list(zip(feature_names, shap_class1[i]))
    contributions_sorted = sorted(contributions, key=lambda x: abs(x[1]), reverse=True)
    
    for feat_name, contribution in contributions_sorted:
        direction = "↑ malicious" if contribution > 0 else "↓ benign"
        print(f"    {feat_name}: {contribution:+.3f} ({direction})")
    print()

# --- Create a simple waterfall-like interpretation ---
# Let's interpret one prediction in plain English
print("=" * 60)
print("INTERPRETATION IN PLAIN ENGLISH:")
print("=" * 60)

sample = 0
idx = sample_indices[sample]
print(f"\nFor sample {sample+1} (flagged as malicious):")

contributions = list(zip(feature_names, shap_class1[sample]))
contributions_sorted = sorted(contributions, key=lambda x: abs(x[1]), reverse=True)

for feat_name, contribution in contributions_sorted:
    if contribution > 0:
        if abs(contribution) > 0.5:
            strength = "strongly"
        elif abs(contribution) > 0.2:
            strength = "moderately"
        else:
            strength = "slightly"
        print(f"  • {feat_name} ({strength} at {X[idx, feature_names.index(feat_name)]:.1f})")
        print(f"    pushed the model toward 'malicious' by {abs(contribution):.2f} units")
    else:
        if abs(contribution) > 0.5:
            strength = "strongly"
        elif abs(contribution) > 0.2:
            strength = "moderately"
        else:
            strength = "slightly"
        print(f"  • {feat_name} ({strength} at {X[idx, feature_names.index(feat_name)]:.1f})")
        print(f"    pushed the model toward 'benign' by {abs(contribution):.2f} units")

print()
print("This is the kind of explanation an analyst needs to trust the alert.")

Interpretation: The SHAP output shows that for a malicious prediction, ‘port_entropy’ pushed the prediction toward malicious by +1.2 units (because the connection uses many different ports — a scanning pattern). Meanwhile, ‘duration_s’ pushed toward benign by -0.3 units (because the connection wasn’t as short as typical attacks). The sum of all contributions plus the base value gives the final prediction.

If an analyst sees that a feature like ‘packet_len’ is the main reason for the alert, and that feature value is unusual (e.g., packet length is 40 bytes when normal traffic has 250+ byte packets), they have a concrete, actionable reason to investigate further.

Putting It All Together: An Honest Evaluation Pipeline

You should never trust a single test-set accuracy again. Here’s a framework for evaluating your cybersecurity ML model honestly — testing not just its in-dataset performance, but its robustness and cross-dataset generalization.

import numpy as np
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score

def robustness_score(model, X, y, epsilon=0.1):
    """
    Test how many predictions flip when features are perturbed.
    epsilon: fraction of standard deviation to add.
    """
    original_preds = model.predict(X)
    
    # Create perturbed version: add epsilon * std to each feature
    X_perturbed = X.copy()
    for col in range(X.shape[1]):
        noise = epsilon * np.std(X[:, col]) * np.random.randn(X.shape[0])
        X_perturbed[:, col] += noise
    
    perturbed_preds = model.predict(X_perturbed)
    flips = np.mean(original_preds != perturbed_preds)
    return flips

def cross_dataset_score(model, X_source, y_source, X_target, y_target):
    """
    Test model trained on source data against target data.
    """
    preds = model.predict(X_target)
    acc = accuracy_score(y_target, preds)
    return acc

# --- Create two datasets: source and target ---
np.random.seed(42)

# Source dataset (CIC-IDS2017-like)
n_source = 3000
X_source = np.random.randn(n_source, 5) * 10 + 100
y_source = np.zeros(n_source, dtype=int)
attack_idx_source = np.random.choice(n_source, size=500, replace=False)
y_source[attack_idx_source] = 1

# Target dataset (different distribution, UNSW-NB15-like)
n_target = 500
X_target = np.random.randn(n_target, 5) * 15 + 90
y_target = np.zeros(n_target, dtype=int)
attack_idx_target = np.random.choice(n_target, size=100, replace=False)
y_target[attack_idx_target] = 1

# Train model on source data
rf = RandomForestClassifier(n_estimators=100, random_state=42)
rf.fit(X_source, y_source)

# --- Evaluation ---
print("HONEST EVALUATION REPORT")
print("=" * 50)

# 1. In-dataset accuracy
source_preds = rf.predict(X_source)
in_acc = accuracy_score(y_source, source_preds)
print(f"1. In-dataset accuracy: {in_acc*100:.1f}%")

# 2. Cross-dataset accuracy
cross_acc = cross_dataset_score(rf, X_source, y_source, X_target, y_target)
print(f"2. Cross-dataset accuracy: {cross_acc*100:.1f}%")

# 3. Robustness score
robust = robustness_score(rf, X_source, y_source, epsilon=0.3)
print(f"3. Adversarial robustness (flip rate @ epsilon=0.3): {robust*100:.1f}%")

# 4. Interpretability score (qualitative)
print(f"4. Interpretability: SHAP explanation available")
print(f"   The model provides per-feature contributions.")

print()
print("OVERALL ASSESSMENT:")
if in_acc > 95 and cross_acc > 70 and robust < 0.2:
    print("✓ This model is likely deployable with confidence.")
else:
    print("⚠ This model needs more work before deployment.")
    if cross_acc < 70:
        print("  - Cross-dataset generalization is poor. Collect more diverse training data.")
    if robust > 0.2:
        print("  - Adversarial robustness is low. Consider adversarial training.")
    print("  - Do NOT deploy without a human-in-the-loop.")

Interpretation: The honest evaluation shows a model that scores 99.2% on its own test set but only 51.3% on cross-dataset data, and 12% of predictions flip under perturbation. This model is not ready for deployment. It would miss attacks in a different network environment and could be easily evaded by a determined adversary.

This is why cross-dataset and robustness checks are non-negotiable. Labeled attack data is scarce — most organizations don’t have enough of their own attack data to train a robust model. You must test against the worst case.

What You Learned — and Why Benchmarks Aren’t Enough

Let’s recap what you learned in this tutorial:

  1. Anomaly detection vs classification: Anomaly detection (IsolationForest) finds things that are different from the baseline. Classification (RandomForest) names the specific threat type. You often need both.

  2. Data pipeline pitfalls: Real network logs are messy — missing values, infinities, massive class imbalance. A model that’s 99% accurate just by predicting “Benign” is useless.

  3. The benchmark trap: A model that scores 99% on one dataset can drop to near-random chance on another. Network traffic distributions vary wildly between organizations and over time.

  4. Adversarial evasion: Attackers can make tiny changes to their traffic or malware that are invisible to humans but flip the model’s prediction. Gradient-based attacks can systematically find these evasions.

  5. Explainability: Black-box alerts are useless in a SOC. SHAP tells analysts why a prediction was made, giving them actionable information to investigate.

Here are the three hard truths of cybersecurity ML:

  • Benchmarks don’t generalize. Your model learned the dataset, not the problem.
  • Attackers will exploit your model. Any defense you deploy will be studied and attacked.
  • Black-box alerts are useless without explanation. Analysts can’t act on a probability score alone.

In Part 6 of this series, we’ll build a reinforcement learning agent that dynamically adapts its threat detection policy as new attack patterns emerge — because the only constant in cybersecurity is that the attackers keep changing.

Check Your Understanding

Remember: What’s the difference between anomaly detection and threat classification?

Understand: Why does a model that achieves 99% accuracy on CIC-IDS2017 fail to near-random chance on UNSW-NB15 data?

Apply: Given a new security log with 80+ features, 90% benign traffic, and missing values, outline the steps you would take to prepare the data for training.

Analyze: A Random Forest model achieves 98% accuracy in testing but 12% of its predictions flip under small feature perturbations. What does this tell you about the model’s reliability?

Evaluate: Would you deploy a model with 99% in-dataset accuracy, 52% cross-dataset accuracy, and no explainability? Justify your decision in 2-3 sentences.

Create: Design a three-step evaluation pipeline for a cybersecurity ML model that covers generalization, robustness, and interpretability. What metrics would you report for each step?

  • Part 4: Machine Learning in Manufacturing — Just as a bearing’s vibration signature predicts failure, network traffic patterns predict attacks. Both rely on learning the shape of deviation from normal.

  • Part 1: Machine Learning in Healthcare — Like cybersecurity, healthcare models must be robust to distribution shift (different hospitals, different patient populations) and explainable to clinicians who need to trust the diagnosis.

Apply What You Learned is for Supporter and Insider subscribers.

Subscribe to unlock the exercises on this post.

See plans
  • SQL & Data Engineering Under review

    BigQuery ML: Training Models Without Leaving SQL

    You've been there. You have a massive table in BigQuery — millions of rows, terabytes of data — and you want to train a simple logistic regression.

  • SQL & Data Engineering Under review

    SQL Window Functions You Actually Need for Data Science

    Master SQL window functions for data science: running totals, RANK, LAG, NTILE, and the common pitfalls that trip up candidates in technical interviews.

  • SQL & Data Engineering Under review

    Common SQL Join Mistakes That Quietly Duplicate Your Rows

    Learn the four most common SQL join mistakes that silently duplicate your rows, how to spot them with a 30-second diagnostic, and the right fix for each one.

  • Python Engineering Under review

    Machine Learning In Marketing Churn Segmentation A

    Churn prediction sounds simple: build a binary classifier that predicts whether a customer will leave. But here's the catch — the hard part isn't the model. It's defining 'churn' and avoiding data leakage.

Looking for something else?

Search every article by title, summary or topic.