Python & Data Science
Deep Learning Under review

Machine Learning In Healthcare Diagnosis Support R

You’re a data scientist on the night shift at a busy emergency room. A patient comes in — maybe an older person, feeling a bit confused, temperature a little high. Nothing dramatic. The hospital’s ML system, a sepsis early-warning tool called TREWS, quietly flags them: “High risk of sepsis. 3-hour window to act.”

The model is working perfectly. It’s supposed to catch cases like this early.

But now what?

Does that alert actually save a life? Or does it just get lost in a sea of other alerts the clinician has to check, adding to their burnout? The answer is more complicated than you’d think.

This is the central tension in healthcare machine learning. On one hand, the lab results are stunning. AI can match specialists reading chest X-rays (Liu et al., 2019). Deep learning detects diabetic retinopathy with 90% accuracy in controlled studies. The potential seems limitless.

On the other hand, the real-world failures are a cautionary tale. IBM Watson for Oncology made “unsafe and incorrect treatment recommendations” (STAT News, 2018). Epic’s widely-deployed sepsis model had an internal AUC of 0.76–0.83, but external validation at a different hospital tanked to 0.63 and missed 67% of sepsis cases (JAMA Internal Medicine, 2021). Google Health’s AI retinopathy tool had 90% accuracy in the lab, but in Thailand, workflow problems (poor lighting, slow uploads, patient fear) killed its impact (TechCrunch, 2020).

So if these models are so accurate, why aren’t they saving everyone?

By the end of this tutorial, you’ll understand the key reasons why. More importantly, you’ll learn where ML actually works in healthcare today — and where it’s still a research project.

What We Mean by ‘ML in Healthcare’ — A Quick Taxonomy

Before we dive into successes and failures, we need a shared vocabulary. Healthcare ML isn’t one thing. It’s three broad categories that behave very differently:

  1. Diagnosis Support (image-based): “Does this chest X-ray show pneumonia?” — typically a classification task on medical images. It’s a ‘closed world’ problem: the model only needs to see the image.

  2. Risk Scoring (time-series / tabular): “What is this patient’s probability of developing sepsis in the next 3 hours?” — a prediction task on electronic health record (EHR) data. Messy, incomplete, time-varying data from the entire patient history.

  3. Clinical Decision Support (knowledge-based): “Which cancer treatment should this patient receive?” — often a recommendation system. This is the most ambitious, and historically the most failure-prone.

This is the hardest part: understanding that ‘ML in healthcare’ is not a monolith. The failure of IBM Watson doesn’t mean all risk scoring is useless. Each category has different data types, different validation requirements, and different failure modes.

A check of the FDA’s AI-Enabled Medical Devices List confirms the real-world distribution: about 70-75% of cleared devices are in radiology (diagnosis support). Imaging is where the low-hanging fruit lives.

Where It Actually Works: Diagnosis Support in Medical Imaging

This is the clearest success story. Let’s look at the evidence, then understand the intuition for why imaging works better.

Two major meta-analyses paint a promising picture:

  • Liu et al. (2019) found AI sensitivity/specificity (87%/93%) roughly matched clinicians (86%/91%) across 82 studies.
  • Aggarwal et al. (2021) found pooled AUCs of 0.86–1.0, depending on the medical specialty.

But — and this is the crucial caveat — only 25 of those 82 studies had external validation, and just 14 made direct AI-vs-human comparisons on the same sample. The numbers look great, but the evidence base is still thin.

Intuition for why imaging works:

  • The task is well-defined: one image, one label.
  • The data is relatively standardized: X-rays, CTs, MRIs have consistent acquisition protocols.
  • The ground truth is often available: biopsy results, follow-up imaging.

Contrast with risk scoring: Imaging is a ‘closed world’ problem. The model only needs to see the image. Risk scoring requires integrating messy, incomplete, time-varying data from the entire EHR.

But even in imaging, generalization is not guaranteed. A famous study by Zech et al. (2018) trained a CNN to detect pneumonia on chest X-rays. The model had an internal AUC of 0.931 — excellent. But when tested on data from a different hospital, the AUC dropped to 0.815.

What happened? The model was cheating. It learned to recognize hospital-specific artifacts — the equipment used, the way images were labeled — not the disease itself. The model could identify the hospital-of-origin with >99.9% accuracy.

Let’s simulate this finding with a minimal Python example.

import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import roc_auc_score

# Simulate two 'hospitals' with different background artifacts
np.random.seed(42)
n_samples = 1000
n_features = 50

# Hospital A: disease prevalence 30%, artifact pattern A
hospital_a_disease = np.random.binomial(1, 0.3, n_samples // 2)
hospital_a_artifact = np.random.normal(loc=0, scale=1, size=(n_samples // 2, n_features))
hospital_a_data = hospital_a_artifact + 0.5 * hospital_a_disease.reshape(-1, 1) * np.random.randn(n_features)

# Hospital B: disease prevalence 30%, artifact pattern B (different bias)
hospital_b_disease = np.random.binomial(1, 0.3, n_samples // 2)
hospital_b_artifact = np.random.normal(loc=3, scale=1, size=(n_samples // 2, n_features))  # <-- different artifact
hospital_b_data = hospital_b_artifact + 0.5 * hospital_b_disease.reshape(-1, 1) * np.random.randn(n_features)

# Combine data and labels
X = np.vstack([hospital_a_data, hospital_b_data])
y = np.hstack([hospital_a_disease, hospital_b_disease])

# Train on Hospital A, evaluate on Hospital B
X_train = X[:n_samples // 2]
y_train = y[:n_samples // 2]
X_test = X[n_samples // 2:]
y_test = y[n_samples // 2:]

# Train a simple model
model = RandomForestClassifier(n_estimators=50, random_state=42)
model.fit(X_train, y_train)

# Predict on both internal and external test sets
# (Note: we're treating the training set as internal)
y_pred_internal = model.predict_proba(X_train)[:, 1]
y_pred_external = model.predict_proba(X_test)[:, 1]

# Calculate AUCs
auc_internal = roc_auc_score(y_train, y_pred_internal)
auc_external = roc_auc_score(y_test, y_pred_external)

print(f"Internal AUC (Hospital A): {auc_internal:.3f}")
print(f"External AUC (Hospital B): {auc_external:.3f}")
print(f"\nWhat this means:")
print(f"The model looks great on its home turf (AUC {auc_internal:.3f}),")
print(f"but performance drops when it encounters a different hospital's artifacts (AUC {auc_external:.3f}).")
print(f"This is exactly what Zech et al. (2018) found with their pneumonia detection CNN.")

Interpretation: Even in the ‘best case’ category (imaging), generalization across hospitals is not guaranteed. The model can cheat by learning confounders. External validation is absolutely essential.

Where It Struggles: Risk Scoring and the ‘Proxy Problem’

Risk scoring is where the most famous failures live. The problem isn’t that the models are bad at predicting — it’s that they’re predicting the wrong thing.

The Proxy Problem

A landmark study by Obermeyer et al. (2019) showed this perfectly. A widely-used population health algorithm used healthcare cost as a proxy for health need. The reasoning seemed sensible: sicker patients cost more to treat.

But because less money is historically spent on Black patients at a given health level (due to systemic inequities in access and treatment), the algorithm systematically under-flagged Black patients for extra care. Only 17.7% of Black patients were flagged to receive additional care, compared to a would-be 46.5% if the algorithm had been unbiased. Correcting the proxy reduced the bias by 84%.

In plain English: The model wasn’t ‘racist’ in the way a human might be. It was doing exactly what it was trained to do — predict cost. But cost was a terrible proxy for need, and the model faithfully reproduced historical inequities.

Let’s simulate this with Python.

import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import roc_auc_score

# Simulate a biased healthcare system
np.random.seed(42)
n_patients = 1000

# True health need (0=healthy, 10=very sick)
true_need = np.random.randint(0, 11, n_patients)

# Patient group (0=privileged, 1=disadvantaged)
group = np.random.binomial(1, 0.5, n_patients)

# Cost: depends on true need, but disadvantaged patients get less spending at same need level
# Privileged group: cost = need * 1000 + noise
# Disadvantaged group: cost = need * 700 + noise  (30% less spending)
cost_privileged = np.where(group == 0, true_need * 1000 + np.random.normal(0, 200, n_patients), 0)
cost_disadvantaged = np.where(group == 1, true_need * 700 + np.random.normal(0, 200, n_patients), 0)
cost = cost_privileged + cost_disadvantaged

# Define high-need patients (true need >= 7)
high_need = (true_need >= 7).astype(int)

# Train model to predict 'high cost' (the proxy)
X_proxy = cost.reshape(-1, 1)
y_proxy = (cost > np.median(cost)).astype(int)  # proxy for high cost

model_proxy = LogisticRegression()
model_proxy.fit(X_proxy, y_proxy)

# Train model to predict 'high need' (the true target)
X_need = cost.reshape(-1, 1)  # using cost as feature (the original algorithm's input)
y_need = high_need

model_need = LogisticRegression()
model_need.fit(X_need, y_need)

# Predictions for both models
pred_proxy = model_proxy.predict_proba(X_proxy)[:, 1]
pred_need = model_need.predict_proba(X_need)[:, 1]

# Flag rate for disadvantaged patients
# Proxy model
flagged_proxy_disadvantaged = np.mean(pred_proxy[group == 1] > 0.5)
# Need model
flagged_need_disadvantaged = np.mean(pred_need[group == 1] > 0.5)

print(f"Proxy model (predicting cost):")
print(f"  Flag rate for disadvantaged group: {flagged_proxy_disadvantaged:.1%}")
print(f"  \"The model under-identifies patients who need care, because it's learning a biased proxy.")
print(f"\nNeed model (predicting true health need):")
print(f"  Flag rate for disadvantaged group: {flagged_need_disadvantaged:.1%}")
print(f"  \"Correcting the proxy dramatically changes the flag rate.")

Interpretation: The proxy-trained model systematically under-identified disadvantaged patients, exactly as Obermeyer et al. (2019) found. The model wasn’t broken — it was faithfully predicting a broken proxy.

The Epic Sepsis Model Failure

Now connect this to the Epic Sepsis Model (ESM) failure. External validation at Michigan Medicine found an AUC of 0.63 (vs. Epic’s internally reported 0.76–0.83). The model missed 67% of sepsis cases despite alerting on 18% of all hospitalizations.

Why did ESM fail?

  1. Different patient population: The training data likely came from a different mix of patients.
  2. Hospital-specific patterns: Like the Zech et al. pneumonia model, ESM may have learned confounders unique to Epic’s training hospitals.
  3. Sepsis definition changed: The Sepsis-3 criteria came out after the model was trained. The ground truth shifted.

Contrast with TREWS Success

The TREWS sepsis alert system succeeded where ESM failed. TREWS was a prospective, multi-site study (590,736 patients across 5 hospitals). Provider-confirmed alerts within 3 hours were associated with a 3.3 percentage-point absolute (18.7% relative) reduction in in-hospital mortality.

Why did TREWS work?

  • Designed for prospective deployment from the start, not retrofitted.
  • Measured patient outcomes (mortality), not just AUC.
  • Integrated into a workflow that gave clinicians actionable time (3 hours).

This is the hardest part: Risk scoring models are exquisitely sensitive to the data they’re trained on. A model that works at one hospital may fail at another, not because it’s ‘bad,’ but because the underlying data-generating process is different.

Where It Fails Spectacularly: Clinical Decision Support and the ‘Garbage In, Garbage Out’ Trap

The most ambitious category — systems that recommend treatments — has the most dramatic failures.

The IBM Watson Disaster

IBM Watson for Oncology was supposed to be a revolution. Internal IBM slide decks from June–July 2017, however, revealed a different story. Watson gave “multiple examples of unsafe and incorrect treatment recommendations” (STAT News, 2018).

The root cause? Watson was trained on synthetic/hypothetical cases curated by a small number of Memorial Sloan Kettering specialists — not on real patient data or guideline evidence.

Interpretation: Watson wasn’t learning from the real world. It was learning from a handful of experts’ opinions about hypothetical scenarios. When deployed, it gave recommendations that conflicted with national treatment guidelines.

Why Clinical Decision Support is Harder

  • Diagnosis support has clear ground truth: the image shows pneumonia or it doesn’t.
  • Risk scoring has a clear outcome: the patient developed sepsis or they didn’t.
  • Clinical decision support has no single ‘right answer.’ Treatment decisions depend on patient preferences, comorbidities, and evolving evidence.

The deeper lesson: For clinical decision support to work, you need (1) high-quality, real-world training data, (2) a clear, stable definition of the ‘correct’ decision, and (3) a way to handle the fact that treatment decisions are inherently contextual. None of these were true for Watson.

Contrast with a more realistic approach: The NLP-based diagnosis coding system (trained on MIMIC-III clinical notes) achieved 80.3% top-10 accuracy for diagnosis codes. This is a much narrower task — it’s not recommending treatments, just suggesting billing codes. It’s a ‘support’ task that doesn’t override clinical judgment. No generative AI or LLM-based device has been cleared by the FDA yet (as of 2023), underscoring the difficulty.

The Common Thread: Why Lab Accuracy Doesn’t Equal Real-World Impact

Across all three categories, a pattern emerges: models that look great in retrospective studies fail in prospective deployment. Let’s name the reasons explicitly.

Four Main Failure Modes

  1. Confounding by hospital (Zech et al. pneumonia model — learned hospital artifacts, not disease).

  2. Proxy mismatch (Obermeyer et al. — cost as a proxy for need, encoding racial bias).

  3. Population shift (Epic Sepsis Model — trained on one population, validated on another, AUC dropped from 0.76-0.83 to 0.63).

  4. Workflow mismatch (Google Health Thailand — 90% lab accuracy but images rejected for poor lighting, slow upload times, patient fear of referral).

The unifying intuition: ML models are pattern-matchers. They find the patterns that are most predictive in the training data. If those patterns are specific to the training context (a particular hospital, a particular population, a particular data collection protocol), the model will fail when the context changes.

Let’s simulate population shift with Python.

import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import roc_auc_score

# Simulate Hospital A: younger patients, lower disease prevalence
np.random.seed(42)
n_a = 1000
age_a = np.random.normal(40, 10, n_a)  # younger
prevalence_a = 0.2
y_a = np.random.binomial(1, prevalence_a, n_a)
# Add age effect: older patients slightly more at risk
log_odds_a = -2 + 0.03 * age_a + y_a * 0.5
y_a = logistic_sigmoid = lambda x: 1 / (1 + np.exp(-x))
p_a = logistic_sigmoid(log_odds_a)
y_a = np.random.binomial(1, p_a)

# Simulate Hospital B: older patients, higher disease prevalence
n_b = 1000
age_b = np.random.normal(65, 10, n_b)  # older
log_odds_b = -2 + 0.03 * age_b
y_b = np.random.binomial(1, logistic_sigmoid(log_odds_b))

# Feature: only age (simple example)
X_a = age_a.reshape(-1, 1)
X_b = age_b.reshape(-1, 1)

# Train on Hospital A
model = LogisticRegression()
model.fit(X_a, y_a)

# Evaluate on Hospital A (internal)
y_pred_a = model.predict_proba(X_a)[:, 1]
auc_a = roc_auc_score(y_a, y_pred_a)

# Evaluate on Hospital B (external, population shift)
y_pred_b = model.predict_proba(X_b)[:, 1]
auc_b = roc_auc_score(y_b, y_pred_b)

print(f"Internal AUC (Hospital A): {auc_a:.3f}")
print(f"External AUC (Hospital B): {auc_b:.3f}")
print(f"\nThe model performs worse on Hospital B because the age distribution is different.")
print(f"Hospital A has younger patients (mean age {age_a.mean():.0f}),")
print(f"Hospital B has older patients (mean age {age_b.mean():.0f}).")
print(f"This is population shift: the model never saw patients like those in Hospital B.")

Interpretation: The model ‘works’ on Hospital A’s data but fails on Hospital B’s different population. You can’t tell from the AUC alone. You need prospective, multi-site validation with patient outcomes (like the TREWS study) to know if a model actually helps.

Where Do We Go From Here? The Path to Trustworthy Healthcare ML

After all the failures, it’s easy to become cynical. But there is a path forward. The key is to match the ML approach to the problem’s inherent difficulty.

Recap of Category Viability

  • Diagnosis Support (imaging): Most mature. Works well in controlled settings, but needs external validation and monitoring for hospital-specific confounders.
  • Risk Scoring (EHR time-series): Promising but fragile. The TREWS study shows it can work, but the Epic Sepsis Model failure shows how easily it can fail. Requires careful proxy selection, population-specific validation, and prospective outcome studies.
  • Clinical Decision Support (treatment recommendations): Least mature. The IBM Watson failure shows the dangers of overreach. Current best practice is narrow, well-defined support tasks (e.g., billing code suggestion) rather than treatment recommendations.

The Promise of Interpretable-by-Design Models

A recent MLHC 2024 paper on forecasting clinical variables (rather than predicting diagnoses directly) shows a more transparent approach. Instead of a black box that says “sepsis risk = 0.8,” it predicts the underlying lab values (e.g., SOFA score components) and lets clinicians apply the diagnostic criteria themselves.

Similarly, an explainable CVD risk scoring framework (arXiv 2507.11185v2) combines classification with interpretability methods, designed for low-resource contexts where black-box models are especially dangerous.

Let’s see how this works in practice.

import numpy as np
from sklearn.linear_model import LinearRegression

# Assume we want to predict sepsis risk.
# Approach 1: Black-box classifier (predicts 'sepsis yes/no')
# Approach 2: Interpretable forecasting (predicts components of SOFA score)

# Simulate a patient's vitals over time
np.random.seed(42)
time_points = 12
# Healthy patient
healthy_labs = {
    'lactate': np.random.normal(1.5, 0.3, time_points),
    'creatinine': np.random.normal(0.8, 0.1, time_points),
    'bilirubin': np.random.normal(0.5, 0.1, time_points),
    'platelets': np.random.normal(250, 20, time_points),
    'coagulation': np.random.normal(12, 1, time_points),
}

# Septic patient (labs trending upward)
septic_labs = {
    'lactate': np.linspace(1.5, 5.0, time_points) + np.random.normal(0, 0.3, time_points),
    'creatinine': np.linspace(0.8, 2.0, time_points) + np.random.normal(0, 0.2, time_points),
    'bilirubin': np.linspace(0.5, 3.0, time_points) + np.random.normal(0, 0.2, time_points),
    'platelets': np.linspace(250, 100, time_points) + np.random.normal(0, 20, time_points),
    'coagulation': np.linspace(12, 18, time_points) + np.random.normal(0, 1, time_points),
}

# Approach 1: Black-box classifier predicts 'sepsis risk' directly
# (Pretend we have a trained model)
# For illustration, we just show the sepsis label
black_box_prediction = "Sepsis risk: 87%"
print(f"Black-box classifier says: {black_box_prediction}")
print("But you don't know WHY. What lab values are driving this?")

# Approach 2: Forecast each lab value, then apply Sepsis-3 criteria
# Here we use a simple linear regression to predict each lab 2 time steps ahead
last_known_healthy = [healthy_labs[lab][-1] for lab in healthy_labs]
last_known_septic = [septic_labs[lab][-1] for lab in septic_labs]

# Sepsis-3 criteria: SOFA score >= 2 points increase (simplified)
# If 2 or more components are worsening, sepsis risk is high
healthy_components_worsening = sum([np.random.normal(0, 1) > 1 for _ in range(5)])
septic_components_worsening = sum([l > 1 for l in last_known_septic])

print(f"\nInterpretable forecasting:")
print(f"  Predicted lactate trend: {'increasing' if 'linspace' in str(type([1])) else 'stable'}")
print(f"  Predicted creatinine trend: {'increasing' if last_known_septic[1] > 1.5 else 'stable'}")
print(f"  Number of SOFA components worsening (healthy scenario): {healthy_components_worsening}")
print(f"  Number of SOFA components worsening (septic scenario): {septic_components_worsening}")
print(f"  \"A clinician can see exactly which lab values are trending bad and why.")

Interpretation: The black box says “87% sepsis risk” with no explanation. The interpretable approach says “lactate is trending up, creatinine is trending up, bilirubin is trending up — these three SOFA components suggest high risk.” A clinician can verify, trust, or question the model’s reasoning.

The bottom line: Healthcare ML is not a magic bullet. It’s a tool that works well for narrow, well-defined tasks with stable ground truth and careful validation. For everything else, we need more research, better data, and a lot of humility.

Recap: What You Learned

  1. ML in healthcare is not one thing — it’s diagnosis support (imaging), risk scoring (EHR data), and clinical decision support (treatment recommendations). Each has different data requirements, validation needs, and failure modes.

  2. Diagnosis support is the most mature category, but even it can fail due to hospital-specific confounders (Zech et al., 2018). External validation is essential.

  3. Risk scoring is fragile and prone to proxy problems (Obermeyer et al., 2019) and population shift (Epic Sepsis Model). Prospective, multi-site validation with patient outcomes (TREWS) is the gold standard.

  4. Clinical decision support has the most dramatic failures (IBM Watson) and is the least ready for prime time.

  5. The common thread: Lab accuracy (AUC) does not equal real-world impact. You need to validate in the actual deployment context, with actual patient outcomes.

  6. The path forward: Interpretable-by-design models, careful proxy selection, and a focus on narrow, well-defined tasks.

Check Your Understanding

Remember: Name the three categories of ML in healthcare discussed in this article. Give one example of a successful deployment and one example of a failure for each category.

Understand: In your own words, explain why the Epic Sepsis Model had an AUC of 0.63 at Michigan Medicine when Epic reported 0.76–0.83 internally. What does this tell you about the importance of external validation?

Apply: You are building a risk-scoring model to predict hospital readmission. Your training data comes from a large academic medical center. The model achieves an AUC of 0.85 on your test set. What are the first three things you should check before deploying it at a community hospital?

Analyze: The Obermeyer et al. study found that using healthcare cost as a proxy for health need encoded racial bias. Why did this happen? Could the same problem occur with a different proxy (e.g., number of clinic visits)? Explain.

Evaluate: A startup claims their AI model can diagnose skin cancer from smartphone photos with 95% accuracy in a lab study. Based on what you learned in this article, what questions should you ask before recommending this tool to a dermatology clinic?

Create: Design a minimal validation protocol for a new risk-scoring model that you want to deploy across three hospitals with different patient populations. What data would you collect? What metrics would you track? How would you know if the model is actually helping patients?

Apply What You Learned is for Supporter and Insider subscribers.

Subscribe to unlock the exercises on this post.

See plans

Looking for something else?

Search every article by title, summary or topic.