Python & Data Science
Statistics Under review

Machine Learning In Education Adaptive Learning An

You’re a college freshman, sitting in a lecture hall with 300 other students. The professor is moving through slides faster than you can follow. You raise your hand, but the question you have is about something from last week. The professor nods, says “we’ll cover that later,” and moves on. You don’t ask again. By mid-semester, you’ve stopped checking the online portal. Your grades are slipping. Nobody notices until it’s too late.

Now imagine a different world. Your learning platform adapts to you: if you struggle with a concept, it gives you more practice at a slower pace. If you breeze through, it moves ahead. And before you even realize you’re disengaging, an advisor emails you: “I noticed your login frequency dropped. Want to grab coffee and talk about how things are going?”

That second world isn’t science fiction. It’s what happens when machine learning meets education. Two ML pillars make it possible: adaptive learning (personalizing content) and dropout prediction (early warning systems). Georgia State University’s early-warning system, for example, boosted graduation rates by 23 percentage points (Hechinger Report, 2019). That’s not a small tweak — it’s a transformation.

In this article, we’ll walk through both pillars. You’ll learn the intuition behind how ML “sees” a student, see working code for a Bayesian Knowledge Tracing model (adaptive learning) and a dropout predictor with SHAP explanations, and — most importantly — understand the ethical landmines you must avoid if you ever build one of these systems.


Intuition First: How ML Sees a Student

Before any math, let’s build a mental model. Think of a student as a bundle of latent traits — things you can’t directly observe, like their knowledge of calculus, their motivation level, or their risk of dropping out. What you can observe are visible actions: quiz answers, time spent on a page, login frequency, grade changes.

ML’s job is to infer the hidden traits from the breadcrumbs. It’s like a doctor diagnosing an illness from symptoms. The patient doesn’t walk in saying “I have strep throat” — they say “my throat hurts.” The doctor infers the hidden cause from visible signs.

But here’s the hard part: educational data is sparse, noisy, and non-stationary. A student might get a question right because they guessed (noise), or they might suddenly improve after a tutoring session (non-stationary). The model has to separate signal from noise while the student is changing.

We’re basically trying to read a student’s mind from a few breadcrumbs. That’s both exciting and terrifying.


Adaptive Learning: Knowledge Tracing from BKT to DKT

Adaptive learning systems adjust difficulty, pace, and content based on what a student currently knows. The core technology is knowledge tracing — estimating a student’s mastery of each skill over time.

Bayesian Knowledge Tracing (BKT)

BKT is the classic approach, dating back to the 1990s. It models each skill as a hidden state (mastered or not mastered) and updates that state after each answer. It has four intuitive parameters:

  • Learn rate: probability of transitioning from “not mastered” to “mastered” after a practice opportunity.
  • Guess: probability of getting a correct answer even if the skill is not mastered.
  • Slip: probability of getting a wrong answer even if the skill is mastered.
  • Forget: probability of forgetting a mastered skill (often set to zero in practice).

Let’s see this in action with a simple simulation.

# --- Bayesian Knowledge Tracing: a simple update loop ---
import numpy as np

def bkt_update(p_mastered, correct, learn_rate=0.2, guess=0.15, slip=0.1, forget=0.0):
    """
    Update the probability that a student has mastered a skill
    given their latest answer (correct=1 or 0).
    
    Parameters:
    - p_mastered: prior belief (0 to 1)
    - correct: 1 if answer was correct, 0 if wrong
    - learn_rate, guess, slip, forget: BKT parameters
    
    Returns: updated p_mastered
    """
    # Step 1: Apply forget (if any) before this practice
    p_mastered = p_mastered * (1 - forget) + (1 - p_mastered) * forget
    
    # Step 2: Probability of observing this answer given current state
    # P(correct | mastered) = 1 - slip
    # P(correct | not mastered) = guess
    if correct == 1:
        p_obs_given_mastered = 1 - slip
        p_obs_given_not = guess
    else:
        p_obs_given_mastered = slip
        p_obs_given_not = 1 - guess
    
    # Step 3: Bayes' rule to update belief
    # P(mastered | observation) = P(obs|mastered)*P(mastered) / P(obs)
    p_obs = p_obs_given_mastered * p_mastered + p_obs_given_not * (1 - p_mastered)
    p_mastered_given_obs = (p_obs_given_mastered * p_mastered) / p_obs
    
    # Step 4: Apply learning (transition to mastered if not already)
    # After practice, there's a chance the student learns
    p_mastered_after = p_mastered_given_obs + (1 - p_mastered_given_obs) * learn_rate
    
    return p_mastered_after

# Simulate a student practicing "solving linear equations"
# Start with low belief (10% mastered)
p_mastered = 0.1
print(f"Initial belief of mastery: {p_mastered:.0%}")

# Sequence of answers: correct, correct, wrong, correct, correct
answers = [1, 1, 0, 1, 1]
for i, ans in enumerate(answers, 1):
    p_mastered = bkt_update(p_mastered, ans)
    print(f"After answer {i} ({'correct' if ans else 'wrong'}): belief = {p_mastered:.2%}")

print(f"\nFinal belief after 5 practices: {p_mastered:.2%}")

What the output means: After three correct answers in a row, the model believes the student has mastered the skill with high probability. But a wrong answer (the third) drops belief — though not all the way to zero, because the model accounts for the possibility of a “slip” (the student knew it but made a careless mistake). This is the power of BKT: it’s interpretable. You can see exactly why the model thinks what it thinks.

Deep Knowledge Tracing (DKT)

BKT has limitations: it assumes each skill is independent, and it requires hand-picked parameters. Deep Knowledge Tracing (DKT) replaces the hand-crafted rules with a recurrent neural network (LSTM). The LSTM takes a sequence of question-answer pairs and outputs the probability of answering correctly on the next question. It can capture complex patterns — like a student who struggles with fractions but excels at decimals — without explicit skill definitions.

Here’s a conceptual example using Keras (simplified for illustration).

# --- Deep Knowledge Tracing: conceptual Keras model ---
# This is not meant to be run without real data, but shows the structure.

import tensorflow as tf
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import LSTM, Dense

# Assume we have 100 possible questions (one-hot encoded)
num_questions = 100
# Each input is a vector of length 2*num_questions: one-hot of question + correctness
input_dim = 2 * num_questions

model = Sequential([
    LSTM(200, input_shape=(None, input_dim), return_sequences=True),
    Dense(num_questions, activation='sigmoid')  # probability correct for each question
])

model.compile(optimizer='adam', loss='binary_crossentropy')
# In practice, you'd train on sequences of student interactions
# model.fit(X_train, y_train, ...)

print("DKT model architecture created.")
print("Input: (batch, timesteps, features)")
print("Output: (batch, timesteps, num_questions) — probability correct for each question at each step")

Trade-off: BKT is interpretable (you can see why it thinks a student knows something), but DKT is more accurate. In practice, many systems use a hybrid: BKT for core skills where interpretability matters, DKT for fine-grained prediction.

Real-world example: Duolingo uses a different approach called “half-life regression” — it models the forgetting curve and schedules review at the optimal time. You can explore their open-source implementation on GitHub.


Dropout Prediction: Building a Classifier from Scratch

Now let’s switch to the second pillar: predicting which students are at risk of dropping out. We’ll use a synthetic dataset inspired by the UCI “Predict Students’ Dropout and Academic Success” dataset (4,424 records, 36 features). The real dataset includes demographic, academic, and economic indicators. For reproducibility, we’ll generate a similar dataset with the same structure.

Load and explore the data

# --- Generate synthetic student data (inspired by UCI dataset) ---
import numpy as np
import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler

np.random.seed(42)
n = 4424  # same size as UCI dataset

# Features (simplified version of UCI features)
data = {
    'age_at_enrollment': np.random.randint(18, 60, n),
    'gender': np.random.choice([0, 1], n, p=[0.6, 0.4]),  # 0=male, 1=female
    'previous_grade_avg': np.random.uniform(5, 20, n),  # scale 0-20
    'tuition_fees_up_to_date': np.random.choice([0, 1], n, p=[0.2, 0.8]),
    'scholarship_holder': np.random.choice([0, 1], n, p=[0.7, 0.3]),
    'debtor': np.random.choice([0, 1], n, p=[0.8, 0.2]),
    'curricular_units_1st_sem_approved': np.random.randint(0, 20, n),
    'curricular_units_1st_sem_grade': np.random.uniform(0, 20, n),
    'displaced': np.random.choice([0, 1], n, p=[0.7, 0.3]),
    'educational_special_needs': np.random.choice([0, 1], n, p=[0.95, 0.05]),
}

# Target: dropout (1) vs not dropout (0) — roughly 30% dropout rate
# We'll make dropout more likely for students with low grades, fees not up to date, etc.
log_odds = (
    -0.3 * data['previous_grade_avg']
    - 1.5 * data['tuition_fees_up_to_date']
    + 0.5 * data['debtor']
    - 0.1 * data['curricular_units_1st_sem_approved']
    + 0.2 * data['displaced']
    + np.random.normal(0, 1, n)
)
prob_dropout = 1 / (1 + np.exp(-log_odds))
data['dropout'] = (np.random.uniform(0, 1, n) < prob_dropout).astype(int)

df = pd.DataFrame(data)
print(f"Dataset shape: {df.shape}")
print(f"Dropout rate: {df['dropout'].mean():.1%}")
print(df.head())

Handle class imbalance with SMOTE

Only about 30% of students drop out. If we train a classifier on this imbalanced data, it will learn to always predict “not dropout” and get 70% accuracy — useless. We’ll use SMOTE (Synthetic Minority Oversampling Technique) to create synthetic samples of the minority class.

# --- Train/test split and apply SMOTE ---
from imblearn.over_sampling import SMOTE
from sklearn.linear_model import LogisticRegression
from xgboost import XGBClassifier
from sklearn.metrics import roc_auc_score, classification_report

# Separate features and target
X = df.drop('dropout', axis=1)
y = df['dropout']

# Split (stratified to preserve class proportions)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42, stratify=y
)

print(f"Training set size: {X_train.shape[0]}")
print(f"Test set size: {X_test.shape[0]}")
print(f"Training set dropout rate: {y_train.mean():.1%}")

# Apply SMOTE to training set only
smote = SMOTE(random_state=42)
X_train_resampled, y_train_resampled = smote.fit_resample(X_train, y_train)

print(f"\nAfter SMOTE:")
print(f"Training set size: {X_train_resampled.shape[0]}")
print(f"Dropout rate: {y_train_resampled.mean():.1%}")

Train two models: Logistic Regression (interpretable) and XGBoost (high accuracy)

# --- Train models ---

# Scale features for logistic regression
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train_resampled)
X_test_scaled = scaler.transform(X_test)

# Model 1: Logistic Regression with L1 regularization (LASSO) for sparsity
lr = LogisticRegression(penalty='l1', solver='saga', C=1.0, max_iter=1000, random_state=42)
lr.fit(X_train_scaled, y_train_resampled)

# Predictions
lr_pred_proba = lr.predict_proba(X_test_scaled)[:, 1]
lr_auc = roc_auc_score(y_test, lr_pred_proba)
print(f"Logistic Regression AUC: {lr_auc:.3f}")

# Model 2: XGBoost (no scaling needed)
xgb = XGBClassifier(n_estimators=100, max_depth=4, learning_rate=0.1, random_state=42, eval_metric='logloss')
xgb.fit(X_train_resampled, y_train_resampled)

xgb_pred_proba = xgb.predict_proba(X_test)[:, 1]
xgb_auc = roc_auc_score(y_test, xgb_pred_proba)
print(f"XGBoost AUC: {xgb_auc:.3f}")

# Interpretation: AUC of 0.85 means that 85% of the time, a randomly chosen dropout student
# gets a higher risk score than a randomly chosen non-dropout.

Interpret with SHAP

XGBoost is a black box. Let’s open it with SHAP to see which features drive predictions.

# --- SHAP feature importance ---
import shap

# Create a SHAP explainer for XGBoost (using TreeExplainer)
explainer = shap.TreeExplainer(xgb)
shap_values = explainer.shap_values(X_test)

# Summary plot (top 10 features)
shap.summary_plot(shap_values, X_test, max_display=10, show=False)
print("SHAP summary plot generated (not displayed in this console).")
print("\nTop features by mean absolute SHAP value:")
mean_shap = np.abs(shap_values).mean(axis=0)
feature_importance = pd.DataFrame({
    'feature': X_test.columns,
    'mean_abs_shap': mean_shap
}).sort_values('mean_abs_shap', ascending=False)
print(feature_importance.head(10).to_string(index=False))

What this tells us: The most important features are typically previous grades, tuition fees up to date, and number of approved courses. This makes intuitive sense: students who are already struggling academically and have financial issues are more likely to drop out.


The Hard Part: Fairness, Privacy, and the Risk of False Alarms

Now for the uncomfortable truth. These models can reinforce bias, invade privacy, and cause harm if deployed carelessly.

Fairness: Dropout prediction models can have different false-positive rates across racial or income groups. A model might flag more students from a particular demographic as “at risk” when they’re actually fine, leading to unnecessary interventions that could stigmatize them. Research shows that including protected attributes (like race) in the model doesn’t improve accuracy much and gives only marginal fairness gains (arXiv 2103.15237). The best practice is to audit your model for disparate impact and consider fairness constraints during training.

Privacy: Georgia State’s system tracked wifi logins and library swipes. Students may not know they’re being monitored. The line between helpful intervention and surveillance is thin. Always get informed consent and anonymize data where possible.

False alarms: Labeling a student as at-risk when they’re fine can cause anxiety, stigmatization, or even a self-fulfilling prophecy. The cost of a false positive is real. That’s why these systems should always be paired with human judgment — advisors who can interpret the risk score and decide on an appropriate response.

Let’s compute a simple fairness metric on our synthetic data to see if false-positive rates differ by gender.

# --- Fairness audit: false-positive rate by gender ---
# We'll use the logistic regression predictions
from sklearn.metrics import confusion_matrix

# Get binary predictions (threshold 0.5)
lr_pred_binary = (lr_pred_proba >= 0.5).astype(int)

# Add predictions to test set
test_df = X_test.copy()
test_df['dropout_true'] = y_test.values
test_df['dropout_pred'] = lr_pred_binary
test_df['gender'] = X_test['gender']  # gender is in features

# Compute false-positive rate for each gender
for gender_val in [0, 1]:
    subset = test_df[test_df['gender'] == gender_val]
    tn, fp, fn, tp = confusion_matrix(subset['dropout_true'], subset['dropout_pred']).ravel()
    fpr = fp / (fp + tn) if (fp + tn) > 0 else 0
    gender_label = 'Male' if gender_val == 0 else 'Female'
    print(f"False-positive rate for {gender_label}: {fpr:.2%}")

If the false-positive rates differ significantly, you have a fairness problem. In practice, you’d also check by race, income, and other protected attributes.


Putting It All Together: A Hypothetical Integrated System

Imagine a system that combines adaptive learning and dropout prediction. Here’s how it could work:

  1. Adaptive learning (BKT or DKT) tracks each student’s knowledge state in real time. When a student struggles with a topic, the system offers easier problems or additional resources.

  2. Dropout prediction runs in the background, using features like login frequency, time spent, and recent grades to update a risk score. If the risk score crosses a threshold, the system triggers an intervention: an email from an advisor, a meeting invitation, or a check-in call.

  3. The two systems talk to each other. A student who suddenly stops engaging with adaptive content might be flagged by the dropout model even before their grades drop. Conversely, a student who is flagged as at-risk but is actively practicing on the adaptive platform might be a false alarm — the model can adjust.

Research like the peer-inspired GNN (R²GCN) shows how a graph-based approach could power both tasks from the same interaction network. But for now, even a simple combination of BKT and logistic regression can make a real difference.


Conclusion: What You Learned and Where to Go Next

Let’s recap what you’ve learned:

  • Intuition behind knowledge tracing: BKT uses four intuitive parameters (learn, guess, slip, forget) to update belief about mastery after each answer. DKT uses an LSTM for more accurate but less interpretable predictions.
  • How to build a dropout predictor: You saw a full pipeline — synthetic data generation, SMOTE for class imbalance, logistic regression and XGBoost, evaluation with AUC, and interpretation with SHAP.
  • The ethical landmines: Fairness, privacy, and false alarms are not afterthoughts. They must be built into the system from day one.

You can now explore the real UCI dataset, try different models, compute fairness metrics, or implement a simple BKT on your own educational data. The code you’ve seen is a starting point — adapt it to your context.

This is part 10 of our “Machine Learning in the Real World” series. Next up, we’ll explore ML in healthcare — where the stakes are even higher and the ethical considerations even more complex.


Check Your Understanding

  1. Remember: What are the four parameters of Bayesian Knowledge Tracing? Explain each in one sentence.
  2. Understand: Why does BKT not drop belief to zero after a wrong answer?
  3. Apply: Given a student with initial mastery belief of 0.2, a correct answer, and parameters learn_rate=0.3, guess=0.1, slip=0.05, forget=0.0, compute the updated belief using the BKT update function.
  4. Analyze: Compare the trade-offs between BKT and DKT. When would you choose one over the other?
  5. Evaluate: In the dropout prediction example, the AUC for logistic regression was about 0.80. Is that good enough to deploy? What else would you need to check before deployment?
  6. Create: Propose a simple experiment to test whether a dropout prediction model has different false-positive rates for two demographic groups. What data would you need? What metric would you use?

  • Machine Learning in Healthcare: Diagnosis Support, Risk Scoring, and Where It Actually Works (Part 1 of this series) — Healthcare faces similar challenges of fairness and interpretability when predicting patient outcomes.
  • Machine Learning in Marketing: Churn, Segmentation, and Attribution Modeling (Part 6) — Customer churn prediction is conceptually similar to dropout prediction; many of the same techniques (SMOTE, SHAP) apply.

Apply What You Learned is for Supporter and Insider subscribers.

Subscribe to unlock the exercises on this post.

See plans

Looking for something else?

Search every article by title, summary or topic.