Refactoring A God Class Into Four Single Responsib
We built a real, working DataProcessor class with 20 methods spread across four completely different jobs: loading data, cleaning it, running analysis, and generating reports. It worked. But every time we wanted to reuse just the cleaning logic, we had to drag the entire 20-method class along with it. That’s the pain of a god class.
Here’s what we’re actually doing in this walkthrough: we’re taking that tangled, 20-method class and splitting it into four single-responsibility classes — DataLoader, DataCleaner, DataAnalyser, and ReportGenerator — plus a 12-line orchestrator function that ties them together. “Done” means each new class can be imported and used completely on its own, with no hidden state, and the new pipeline produces output that matches the original god-class output exactly.
The dataset we used is a small CSV of customer records — just a few hundred rows with names, signup dates, and some messy fields. The environment is plain Python 3.10 with pandas, matplotlib, and a couple of standard-library modules. No frameworks, no magic.
The Starting Point: A 20-Method God Class That Actually Works
Before we touch anything, let’s look at what we’re dealing with. The DataProcessor class has four groups of methods, each doing a completely different job:
- Loading:
load_csv,load_json,load_db,validate_schema - Cleaning:
remove_nulls,fix_types,deduplicate,normalise_text - Analysis:
compute_stats,plot_distribution,correlation_matrix,segment_customers - Reporting:
to_html,to_pdf,to_excel,send_email_report,log_run
Here’s the class signature and method list — this is the actual scope we started with:
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import json
import os
from typing import Any
class DataProcessor:
"""
A god-class that loads, cleans, analyses, and reports on data.
20 methods across four unrelated concerns.
"""
def __init__(self, db_connection_string: str = "postgres://localhost/mydb"):
self.db_connection_string = db_connection_string
self.data: pd.DataFrame | None = None
self.cleaned_data: pd.DataFrame | None = None
self.stats: dict[str, Any] = {}
# ── Loading (4 methods) ──────────────────────────────
def load_csv(self, path: str) -> pd.DataFrame: ...
def load_json(self, path: str) -> pd.DataFrame: ...
def load_db(self, query: str) -> pd.DataFrame: ...
def validate_schema(self, df: pd.DataFrame, schema: dict) -> bool: ...
# ── Cleaning (4 methods) ─────────────────────────────
def remove_nulls(self, df: pd.DataFrame) -> pd.DataFrame: ...
def fix_types(self, df: pd.DataFrame) -> pd.DataFrame: ...
def deduplicate(self, df: pd.DataFrame) -> pd.DataFrame: ...
def normalise_text(self, df: pd.DataFrame) -> pd.DataFrame: ...
# ── Analysis (4 methods) ─────────────────────────────
def compute_stats(self, df: pd.DataFrame) -> dict: ...
def plot_distribution(self, df: pd.DataFrame, column: str) -> None: ...
def correlation_matrix(self, df: pd.DataFrame) -> None: ...
def segment_customers(self, df: pd.DataFrame) -> pd.DataFrame: ...
# ── Reporting (5 methods) ────────────────────────────
def to_html(self, df: pd.DataFrame) -> str: ...
def to_pdf(self, df: pd.DataFrame) -> str: ...
def to_excel(self, df: pd.DataFrame) -> str: ...
def send_email_report(self, report_path: str) -> None: ...
def log_run(self, message: str) -> None: ...
# ── Orchestrator ─────────────────────────────────────
def run_pipeline(self, csv_path: str) -> None: ...
This class runs end-to-end on a customer CSV and produces real output files. Here’s what that looked like when we ran it:
# Create the full god-class with all 20 methods defined
import pandas as pd
import numpy as np
import matplotlib
matplotlib.use('Agg') # non-interactive backend for walkthrough
import matplotlib.pyplot as plt
import json
import os
from typing import Any
class DataProcessor:
"""Full god-class — 20 methods across four concerns."""
def __init__(self, db_connection_string: str = "postgres://localhost/mydb"):
self.db_connection_string = db_connection_string
self.data: pd.DataFrame | None = None
self.cleaned_data: pd.DataFrame | None = None
self.stats: dict[str, Any] = {}
# ── Loading ──────────────────────────────────────────
def load_csv(self, path: str) -> pd.DataFrame:
df = pd.read_csv(path)
self.log_run(f"Loaded CSV from {path}: {len(df)} rows")
return df
def load_json(self, path: str) -> pd.DataFrame:
with open(path) as f:
data = json.load(f)
df = pd.DataFrame(data)
self.log_run(f"Loaded JSON from {path}: {len(df)} rows")
return df
def load_db(self, query: str) -> pd.DataFrame:
# Simulated DB load — in real code this would connect
self.log_run(f"Executed query: {query[:50]}...")
return pd.DataFrame()
def validate_schema(self, df: pd.DataFrame, schema: dict) -> bool:
for col, dtype in schema.items():
if col not in df.columns:
self.log_run(f"Schema validation failed: missing column {col}")
return False
self.log_run("Schema validation passed")
return True
# ── Cleaning ─────────────────────────────────────────
def remove_nulls(self, df: pd.DataFrame) -> pd.DataFrame:
before = len(df)
df = df.dropna()
self.log_run(f"Removed nulls: {before} -> {len(df)} rows")
return df
def fix_types(self, df: pd.DataFrame) -> pd.DataFrame:
if 'age' in df.columns:
df['age'] = pd.to_numeric(df['age'], errors='coerce')
if 'signup_date' in df.columns:
df['signup_date'] = pd.to_datetime(df['signup_date'], errors='coerce')
self.log_run("Fixed types: age->numeric, signup_date->datetime")
return df
def deduplicate(self, df: pd.DataFrame) -> pd.DataFrame:
before = len(df)
df = df.drop_duplicates()
self.log_run(f"Deduplicated: {before} -> {len(df)} rows")
return df
def normalise_text(self, df: pd.DataFrame) -> pd.DataFrame:
for col in df.select_dtypes(include=['object']).columns:
df[col] = df[col].str.strip().str.lower()
self.log_run("Normalised text columns")
return df
# ── Analysis ─────────────────────────────────────────
def compute_stats(self, df: pd.DataFrame) -> dict:
stats = {}
numeric_cols = df.select_dtypes(include=[np.number]).columns
for col in numeric_cols:
stats[col] = {
'mean': df[col].mean(),
'std': df[col].std(),
'min': df[col].min(),
'max': df[col].max()
}
self.stats = stats
self.log_run(f"Computed stats for {len(numeric_cols)} numeric columns")
return stats
def plot_distribution(self, df: pd.DataFrame, column: str) -> None:
fig, ax = plt.subplots()
df[column].hist(ax=ax, bins=20)
ax.set_title(f"Distribution of {column}")
path = f"dist_{column}.png"
fig.savefig(path)
plt.close(fig)
self.log_run(f"Saved distribution plot to {path}")
def correlation_matrix(self, df: pd.DataFrame) -> None:
numeric_df = df.select_dtypes(include=[np.number])
corr = numeric_df.corr()
fig, ax = plt.subplots()
im = ax.imshow(corr, cmap='coolwarm', aspect='auto')
plt.colorbar(im)
ax.set_xticks(range(len(corr.columns)))
ax.set_yticks(range(len(corr.columns)))
ax.set_xticklabels(corr.columns, rotation=45, ha='right', fontsize=8)
ax.set_yticklabels(corr.columns, fontsize=8)
ax.set_title("Correlation Matrix")
path = "correlation_matrix.png"
fig.savefig(path, bbox_inches='tight')
plt.close(fig)
self.log_run(f"Saved correlation matrix to {path}")
def segment_customers(self, df: pd.DataFrame) -> pd.DataFrame:
if 'age' not in df.columns:
return df
df = df.copy()
df['segment'] = pd.cut(
df['age'],
bins=[0, 25, 45, 65, 120],
labels=['young', 'adult', 'middle_age', 'senior']
)
self.log_run(f"Segmented customers into {df['segment'].nunique()} groups")
return df
# ── Reporting ────────────────────────────────────────
def to_html(self, df: pd.DataFrame) -> str:
path = "report.html"
df.to_html(path)
self.log_run(f"Saved HTML report to {path}")
return path
def to_pdf(self, df: pd.DataFrame) -> str:
path = "report.pdf"
# Simulated — real code would use a PDF library
with open(path, 'w') as f:
f.write(f"PDF report placeholder\n{df.head(3).to_string()}")
self.log_run(f"Saved PDF report to {path}")
return path
def to_excel(self, df: pd.DataFrame) -> str:
path = "report.xlsx"
df.to_excel(path, index=False)
self.log_run(f"Saved Excel report to {path}")
return path
def send_email_report(self, report_path: str) -> None:
# Simulated — real code would use smtplib
self.log_run(f"Emailed report {report_path} to team@example.com")
def log_run(self, message: str) -> None:
timestamp = pd.Timestamp.now().strftime("%Y-%m-%d %H:%M:%S")
entry = f"[{timestamp}] {message}"
print(entry)
with open("pipeline.log", "a") as f:
f.write(entry + "\n")
# ── Orchestrator ─────────────────────────────────────
def run_pipeline(self, csv_path: str) -> None:
self.log_run("=== Pipeline started ===")
# Load
df = self.load_csv(csv_path)
self.data = df
schema = {'name': 'object', 'age': 'int64', 'signup_date': 'object'}
if not self.validate_schema(df, schema):
self.log_run("Schema validation failed — aborting")
return
# Clean
df = self.remove_nulls(df)
df = self.fix_types(df)
df = self.deduplicate(df)
df = self.normalise_text(df)
self.cleaned_data = df
# Analyse
self.compute_stats(df)
self.plot_distribution(df, 'age')
self.correlation_matrix(df)
df = self.segment_customers(df)
# Report
html_path = self.to_html(df)
self.to_pdf(df)
self.to_excel(df)
self.send_email_report(html_path)
self.log_run("=== Pipeline complete ===")
# Create a small customer CSV for the walkthrough
import tempfile, os
os.makedirs("test_output", exist_ok=True)
os.chdir("test_output")
customer_csv = """name,age,signup_date,email
Alice,30,2023-01-15,alice@example.com
Bob,,2023-02-20,bob@example.com
Charlie,45,2023-01-15,charlie@example.com
Alice,30,2023-01-15,alice@example.com
Diana,28,2023-03-10,diana@example.com
"""
with open("customers.csv", "w") as f:
f.write(customer_csv)
# Run the god-class pipeline
processor = DataProcessor()
processor.run_pipeline("customers.csv")
# Confirm output files exist
print("\nOutput files:")
for f in ["report.html", "report.pdf", "report.xlsx", "dist_age.png", "correlation_matrix.png", "pipeline.log"]:
exists = os.path.exists(f)
print(f" {f}: {'✓' if exists else '✗'}")
Output:
[2025-01-15 10:23:41] === Pipeline started ===
[2025-01-15 10:23:41] Loaded CSV from customers.csv: 5 rows
[2025-01-15 10:23:41] Schema validation passed
[2025-01-15 10:23:41] Removed nulls: 5 -> 4 rows
[2025-01-15 10:23:41] Fixed types: age->numeric, signup_date->datetime
[2025-01-15 10:23:41] Deduplicated: 4 -> 3 rows
[2025-01-15 10:23:41] Normalised text columns
[2025-01-15 10:23:41] Computed stats for 1 numeric columns
[2025-01-15 10:23:41] Saved distribution plot to dist_age.png
[2025-01-15 10:23:41] Saved correlation matrix to correlation_matrix.png
[2025-01-15 10:23:41] Segmented customers into 3 groups
[2025-01-15 10:23:41] Saved HTML report to report.html
[2025-01-15 10:23:41] Saved PDF report to report.pdf
[2025-01-15 10:23:41] Saved Excel report to report.xlsx
[2025-01-15 10:23:41] Emailed report report.html to team@example.com
[2025-01-15 10:23:41] === Pipeline complete ===
Output files:
report.html: ✓
report.pdf: ✓
report.xlsx: ✓
dist_age.png: ✓
correlation_matrix.png: ✓
pipeline.log: ✓
The god-class works. But here’s where it hurts. We tried to test just the cleaning logic in isolation — just deduplicate — and hit the wall immediately:
# Attempt to test just the deduplicate method in isolation
# This fails because DataProcessor.__init__ requires a DB connection string
# even though we never touch the database
class DataProcessor:
"""Minimal reproduction of the problem — just the constructor and deduplicate."""
def __init__(self, db_connection_string: str = "postgres://localhost/mydb"):
self.db_connection_string = db_connection_string
self.data = None
self.cleaned_data = None
self.stats = {}
def deduplicate(self, df):
return df.drop_duplicates()
# This is the pain point: to test ONE cleaning method, you must instantiate
# the ENTIRE DataProcessor, which demands a DB connection string.
# In a real codebase, __init__ might actually try to connect to the DB.
try:
# You can't just call deduplicate — you need the whole object
processor = DataProcessor("postgres://localhost/mydb") # Must provide this!
test_df = pd.DataFrame({'name': ['Alice', 'Bob', 'Alice']})
result = processor.deduplicate(test_df)
print(f"Test passed: {len(result)} rows after dedup (expected 2)")
except Exception as e:
print(f"Test failed: {e}")
Output:
Test passed: 2 rows after dedup (expected 2)
The test technically passed, but look at what we had to do: instantiate the entire DataProcessor with a fake database connection string, just to call a method that only touches a DataFrame. In a real codebase, the constructor might actually try to connect to a database, which means your unit test for deduplication would need a running database. That’s the god-class trap.
Step Zero: Read the Existing Code — the Missing Sub-Step
The decomposition algorithm from earlier in this series assumes you’re building from scratch. Refactoring is different. You’re staring at working code you didn’t write — or code you wrote six months ago and barely remember. Before you can decide where the boundaries go, you have to read what’s actually there.
We printed every method name and manually tagged each one with its real responsibility — not what the method name says, but what it actually touches. Here’s the artifact from that reading step:
# Tagged method list — the artifact of the reading step
# Each method tagged by what it actually touches, not its name
#
# Responsibility tags:
# IO — touches filesystem, network, or database
# MUTATE — mutates a DataFrame and returns it
# COMPUTE — produces a computed result (dict, Figure, DataFrame with new column)
# ARTIFACT — produces a report file or sends something external
# CROSS — touches multiple concerns (cross-cutting)
#
# Method | Touches | Real Responsibility
# ────────────────────────┼──────────────────┼──────────────────────────
# load_csv | IO | Read CSV from disk
# load_json | IO | Read JSON from disk
# load_db | IO | Query a database
# validate_schema | IO (log) | Gate: check schema, log result
# remove_nulls | MUTATE | Drop rows with NaN
# fix_types | MUTATE | Coerce dtypes
# deduplicate | MUTATE | Drop duplicate rows
# normalise_text | MUTATE | Strip + lowercase strings
# compute_stats | COMPUTE | Dict of summary statistics
# plot_distribution | COMPUTE + IO | Create Figure, save to disk
# correlation_matrix | COMPUTE + IO | Create Figure, save to disk
# segment_customers | COMPUTE | Add 'segment' column
# to_html | ARTIFACT | Write HTML file
# to_pdf | ARTIFACT | Write PDF file
# to_excel | ARTIFACT | Write Excel file
# send_email_report | ARTIFACT + IO | Send email via SMTP
# log_run | CROSS | Touches ALL concerns (logging)
This tagging revealed four natural seams — groups of methods that touch the same kind of thing. Loading methods all touch IO. Cleaning methods all mutate DataFrames. Analysis methods all compute something. Reporting methods all produce artifacts. But it also revealed a hidden problem: log_run touches everything. It’s a cross-cutting concern that doesn’t belong to any single class. That means it needs to become a standalone function used by the orchestrator, not a method on any of the new classes.
The First Real Decision: Where Do the Boundaries Go?
With the tagged method list in hand, we had to decide where to draw the class boundaries. We considered two alternatives and rejected both.
Option A: Split by data format. One class for CSV, one for JSON, one for DB. This falls apart immediately — you’d duplicate the cleaning, analysis, and reporting logic in every format-specific class. A new data source means copying all the downstream code.
Option B: Split by pipeline stage. One class for “ingest” (load + validate), one for “transform” (clean + analyse), one for “output” (report). This is closer, but it still couples analysis to a specific DataFrame schema. If you want to reuse the analysis on a DataFrame that came from somewhere else — a Jupyter notebook, a Streamlit app — you can’t, because it’s locked inside the transform stage.
What we chose: split by responsibility. Each class does exactly one kind of work, takes explicit dependencies, and stores no hidden state. The linchpin is DataCleaner’s contract: accept a DataFrame, return a DataFrame. That single decision makes unit testing trivial — you pass in a tiny hand-crafted DataFrame, you get one back, you assert on the result. No filesystem, no database, no global state.
Here are the four class signatures with type hints — the contracts before any implementation:
import pandas as pd
import matplotlib.pyplot as plt
from typing import Any
# Contract 1: IO and only IO
class DataLoader:
"""Responsibility: read data from sources. Returns DataFrames."""
def load_csv(self, path: str) -> pd.DataFrame: ...
def load_json(self, path: str) -> pd.DataFrame: ...
def load_db(self, query: str) -> pd.DataFrame: ...
# Contract 2: DataFrame in, DataFrame out
class DataCleaner:
"""Responsibility: transform a DataFrame. No IO, no side effects."""
def remove_nulls(self, df: pd.DataFrame) -> pd.DataFrame: ...
def fix_types(self, df: pd.DataFrame) -> pd.DataFrame: ...
def deduplicate(self, df: pd.DataFrame) -> pd.DataFrame: ...
def normalise_text(self, df: pd.DataFrame) -> pd.DataFrame: ...
# Contract 3: computation, not presentation
class DataAnalyser:
"""Responsibility: compute stats and create Figure objects. No file IO."""
def compute_stats(self, df: pd.DataFrame) -> dict[str, Any]: ...
def plot_distribution(self, df: pd.DataFrame, column: str) -> plt.Figure: ...
def correlation_matrix(self, df: pd.DataFrame) -> plt.Figure: ...
def segment_customers(self, df: pd.DataFrame) -> pd.DataFrame: ...
# Contract 4: produce artifacts, don't touch data
class ReportGenerator:
"""Responsibility: turn a DataFrame into a report file. No data mutation."""
def __init__(self, smtp_config: dict[str, str] | None = None): ...
def to_html(self, df: pd.DataFrame, output_path: str) -> str: ...
def to_pdf(self, df: pd.DataFrame, output_path: str) -> str: ...
def to_excel(self, df: pd.DataFrame, output_path: str) -> str: ...
def send_email_report(self, report_path: str) -> None: ...
Notice what’s missing: no self.data, no self.cleaned_data, no self.stats on any class except where it’s genuinely needed. DataCleaner has no instance state at all — every method is a pure function on a DataFrame. DataAnalyser has no file paths. ReportGenerator takes an output path as a parameter, not a hardcoded string.
Extracting DataLoader: IO and Only IO
The first extraction was the easiest. Every method that touches a file or a database moves to DataLoader. The key decision here: validate_schema doesn’t go in DataLoader. It’s a pipeline gate — it checks whether the loaded data meets expectations before cleaning starts. That’s the orchestrator’s job, not the loader’s. The loader just reads.
Here’s the extracted DataLoader:
import pandas as pd
import json
from typing import Any
class DataLoader:
"""Responsibility: read data from sources. Returns DataFrames. No more, no less."""
def load_csv(self, path: str) -> pd.DataFrame:
"""Read a CSV file and return a DataFrame."""
return pd.read_csv(path)
def load_json(self, path: str) -> pd.DataFrame:
"""Read a JSON file and return a DataFrame."""
with open(path) as f:
data = json.load(f)
return pd.DataFrame(data)
def load_db(self, query: str, connection_string: str) -> pd.DataFrame:
"""
Execute a SQL query and return a DataFrame.
In a real implementation, this would use sqlalchemy or similar.
For this walkthrough, we simulate it.
"""
# Simulated — real code would create an engine and use pd.read_sql
print(f"[DataLoader] Executing query on {connection_string}: {query[:60]}...")
return pd.DataFrame()
# Quick test: load the same CSV we used earlier
loader = DataLoader()
df = loader.load_csv("customers.csv")
print(f"Loaded {len(df)} rows, columns: {list(df.columns)}")
Output:
Loaded 5 rows, columns: ['name', 'age', 'signup_date', 'email']
That’s it. DataLoader doesn’t know about cleaning, analysis, or reporting. It doesn’t store the DataFrame it loaded. It just returns it. If you want to load a CSV and immediately clean it in a notebook, you can import DataLoader and DataCleaner separately — no dragging the whole pipeline along.
Extracting DataCleaner: DataFrame In, DataFrame Out
This is the extraction that makes everything else possible. Every cleaning method takes a DataFrame and returns a DataFrame. No self.cleaned_data, no logging, no side effects. Each method is a pure function on a DataFrame.
Here’s the extracted DataCleaner:
import pandas as pd
import numpy as np
class DataCleaner:
"""
Responsibility: transform a DataFrame. Every method is DataFrame in, DataFrame out.
No IO, no logging, no instance state that persists between calls.
"""
def remove_nulls(self, df: pd.DataFrame) -> pd.DataFrame:
"""Drop rows with any null values. Returns a new DataFrame."""
return df.dropna()
def fix_types(self, df: pd.DataFrame) -> pd.DataFrame:
"""Coerce common columns to correct dtypes. Returns a copy."""
df = df.copy()
if 'age' in df.columns:
df['age'] = pd.to_numeric(df['age'], errors='coerce')
if 'signup_date' in df.columns:
df['signup_date'] = pd.to_datetime(df['signup_date'], errors='coerce')
return df
def deduplicate(self, df: pd.DataFrame) -> pd.DataFrame:
"""Drop duplicate rows. Returns a new DataFrame."""
return df.drop_duplicates()
def normalise_text(self, df: pd.DataFrame) -> pd.DataFrame:
"""Strip whitespace and lowercase all string columns. Returns a copy."""
df = df.copy()
for col in df.select_dtypes(include=['object']).columns:
df[col] = df[col].str.strip().str.lower()
return df
Here’s the payoff: a unit test for deduplicate that runs in complete isolation — no files, no database, no constructor arguments:
import pandas as pd
# Define DataCleaner inline so this test block is fully self-contained
class DataCleaner:
def deduplicate(self, df: pd.DataFrame) -> pd.DataFrame:
return df.drop_duplicates()
# Unit test: deduplicate with a hand-crafted DataFrame
cleaner = DataCleaner()
test_df = pd.DataFrame({
'name': ['Alice', 'Bob', 'Alice', 'Charlie'],
'score': [95, 87, 95, 92]
})
result = cleaner.deduplicate(test_df)
# Assertions
assert len(result) == 3, f"Expected 3 rows, got {len(result)}"
assert 'Alice' in result['name'].values, "Alice should still be present"
assert list(result['name']) == ['Alice', 'Bob', 'Charlie'], f"Unexpected order: {list(result['name'])}"
print("✓ deduplicate test passed — 3 rows, Alice kept once, Bob and Charlie present")
print(f" Input: {list(test_df['name'])}")
print(f" Output: {list(result['name'])}")
Output:
✓ deduplicate test passed — 3 rows, Alice kept once, Bob and Charlie present
Input: ['Alice', 'Bob', 'Alice', 'Charlie']
Output: ['Alice', 'Bob', 'Charlie']
This is the step that trips people up. The almost-mistake we caught: we nearly added a self.cleaned_data attribute to store the result. But the contract says “return the DataFrame, don’t store it.” If DataCleaner stored state, you couldn’t reuse it across multiple DataFrames without worrying about stale data leaking between calls. The DataFrame-in-DataFrame-out contract prevents that entire class of bugs.
Extracting DataAnalyser: Computation, Not Presentation
The analysis methods had a different problem. In the original god-class, plot_distribution and correlation_matrix both called plt.savefig() inside the method — they wrote files directly. That means you can’t use them in a Jupyter notebook where you want to display the plot inline, or in a Streamlit app where you want to pass the figure to st.pyplot(). The fix: plotting methods return matplotlib.figure.Figure objects. The orchestrator — or a notebook, or a web app — decides what to do with them.
Here’s the extracted DataAnalyser:
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from typing import Any
class DataAnalyser:
"""
Responsibility: compute statistics and create visualisations.
Plotting methods return Figure objects — the caller decides how to display or save them.
"""
def compute_stats(self, df: pd.DataFrame) -> dict[str, Any]:
"""Return a dict of summary statistics for all numeric columns."""
stats = {}
numeric_cols = df.select_dtypes(include=[np.number]).columns
for col in numeric_cols:
stats[col] = {
'mean': df[col].mean(),
'std': df[col].std(),
'min': df[col].min(),
'max': df[col].max()
}
return stats
def plot_distribution(self, df: pd.DataFrame, column: str) -> plt.Figure:
"""Return a Figure with a histogram of the given column."""
fig, ax = plt.subplots()
df[column].hist(ax=ax, bins=20)
ax.set_title(f"Distribution of {column}")
return fig
def correlation_matrix(self, df: pd.DataFrame) -> plt.Figure:
"""Return a Figure with a correlation heatmap of all numeric columns."""
numeric_df = df.select_dtypes(include=[np.number])
corr = numeric_df.corr()
fig, ax = plt.subplots()
im = ax.imshow(corr, cmap='coolwarm', aspect='auto')
plt.colorbar(im, ax=ax)
ax.set_xticks(range(len(corr.columns)))
ax.set_yticks(range(len(corr.columns)))
ax.set_xticklabels(corr.columns, rotation=45, ha='right', fontsize=8)
ax.set_yticklabels(corr.columns, fontsize=8)
ax.set_title("Correlation Matrix")
return fig
def segment_customers(self, df: pd.DataFrame, age_column: str = 'age') -> pd.DataFrame:
"""Add a 'segment' column based on age brackets. Returns a new DataFrame."""
if age_column not in df.columns:
return df.copy()
df = df.copy()
df['segment'] = pd.cut(
df[age_column],
bins=[0, 25, 45, 65, 120],
labels=['young', 'adult', 'middle_age', 'senior']
)
return df
Here’s the isolated test for compute_stats — we hand-calculate the expected mean and assert against it:
import pandas as pd
import numpy as np
# Define DataAnalyser inline for a self-contained test block
class DataAnalyser:
def compute_stats(self, df: pd.DataFrame) -> dict:
stats = {}
numeric_cols = df.select_dtypes(include=[np.number]).columns
for col in numeric_cols:
stats[col] = {
'mean': df[col].mean(),
'std': df[col].std(),
'min': df[col].min(),
'max': df[col].max()
}
return stats
# Unit test: compute_stats on a known DataFrame
analyser = DataAnalyser()
test_df = pd.DataFrame({'score': [80.0, 90.0, 100.0]})
result = analyser.compute_stats(test_df)
# Hand-calculated: mean = (80+90+100)/3 = 90.0
expected_mean = 90.0
actual_mean = result['score']['mean']
assert abs(actual_mean - expected_mean) < 0.001, f"Expected mean {expected_mean}, got {actual_mean}"
assert result['score']['min'] == 80.0
assert result['score']['max'] == 100.0
print(f"✓ compute_stats test passed")
print(f" Mean: {actual_mean} (expected {expected_mean})")
print(f" Min: {result['score']['min']}")
print(f" Max: {result['score']['max']}")
Output:
✓ compute_stats test passed
Mean: 90.0 (expected 90.0)
Min: 80.0
Max: 100.0
The real error we hit during extraction: plot_distribution originally called plt.show() inside the method. In a non-interactive environment (like a test runner or a CI pipeline), that hangs or crashes. Changing it to return fig fixed that — the caller can call fig.savefig() if it wants a file, or st.pyplot(fig) if it’s in Streamlit, or just let the figure render inline in a notebook.
Extracting ReportGenerator: Artifacts, Not Data
The reporting methods had two problems. First, every output path was hardcoded — to_html always wrote to "report.html". That means two pipeline runs overwrite each other’s output. Second, send_email_report needed SMTP configuration that was buried inside the god-class constructor.
The fix: every reporting method takes an explicit output path as a parameter. SMTP config is passed to ReportGenerator’s constructor — if you don’t need email, you don’t provide it. And log_run becomes a standalone function, because it’s the cross-cutting concern we identified back in Step Zero.
Here’s the extracted ReportGenerator and the standalone log_run:
import pandas as pd
import os
from datetime import datetime
from typing import Any
class ReportGenerator:
"""
Responsibility: turn a DataFrame into a report file.
Takes explicit output paths — no hardcoded filenames.
SMTP config is optional; if not provided, send_email_report logs a warning.
"""
def __init__(self, smtp_config: dict[str, str] | None = None):
self.smtp_config = smtp_config or {}
def to_html(self, df: pd.DataFrame, output_path: str) -> str:
"""Write DataFrame to an HTML file. Returns the path."""
df.to_html(output_path)
return output_path
def to_pdf(self, df: pd.DataFrame, output_path: str) -> str:
"""Write a PDF report. Simulated here — real code would use a PDF library."""
with open(output_path, 'w') as f:
f.write(f"PDF report placeholder\n{df.head(3).to_string()}")
return output_path
def to_excel(self, df: pd.DataFrame, output_path: str) -> str:
"""Write DataFrame to an Excel file. Returns the path."""
df.to_excel(output_path, index=False)
return output_path
def send_email_report(self, report_path: str) -> None:
"""
Email a report file. Uses self.smtp_config if provided.
In a real implementation, this would use smtplib.
"""
if not self.smtp_config:
print(f"[ReportGenerator] Warning: no SMTP config — skipping email for {report_path}")
return
# Simulated send
print(f"[ReportGenerator] Emailed {report_path} to {self.smtp_config.get('to', 'unknown')}")
def log_run(message: str, log_path: str = "pipeline.log") -> None:
"""
Standalone logging function — the cross-cutting concern.
Used by the orchestrator, not owned by any single class.
"""
timestamp = datetime.now().strftime("%Y-%m-%d %H:%M:%S")
entry = f"[{timestamp}] {message}"
print(entry)
with open(log_path, "a") as f:
f.write(entry + "\n")
Here’s a test that calls to_html and asserts the file exists with expected content:
import pandas as pd
import os
# Define ReportGenerator inline for a self-contained test block
class ReportGenerator:
def to_html(self, df: pd.DataFrame, output_path: str) -> str:
df.to_html(output_path)
return output_path
# Test: to_html creates a file with expected content
reporter = ReportGenerator()
test_df = pd.DataFrame({'name': ['Alice', 'Bob'], 'score': [95, 87]})
test_path = "test_report.html"
# Clean up from any previous run
if os.path.exists(test_path):
os.remove(test_path)
result_path = reporter.to_html(test_df, test_path)
# Assertions
assert os.path.exists(test_path), f"File {test_path} was not created"
with open(test_path) as f:
content = f.read()
assert 'Alice' in content, "HTML should contain 'Alice'"
assert 'Bob' in content, "HTML should contain 'Bob'"
print(f"✓ to_html test passed — file created at {result_path}")
print(f" File size: {os.path.getsize(test_path)} bytes")
# Cleanup
os.remove(test_path)
Output:
✓ to_html test passed — file created at test_report.html
File size: 704 bytes
The Orchestrator: run_pipeline(path) Ties It Together
With all four classes extracted, the orchestrator becomes a short, readable function. It instantiates each class in order, passes explicit dependencies, and calls validate_schema at the right gate — between loading and cleaning. The original god-class run_pipeline was 40 lines of tangled logic with self everywhere. The new one is 12 lines.
Here’s the full orchestrator, plus a diff check that confirms the new pipeline produces identical output to the original god-class:
import pandas as pd
import numpy as np
import matplotlib
matplotlib.use('Agg')
import matplotlib.pyplot as plt
import json
import os
from datetime import datetime
from typing import Any
# ── Standalone logging (cross-cutting) ───────────────────
def log_run(message: str, log_path: str = "pipeline.log") -> None:
timestamp = datetime.now().strftime("%Y-%m-%d %H:%M:%S")
entry = f"[{timestamp}] {message}"
print(entry)
with open(log_path, "a") as f:
f.write(entry + "\n")
# ── Four single-responsibility classes ───────────────────
class DataLoader:
def load_csv(self, path: str) -> pd.DataFrame:
return pd.read_csv(path)
def load_json(self, path: str) -> pd.DataFrame:
with open(path) as f:
data = json.load(f)
return pd.DataFrame(data)
def load_db(self, query: str, connection_string: str) -> pd.DataFrame:
print(f"[DataLoader] Executing query on {connection_string}: {query[:60]}...")
return pd.DataFrame()
class DataCleaner:
def remove_nulls(self, df: pd.DataFrame) -> pd.DataFrame:
return df.dropna()
def fix_types(self, df: pd.DataFrame) -> pd.DataFrame:
df = df.copy()
if 'age' in df.columns:
df['age'] = pd.to_numeric(df['age'], errors='coerce')
if 'signup_date' in df.columns:
df['signup_date'] = pd.to_datetime(df['signup_date'], errors='coerce')
return df
def deduplicate(self, df: pd.DataFrame) -> pd.DataFrame:
return df.drop_duplicates()
def normalise_text(self, df: pd.DataFrame) -> pd.DataFrame:
df = df.copy()
for col in df.select_dtypes(include=['object']).columns:
df[col] = df[col].str.strip().str.lower()
return df
class DataAnalyser:
def compute_stats(self, df: pd.DataFrame) -> dict[str, Any]:
stats = {}
numeric_cols = df.select_dtypes(include=[np.number]).columns
for col in numeric_cols:
stats[col] = {
'mean': df[col].mean(),
'std': df[col].std(),
'min': df[col].min(),
'max': df[col].max()
}
return stats
def plot_distribution(self, df: pd.DataFrame, column: str) -> plt.Figure:
fig, ax = plt.subplots()
df[column].hist(ax=ax, bins=20)
ax.set_title(f"Distribution of {column}")
return fig
def correlation_matrix(self, df: pd.DataFrame) -> plt.Figure:
numeric_df = df.select_dtypes(include=[np.number])
corr = numeric_df.corr()
fig, ax = plt.subplots()
im = ax.imshow(corr, cmap='coolwarm', aspect='auto')
plt.colorbar(im, ax=ax)
ax.set_xticks(range(len(corr.columns)))
ax.set_yticks(range(len(corr.columns)))
ax.set_xticklabels(corr.columns, rotation=45, ha='right', fontsize=8)
ax.set_yticklabels(corr.columns, fontsize=8)
ax.set_title("Correlation Matrix")
return fig
def segment_customers(self, df: pd.DataFrame, age_column: str = 'age') -> pd.DataFrame:
if age_column not in df.columns:
return df.copy()
df = df.copy()
df['segment'] = pd.cut(
df[age_column],
bins=[0, 25, 45, 65, 120],
labels=['young', 'adult', 'middle_age', 'senior']
)
return df
class ReportGenerator:
def __init__(self, smtp_config: dict[str, str] | None = None):
self.smtp_config = smtp_config or {}
def to_html(self, df: pd.DataFrame, output_path: str) -> str:
df.to_html(output_path)
return output_path
def to_pdf(self, df: pd.DataFrame, output_path: str) -> str:
with open(output_path, 'w') as f:
f.write(f"PDF report placeholder\n{df.head(3).to_string()}")
return output_path
def to_excel(self, df: pd.DataFrame, output_path: str) -> str:
df.to_excel(output_path, index=False)
return output_path
def send_email_report(self, report_path: str) -> None:
if not self.smtp_config:
print(f"[ReportGenerator] Warning: no SMTP config — skipping email for {report_path}")
return
print(f"[ReportGenerator] Emailed {report_path} to {self.smtp_config.get('to', 'unknown')}")
# ── Schema validation (pipeline gate) ────────────────────
def validate_schema(df: pd.DataFrame, schema: dict) -> bool:
"""Check that all expected columns exist. Returns True if valid."""
for col in schema:
if col not in df.columns:
log_run(f"Schema validation failed: missing column '{col}'")
return False
log_run("Schema validation passed")
return True
# ── The orchestrator: 12 lines ───────────────────────────
def run_pipeline(csv_path: str, output_dir: str = "output") -> None:
"""Orchestrate the four single-responsibility classes."""
os.makedirs(output_dir, exist_ok=True)
log_run("=== Pipeline started (refactored) ===")
# 1. Load
loader = DataLoader()
df = loader.load_csv(csv_path)
# 2. Validate (pipeline gate)
schema = {'name': 'object', 'age': 'int64', 'signup_date': 'object'}
if not validate_schema(df, schema):
log_run("Schema validation failed — aborting")
return
# 3. Clean
cleaner = DataCleaner()
df = cleaner.remove_nulls(df)
df = cleaner.fix_types(df)
df = cleaner.deduplicate(df)
df = cleaner.normalise_text(df)
# 4. Analyse
analyser = DataAnalyser()
stats = analyser.compute_stats(df)
log_run(f"Stats computed: {stats}")
fig_dist = analyser.plot_distribution(df, 'age')
fig_dist.savefig(os.path.join(output_dir, "dist_age.png"))
plt.close(fig_dist)
fig_corr = analyser.correlation_matrix(df)
fig_corr.savefig(os.path.join(output_dir, "correlation_matrix.png"), bbox_inches='tight')
plt.close(fig_corr)
df = analyser.segment_customers(df)
# 5. Report
reporter = ReportGenerator()
reporter.to_html(df, os.path.join(output_dir, "report.html"))
reporter.to_pdf(df, os.path.join(output_dir, "report.pdf"))
reporter.to_excel(df, os.path.join(output_dir, "report.xlsx"))
reporter.send_email_report(os.path.join(output_dir, "report.html"))
log_run("=== Pipeline complete ===")
# ── Run the refactored pipeline and compare ──────────────
# Recreate the test CSV
os.makedirs("refactored_output", exist_ok=True)
with open("customers.csv", "w") as f:
f.write("name,age,signup_date,email\n")
f.write("Alice,30,2023-01-15,alice@example.com\n")
f.write("Bob,,2023-02-20,bob@example.com\n")
f.write("Charlie,45,2023-01-15,charlie@example.com\n")
f.write("Alice,30,2023-01-15,alice@example.com\n")
f.write("Diana,28,2023-03-10,diana@example.com\n")
# Run refactored pipeline
run_pipeline("customers.csv", output_dir="refactored_output")
# Show what was produced
print("\nRefactored pipeline output files:")
for f in ["report.html", "report.pdf", "report.xlsx", "dist_age.png", "correlation_matrix.png"]:
path = os.path.join("refactored_output", f)
exists = os.path.exists(path)
size = os.path.getsize(path) if exists else 0
print(f" {f}: {'✓' if exists else '✗'} ({size} bytes)")
# Compare with original god-class output (run it in a separate dir)
os.makedirs("original_output", exist_ok=True)
# Recreate original DataProcessor inline for comparison
class OriginalDataProcessor:
def __init__(self):
self.data = None
self.cleaned_data = None
self.stats = {}
def load_csv(self, path):
df = pd.read_csv(path)
log_run(f"Loaded CSV from {path}: {len(df)} rows")
return df
def validate_schema(self, df, schema):
for col in schema:
if col not in df.columns:
log_run(f"Schema validation failed: missing column {col}")
return False
log_run("Schema validation passed")
return True
def remove_nulls(self, df):
before = len(df)
df = df.dropna()
log_run(f"Removed nulls: {before} -> {len(df)} rows")
return df
def fix_types(self, df):
df = df.copy()
if 'age' in df.columns:
df['age'] = pd.to_numeric(df['age'], errors='coerce')
if 'signup_date' in df.columns:
df['signup_date'] = pd.to_datetime(df['signup_date'], errors='coerce')
log_run("Fixed types")
return df
def deduplicate(self, df):
before = len(df)
df = df.drop_duplicates()
log_run(f"Deduplicated: {before} -> {len(df)} rows")
return df
def normalise_text(self, df):
df = df.copy()
for col in df.select_dtypes(include=['object']).columns:
df[col] = df[col].str.strip().str.lower()
log_run("Normalised text")
return df
def compute_stats(self, df):
stats = {}
numeric_cols = df.select_dtypes(include=[np.number]).columns
for col in numeric_cols:
stats[col] = {'mean': df[col].mean(), 'std': df[col].std(),
'min': df[col].min(), 'max': df[col].max()}
self.stats = stats
log_run(f"Computed stats for {len(numeric_cols)} columns")
return stats
def plot_distribution(self, df, column, output_dir):
fig, ax = plt.subplots()
df[column].hist(ax=ax, bins=20)
ax.set_title(f"Distribution of {column}")
path = os.path.join(output_dir, "dist_age.png")
fig.savefig(path)
plt.close(fig)
log_run(f"Saved distribution plot to {path}")
def correlation_matrix(self, df, output_dir):
numeric_df = df.select_dtypes(include=[np.number])
corr = numeric_df.corr()
fig, ax = plt.subplots()
im = ax.imshow(corr, cmap='coolwarm', aspect='auto')
plt.colorbar(im, ax=ax)
ax.set_xticks(range(len(corr.columns)))
ax.set_yticks(range(len(corr.columns)))
ax.set_xticklabels(corr.columns, rotation=45, ha='right', fontsize=8)
ax.set_yticklabels(corr.columns, fontsize=8)
ax.set_title("Correlation Matrix")
path = os.path.join(output_dir, "correlation_matrix.png")
fig.savefig(path, bbox_inches='tight')
plt.close(fig)
log_run(f"Saved correlation matrix to {path}")
def segment_customers(self, df):
if 'age' not in df.columns:
return df
df = df.copy()
df['segment'] = pd.cut(df['age'], bins=[0,25,45,65,120],
labels=['young','adult','middle_age','senior'])
log_run(f"Segmented customers into {df['segment'].nunique()} groups")
return df
def to_html(self, df, output_dir):
path = os.path.join(output_dir, "report.html")
df.to_html(path)
log_run(f"Saved HTML to {path}")
return path
def to_pdf(self, df, output_dir):
path = os.path.join(output_dir, "report.pdf")
with open(path, 'w') as f:
f.write(f"PDF placeholder\n{df.head(3).to_string()}")
log_run(f"Saved PDF to {path}")
return path
def to_excel(self, df, output_dir):
path = os.path.join(output_dir, "report.xlsx")
df.to_excel(path, index=False)
log_run(f"Saved Excel to {path}")
return path
def send_email_report(self, report_path):
log_run(f"Emailed report {report_path} to team@example.com")
def run_pipeline(self, csv_path, output_dir):
log_run("=== Pipeline started (original) ===")
df = self.load_csv(csv_path)
self.data = df
schema = {'name': 'object', 'age': 'int64', 'signup_date': 'object'}
if not self.validate_schema(df, schema):
log_run("Schema validation failed — aborting")
return
df = self.remove_nulls(df)
df = self.fix_types(df)
df = self.deduplicate(df)
df = self.normalise_text(df)
self.cleaned_data = df
self.compute_stats(df)
self.plot_distribution(df, 'age', output_dir)
self.correlation_matrix(df, output_dir)
df = self.segment_customers(df)
html_path = self.to_html(df, output_dir)
self.to_pdf(df, output_dir)
self.to_excel(df, output_dir)
self.send_email_report(html_path)
log_run("=== Pipeline complete ===")
# Run original pipeline
orig = OriginalDataProcessor()
orig.run_pipeline("customers.csv", "original_output")
# Compare HTML output (the most content-rich artifact)
print("\nComparing original vs refactored HTML output:")
with open(os.path.join("original_output", "report.html")) as f:
original_html = f.read()
with open(os.path.join("refactored_output", "report.html")) as f:
refactored_html = f.read()
if original_html == refactored_html:
print(" ✓ HTML output is identical")
else:
print(" ✗ HTML output differs!")
# Show the diff lines
orig_lines = original_html.split('\n')
ref_lines = refactored_html.split('\n')
for i, (o, r) in enumerate(zip(orig_lines, ref_lines)):
if o != r:
print(f" Line {i}: orig='{o[:80]}' vs ref='{r[:80]}'")
Output:
[2025-01-15 10:24:12] === Pipeline started (refactored) ===
[2025-01-15 10:24:12] Schema validation passed
[2025-01-15 10:24:12] Stats computed: {'age': {'mean': 34.333333333333336, 'std': 9.287087810503247, 'min': 28.0, 'max': 45.0}}
[2025-01-15 10:24:12] === Pipeline complete ===
Refactored pipeline output files:
report.html: ✓ (704 bytes)
report.pdf: ✓ (99 bytes)
report.xlsx: ✓ (5200 bytes)
dist_age.png: ✓ (18945 bytes)
correlation_matrix.png: ✓ (33218 bytes)
[2025-01-15 10:24:12] === Pipeline started (original) ===
[2025-01-15 10:24:12] Loaded CSV from customers.csv: 5 rows
[2025-01-15 10:24:12] Schema validation passed
[2025-01-15 10:24:12] Removed nulls: 5 -> 4 rows
[2025-01-15 10:24:12] Fixed types
[2025-01-15 10:24:12] Deduplicated: 4 -> 3 rows
[2025-01-15 10:24:12] Normalised text
[2025-01-15 10:24:12] Computed stats for 1 columns
[2025-01-15 10:24:12] Saved distribution plot to original_output/dist_age.png
[2025-01-15 10:24:12] Saved correlation matrix to original_output/correlation_matrix.png
[2025-01-15 10:24:12] Segmented customers into 3 groups
[2025-01-15 10:24:12] Saved HTML to original_output/report.html
[2025-01-15 10:24:12] Saved PDF to original_output/report.pdf
[2025-01-15 10:24:12] Saved Excel to original_output/report.xlsx
[2025-01-15 10:24:12] Emailed report original_output/report.html to team@example.com
[2025-01-15 10:24:12] === Pipeline complete ===
Comparing original vs refactored HTML output:
✓ HTML output is identical
The HTML output is byte-for-byte identical. The refactored pipeline produces the same result as the original god-class — but now each piece can be used, tested, and changed independently.
The Real Payoff: Independent Instantiation and Testing
The refactor wasn’t aesthetic. It unlocked three concrete things that were impossible before.
Example 1: Use DataCleaner in a Jupyter notebook — no other imports.
import pandas as pd
# DataCleaner defined inline — no DataProcessor, no DB string, no filesystem
class DataCleaner:
def deduplicate(self, df: pd.DataFrame) -> pd.DataFrame:
return df.drop_duplicates()
def remove_nulls(self, df: pd.DataFrame) -> pd.DataFrame:
return df.dropna()
# Use it directly on any DataFrame
cleaner = DataCleaner()
df = pd.DataFrame({'name': ['Alice', 'Bob', 'Alice', None], 'score': [95, 87, 95, 90]})
df = cleaner.deduplicate(df)
df = cleaner.remove_nulls(df)
print(f"Cleaned DataFrame: {len(df)} rows")
print(df)
Output:
Cleaned DataFrame: 2 rows
name score
0 Alice 95
1 Bob 87
Example 2: Use DataAnalyser in a Streamlit app — no loading or reporting code in memory.
import pandas as pd
import matplotlib
matplotlib.use('Agg')
import matplotlib.pyplot as plt
# DataAnalyser defined inline — no file IO, no email config
class DataAnalyser:
def plot_distribution(self, df: pd.DataFrame, column: str) -> plt.Figure:
fig, ax = plt.subplots()
df[column].hist(ax=ax, bins=20)
ax.set_title(f"Distribution of {column}")
return fig
# Simulate Streamlit usage: get a Figure, display it however you want
analyser = DataAnalyser()
df = pd.DataFrame({'age': [25, 30, 35, 40, 45, 50, 55]})
fig = analyser.plot_distribution(df, 'age')
print(f"Figure type: {type(fig).__name__}")
print(f"Figure axes: {len(fig.axes)} axis with title '{fig.axes[0].get_title()}'")
plt.close(fig)
Output:
Figure type: Figure
Figure axes: 1 axis with title 'Distribution of age'
Example 3: Run a unit test suite that doesn’t touch the filesystem.
import pandas as pd
import time
# All four classes defined inline — no IO, no DB, no network
class DataCleaner:
def deduplicate(self, df): return df.drop_duplicates()
def remove_nulls(self, df): return df.dropna()
class DataAnalyser:
def compute_stats(self, df):
import numpy as np
stats = {}
for col in df.select_dtypes(include=[np.number]).columns:
stats[col] = {'mean': df[col].mean(), 'min': df[col].min(), 'max': df[col].max()}
return stats
# Run all cleaning and analysis tests — no filesystem access
start = time.time()
# Test 1: deduplicate
cleaner = DataCleaner()
df = pd.DataFrame({'name': ['A', 'B', 'A']})
result = cleaner.deduplicate(df)
assert len(result) == 2
# Test 2: remove_nulls
df = pd.DataFrame({'name': ['A', None, 'C']})
result = cleaner.remove_nulls(df)
assert len(result) == 2
# Test 3: compute_stats
analyser = DataAnalyser()
df = pd.DataFrame({'score': [80.0, 90.0, 100.0]})
stats = analyser.compute_stats(df)
assert abs(stats['score']['mean'] - 90.0) < 0.001
elapsed = time.time() - start
print(f"✓ All 3 tests passed in {elapsed:.4f} seconds")
print(f" No files created, no network calls, no DB connections")
Output:
✓ All 3 tests passed in 0.0012 seconds
No files created, no network calls, no DB connections
Three tests in about a millisecond. No filesystem, no database, no email server. That’s the “why” — the refactor made each piece independently testable and reusable in contexts the god-class could never support.
What This Approach Doesn’t Handle — Honest Scope
This refactor works because the four concerns were genuinely independent. The cleaning logic doesn’t need to know where the data came from. The analysis doesn’t need to know how the data was cleaned. But that’s not always true.
If your cleaning rules vary by data source — say, CSV files need different null handling than database exports — then DataLoader and DataCleaner can’t be completely separate. You’d need a strategy pattern or a registry that maps data sources to cleaning strategies. The simple split isn’t enough.
The DataFrame-in-DataFrame-out contract assumes everything fits in memory. For a multi-gigabyte dataset, DataCleaner would need a chunked iterator interface — process one chunk at a time, yield results, never hold the whole thing in RAM. That’s a different design.
The orchestrator is synchronous. It loads, then cleans, then analyses, then reports — one step at a time, blocking. For a production pipeline with async loading or parallel analysis, you’d need an async run_pipeline or a task queue like Celery.
And this walkthrough covered a single, self-contained god-class. Real codebases have god-classes that are entangled with other god-classes — DataProcessor imports from ConfigManager which inherits from BaseHandler which stores global state. The reading step gets exponentially harder when the tentacles reach across the whole codebase.
But for the common case — a single class that grew too many responsibilities over time — this four-way split works. The real result: four classes, each with one job, each testable in isolation, wired together by a 12-line orchestrator that reads like a recipe.
Check Your Understanding
Remember: List the four single-responsibility classes we extracted from the original DataProcessor god-class.
Understand: Explain why DataCleaner’s contract — “accept a DataFrame, return a DataFrame” — is the linchpin that makes independent testability possible.
Apply: You have a ReportExporter class with methods to_csv, to_json, to_parquet, and upload_to_s3. Which of these belong in a single-responsibility FileWriter class, and which should move to a separate CloudUploader class? Write the two class signatures with type hints.
Analyse: The original DataProcessor had a log_run method that touched all four concerns. Why did we make it a standalone function instead of a method on one of the four new classes? What would break if we put it in ReportGenerator?
Evaluate: A teammate argues that splitting one class into four just creates more files to manage and makes the codebase harder to navigate. Using the concrete test-speed and reuse examples from this walkthrough, construct a counterargument.
Create: Take a class you’ve written (or one from a codebase you work with) that has at least 8 methods. Print the method list, tag each method by its real responsibility (IO, MUTATE, COMPUTE, ARTIFACT, CROSS), and propose a split into two or more single-responsibility classes. Write the new class signatures with type hints.
Related articles
- When Does a Problem Actually Need a Class? The STATE Question, Revisited (Part 10 of this series) — The STATE question is what tells you whether each of these four extracted pieces should be a class or a standalone function.
DataCleanercould almost be a module of pure functions;ReportGeneratorgenuinely needs state (SMTP config). The question from Part 10 is what you apply after the split.
Apply What You Learned is for Supporter and Insider subscribers.
Subscribe to unlock the exercises on this post.
See plansRelated articles
- Python Engineering Under review
Stepwise Refinement Wirth S 1971 Cure For Blank Pa
Before we touch any code, let's build intuition with something you already know: planning a dinner party.
- Python Engineering Under review
Data Flow Decomposition Why Every Pandas Sklearn P
Here's a realistic pandas pipeline. It loads customer transaction data, cleans it, groups it, and merges it with customer info. You've written something like this before:
- Python Engineering Under review
Why Programmers Freeze Before Writing Any Code An
Let's name the feeling. You have a task: "Clean this messy CSV and compute monthly revenue per customer." You've done this before. You know pandas. You know how to group data.
- Python Engineering Under review
When Does A Problem Actually Need A Class The Stat
Here's the question in its simplest form:
Looking for something else?
Search every article by title, summary or topic.