Python & Data Science

When Does A Problem Actually Need A Class The Stat

You’ve just finished translating a formula into clean, working Python. The code runs. The numbers match. You feel great. Then you look at the next task on your board: “Refactor the data pipeline.” And suddenly you’re staring at a blank file, wondering: should this be a class or a function? Should I wrap everything in a class just in case? Or keep it as loose functions?

If you’ve ever felt that freeze, you’re not alone. Even experienced developers fall into what Steve Yegge called “the Kingdom of Nouns” — the belief that everything in code must be a class. In his famous 2006 essay, Yegge pointed out that Java (and by extension, many OOP languages) over-privileges classes. Everything becomes a noun, even when a simple verb would do.

But here’s the thing: Yegge’s critique is spot-on, but it doesn’t tell you what to do instead. It just says “stop making everything a class.” That’s like saying “stop eating junk food” without telling you what to eat instead.

This is where the STATE question comes in. It’s the fix. One simple question that cuts through the noun/verb trap and tells you exactly when to reach for a class and when to keep it simple.

By the end of this tutorial, you’ll know the STATE question cold. You’ll see it in action across several examples. And you’ll be able to apply it to your own code this week.

The STATE Question: One Question That Cuts Through the Noise

Here’s the question in its simplest form:

Does anything need to remember information between calls?

That’s it. If the answer is yes, you probably need a class. If the answer is no, a standalone function is cleaner.

Let’s unpack what “remembering information between calls” actually means. In plain English: a function that remembers something from a previous call will give a different output for the same input, depending on what it saw before. A pure function — one that doesn’t remember anything — always gives the same output for the same input.

Think of it this way: a vending machine is stateless. You put in $$2, press B4, and you get a bag of chips. Every time. But a coffee shop loyalty card is stateful. The barista remembers how many stamps you have. The third stamp gives you a free drink, but the first two don’t. Same action (buying a coffee), different result based on remembered history.

Now here’s the interesting part: this question maps directly to the decomposition strategy you’ve been learning in this series. Remember the five-step algorithm from Part 2? The STATE question is Step 3 — the one that separates nouns from verbs. It’s the decision point that tells you whether a “noun” deserves its own class or should just be a function.

Let’s see what happens when we apply it to a simple example.

# Stateless: pure function, no memory needed
def add(a: float, b: float) -> float:
    """Add two numbers. Same input always gives same output."""
    return a + b

# Stateful: tracks running total between calls
class Accumulator:
    """Tracks a running total. Remembers previous calls."""
    
    def __init__(self):
        # This is the state: a variable that persists between calls
        self.total = 0.0
    
    def add(self, value: float) -> float:
        """Add value to running total and return new total."""
        self.total += value
        return self.total

# Let's see the difference in action
print("Pure function:")
print(f"  add(3, 4) = {add(3, 4)}")  # Always 7
print(f"  add(3, 4) = {add(3, 4)}")  # Still 7
print()

acc = Accumulator()
print("Stateful class:")
print(f"  acc.add(3) = {acc.add(3)}")   # First call: total = 3
print(f"  acc.add(4) = {acc.add(4)}")   # Second call: total = 7 (remembered the 3)
print(f"  acc.add(3) = {acc.add(3)}")   # Third call: total = 10 (remembered the 7)

Output:

Pure function:
  add(3, 4) = 7
  add(3, 4) = 7

Stateful class:
  acc.add(3) = 3
  acc.add(4) = 7
  acc.add(3) = 10

See the difference? The pure function add always returns 7 for inputs (3, 4). But the Accumulator class remembers the previous total. Calling acc.add(3) gives different results depending on what happened before.

This is the hardest part of the class-vs-function decision, but the question makes it simple. If you need memory between calls, use a class. If not, use a function.

Stateless vs. Stateful: A Worked Contrast That Makes It Click

Let’s look at three examples that seem similar on the surface. The STATE question cleanly separates them.

Example 1: Stateless — Z-score normalization

import numpy as np

def z_score_normalize(data: np.ndarray) -> np.ndarray:
    """
    Normalize data to have mean 0 and std 1.
    
    This is a pure function: same input always gives same output.
    No state needed because it doesn't remember anything between calls.
    """
    mean = np.mean(data)
    std = np.std(data)
    return (data - mean) / std

# Test it
sample_data = np.array([1.0, 2.0, 3.0, 4.0, 5.0])
normalized = z_score_normalize(sample_data)
print(f"Original: {sample_data}")
print(f"Normalized: {normalized}")
print(f"Mean of normalized: {np.mean(normalized):.6f}")
print(f"Std of normalized: {np.std(normalized):.6f}")

Output:

Original: [1. 2. 3. 4. 5.]
Normalized: [-1.41421356 -0.70710678  0.          0.70710678  1.41421356]
Mean of normalized: 0.000000
Std of normalized: 1.000000

The STATE question says: does anything need to remember information between calls? No. Each call to z_score_normalize is independent. It takes data, computes mean and std, and returns the result. No memory needed. So a standalone function is perfect.

Example 2: Stateful — TTL Cache

import time
from typing import Any, Optional

class TTLCache:
    """
    A cache that expires entries after a time-to-live (TTL).
    
    Needs state because it must remember:
    - The cached values (self._cache)
    - When each entry was stored (self._timestamps)
    - The TTL duration (self._ttl)
    """
    
    def __init__(self, ttl_seconds: float = 60.0):
        # State: stores the cached key-value pairs
        self._cache: dict[str, Any] = {}
        # State: stores when each key was last set
        self._timestamps: dict[str, float] = {}
        # State: configuration that persists
        self._ttl = ttl_seconds
    
    def set(self, key: str, value: Any) -> None:
        """Store a value with current timestamp."""
        self._cache[key] = value
        self._timestamps[key] = time.time()
    
    def get(self, key: str) -> Optional[Any]:
        """
        Get a value if it exists and hasn't expired.
        Returns None if key doesn't exist or has expired.
        """
        if key not in self._cache:
            return None
        
        # Check if expired
        age = time.time() - self._timestamps[key]
        if age > self._ttl:
            # Expired: remove and return None
            del self._cache[key]
            del self._timestamps[key]
            return None
        
        return self._cache[key]

# Quick test
cache = TTLCache(ttl_seconds=0.1)  # 100ms TTL
cache.set("user_42", {"name": "Alice"})
print(f"Immediately after set: {cache.get('user_42')}")

time.sleep(0.15)  # Wait longer than TTL
print(f"After 150ms: {cache.get('user_42')}")

Output:

Immediately after set: {'name': 'Alice'}
After 150ms: None

The STATE question says: does anything need to remember information between calls? Yes. The cache must remember the stored values and their timestamps across calls to get and set. Without state, a cache is impossible. So a class is the right choice.

Example 3: Stateful — Rate Limiter

import time
from collections import defaultdict

class RateLimiter:
    """
    Limits requests per key within a sliding time window.
    
    Needs state because it must remember:
    - The timestamps of recent requests for each key (self._requests)
    - The window size and max requests (self._window, self._max_requests)
    """
    
    def __init__(self, max_requests: int = 10, window_seconds: float = 60.0):
        # State: maps each key to a list of timestamps of recent requests
        self._requests: dict[str, list[float]] = defaultdict(list)
        # State: configuration that persists
        self._max_requests = max_requests
        self._window = window_seconds
    
    def allow_request(self, key: str) -> bool:
        """
        Check if a request from this key is allowed.
        Returns True if under the limit, False if rate-limited.
        """
        now = time.time()
        
        # Remove timestamps outside the window
        self._requests[key] = [
            ts for ts in self._requests[key]
            if now - ts < self._window
        ]
        
        # Check if under limit
        if len(self._requests[key]) < self._max_requests:
            self._requests[key].append(now)
            return True
        else:
            return False

# Quick test
limiter = RateLimiter(max_requests=3, window_seconds=10.0)

for i in range(5):
    allowed = limiter.allow_request("user_42")
    print(f"Request {i+1}: {'Allowed' if allowed else 'Rate-limited'}")

Output:

Request 1: Allowed
Request 2: Allowed
Request 3: Allowed
Request 4: Rate-limited
Request 5: Rate-limited

The STATE question says: does anything need to remember information between calls? Yes. The rate limiter must remember the timestamps of recent requests for each key. Without that memory, it can’t know if a new request exceeds the limit. A class is the right choice.

Notice something important: all three examples involve time or data. But only two need to remember something between calls. The STATE question correctly separates them.

But Wait — Production Systems Need a More Nuanced STATE Question

The basic STATE question works great for simple scripts. But production systems add complexity. A cache that lives in memory is fine for a single-process app, but what if the process restarts? What if multiple threads access it at once?

This is where we extend the basic question into four sub-questions. Think of them as the “production-ready” version of the STATE question.

Sub-question 3a: In-memory state between calls → use an instance variable

This is the basic case we’ve already seen. If state lives in memory and doesn’t need to survive restarts, use self.something in __init__.

Sub-question 3b: Must survive process restarts → add save/load methods

If the state must persist across restarts, you need to serialize it to disk or a database.

Sub-question 3c: Accessed from multiple threads → add a lock

If multiple threads can access the state simultaneously, you need threading.Lock to prevent race conditions.

Sub-question 3d: Observable for debugging → add a history property

If you need to inspect the state for debugging or monitoring, add a property that exposes it safely.

Let’s see how this plays out with our TTLCache example.

import json
import threading
import time
from typing import Any, Optional

class ProductionTTLCache:
    """
    A TTL cache that's production-ready:
    - Survives restarts (save/load to disk)
    - Thread-safe (lock)
    - Observable (history property)
    """
    
    def __init__(self, ttl_seconds: float = 60.0, filepath: Optional[str] = None):
        # 3a: In-memory state
        self._cache: dict[str, Any] = {}
        self._timestamps: dict[str, float] = {}
        self._ttl = ttl_seconds
        
        # 3c: Thread safety
        self._lock = threading.Lock()
        
        # 3b: Persistence
        self._filepath = filepath
        if filepath:
            self._load_from_disk()
        
        # 3d: Observable history
        self._history: list[dict[str, Any]] = []
    
    def _load_from_disk(self) -> None:
        """Load cache state from disk."""
        try:
            with open(self._filepath, 'r') as f:
                data = json.load(f)
            self._cache = data.get('cache', {})
            self._timestamps = data.get('timestamps', {})
            print(f"Loaded {len(self._cache)} entries from disk")
        except FileNotFoundError:
            print("No existing cache file found, starting fresh")
    
    def _save_to_disk(self) -> None:
        """Save cache state to disk."""
        if not self._filepath:
            return
        data = {
            'cache': self._cache,
            'timestamps': self._timestamps
        }
        with open(self._filepath, 'w') as f:
            json.dump(data, f)
    
    def set(self, key: str, value: Any) -> None:
        """Thread-safe set with persistence."""
        with self._lock:
            self._cache[key] = value
            self._timestamps[key] = time.time()
            self._save_to_disk()
            self._history.append({
                'action': 'set',
                'key': key,
                'timestamp': time.time()
            })
    
    def get(self, key: str) -> Optional[Any]:
        """Thread-safe get with expiry check."""
        with self._lock:
            if key not in self._cache:
                self._history.append({
                    'action': 'get_miss',
                    'key': key,
                    'timestamp': time.time()
                })
                return None
            
            age = time.time() - self._timestamps[key]
            if age > self._ttl:
                del self._cache[key]
                del self._timestamps[key]
                self._save_to_disk()
                self._history.append({
                    'action': 'get_expired',
                    'key': key,
                    'timestamp': time.time()
                })
                return None
            
            self._history.append({
                'action': 'get_hit',
                'key': key,
                'timestamp': time.time()
            })
            return self._cache[key]
    
    @property
    def history(self) -> list[dict[str, Any]]:
        """3d: Observable state for debugging."""
        return list(self._history)
    
    @property
    def cache_size(self) -> int:
        """3d: Observable state for monitoring."""
        with self._lock:
            return len(self._cache)

# Quick test
cache = ProductionTTLCache(ttl_seconds=0.2, filepath="/tmp/test_cache.json")
cache.set("user_42", {"name": "Alice"})
print(f"Cache size: {cache.cache_size}")
print(f"Get immediately: {cache.get('user_42')}")

time.sleep(0.25)
print(f"Get after expiry: {cache.get('user_42')}")
print(f"Cache size after expiry: {cache.cache_size}")
print(f"History entries: {len(cache.history)}")

Output:

No existing cache file found, starting fresh
Cache size: 1
Get immediately: {'name': 'Alice'}
Get after expiry: None
Cache size after expiry: 0
History entries: 3

The basic STATE question still comes first. But these four sub-questions make it production-ready. Start with the basic question, then ask each sub-question to add the features your system needs.

What This Actually Means: Interpreting the Numbers

Let’s walk through a concrete scenario where the STATE question saves you from over-engineering. Imagine you’re writing a script that reads a CSV, normalizes a column, and writes the result.

Here’s the naive approach — the one Yegge warned us about:

import pandas as pd

class DataProcessor:
    """
    Naive class: wraps everything in a class even though no state is needed.
    Notice all the 'self.' references that don't actually use state.
    """
    
    def __init__(self, filepath: str):
        # This looks like state, but it's just configuration passed once
        self.filepath = filepath
        self.data = None  # Will be set later, but only used within one call chain
    
    def read_csv(self) -> pd.DataFrame:
        """Read CSV from filepath."""
        self.data = pd.read_csv(self.filepath)
        return self.data
    
    def normalize_column(self, column: str) -> pd.Series:
        """Normalize a column to z-scores."""
        if self.data is None:
            raise ValueError("No data loaded")
        mean = self.data[column].mean()
        std = self.data[column].std()
        return (self.data[column] - mean) / std
    
    def write_csv(self, output_path: str) -> None:
        """Write processed data to CSV."""
        if self.data is None:
            raise ValueError("No data loaded")
        self.data.to_csv(output_path, index=False)

# Usage
processor = DataProcessor("data.csv")
processor.read_csv()
normalized = processor.normalize_column("value")
processor.data["value_normalized"] = normalized
processor.write_csv("output.csv")
print("Processing complete")

Now apply the STATE question: does anything need to remember information between calls? No. Each step is independent. The read_csv function doesn’t need to remember anything for normalize_column to work — you could pass the data directly. The class adds complexity without benefit.

Here’s the refactored version:

import pandas as pd

def read_csv(filepath: str) -> pd.DataFrame:
    """Read CSV file. Pure function: same input, same output."""
    return pd.read_csv(filepath)

def normalize_column(data: pd.DataFrame, column: str) -> pd.Series:
    """Normalize a column to z-scores. Pure function."""
    mean = data[column].mean()
    std = data[column].std()
    return (data[column] - mean) / std

def write_csv(data: pd.DataFrame, output_path: str) -> None:
    """Write DataFrame to CSV. Pure function (side effect: writes file)."""
    data.to_csv(output_path, index=False)

# Usage
data = read_csv("data.csv")
data["value_normalized"] = normalize_column(data, "value")
write_csv(data, "output.csv")
print("Processing complete")

The refactored version is shorter, easier to test, and more reusable. The class added complexity without any benefit. The STATE question caught it.

The Hardest Part: When State Is Subtle and You Miss It

Now here’s the genuinely difficult edge case: state that isn’t obvious. You look at a function and think “that’s stateless,” but under the hood, it remembers something.

Hidden state: functools.lru_cache

from functools import lru_cache

@lru_cache(maxsize=128)
def expensive_computation(n: int) -> int:
    """
    Looks like a pure function, but has hidden state!
    The cache remembers previous calls and their results.
    """
    print(f"  Computing for n={n}...")
    return n * n

# First call: actually computes
print(f"First call: expensive_computation(5) = {expensive_computation(5)}")

# Second call with same input: uses cache (hidden state)
print(f"Second call: expensive_computation(5) = {expensive_computation(5)}")

# Check the cache info
print(f"Cache info: {expensive_computation.cache_info()}")

Output:

  Computing for n=5...
First call: expensive_computation(5) = 25
Second call: expensive_computation(5) = 25
Cache info: CacheInfo(hits=1, misses=1, maxsize=128, currsize=1)

Notice the second call didn’t print “Computing…” — it used the cached result. The decorator @lru_cache wraps the function in a class that stores previous results. Under the hood, it’s a class with state, even though it looks like a function.

Hidden state: Generator functions

def count_up_to(n: int):
    """
    Generator function: remembers where it left off between calls.
    Each 'yield' pauses execution, and the next call resumes from there.
    """
    i = 0
    while i < n:
        print(f"    Generator: yielding {i}")
        yield i
        i += 1

# Create the generator
gen = count_up_to(3)

# Each call to next() resumes from where it left off
print("First call:")
print(f"  Got: {next(gen)}")

print("Second call:")
print(f"  Got: {next(gen)}")

print("Third call:")
print(f"  Got: {next(gen)}")

Output:

First call:
    Generator: yielding 0
  Got: 0
Second call:
    Generator: yielding 1
  Got: 1
Third call:
    Generator: yielding 2
  Got: 2

The generator remembers where it left off. The variable i persists between calls. That’s state, even though it doesn’t look like a class.

Here’s the reassuring part: these are advanced cases. For 90% of problems, the basic STATE question is enough. But knowing about hidden state helps you spot the edge cases when they appear.

Recap: What You Learned and Where to Go Next

Let’s recap what you learned:

  1. The STATE question is the fix for Yegge’s noun/verb trap. It asks: “Does anything need to remember information between calls?”
  2. Basic form: Yes → class. No → standalone function.
  3. Production systems need four sub-questions: in-memory state (3a), persistence (3b), thread safety (3c), and observability (3d).
  4. Hidden state can appear in decorators like @lru_cache and in generator functions. The STATE question still works — you just need to look for it.

Try applying the STATE question to your own codebase this week. You’ll be surprised how many classes you can simplify into standalone functions.

Explore the next part of the series to see how this fits into testing your decomposed code.

Check Your Understanding

Remember: What does the STATE question ask in its simplest form?

Understand: Explain in plain English why a TTLCache needs a class but a z_score_normalize function does not.

Apply: Given a CountdownTimer that tracks remaining seconds and counts down each time tick() is called, would you implement it as a class or a function? Why?

Analyze: Look at this code snippet. Identify the hidden state and explain why it matters.

from functools import lru_cache

@lru_cache(maxsize=100)
def fibonacci(n):
    if n < 2:
        return n
    return fibonacci(n-1) + fibonacci(n-2)

Evaluate: A teammate argues that every function should be wrapped in a class “just in case we need to add state later.” Using what you learned about the STATE question, evaluate this argument. What’s the cost of over-engineering with classes?

Create: Design a simple ConnectionPool class. Apply the STATE question and all four production sub-questions (3a-3d). What state does it need? How would you make it thread-safe? How would you make it observable?

Apply What You Learned is for Supporter and Insider subscribers.

Subscribe to unlock the exercises on this post.

See plans

Looking for something else?

Search every article by title, summary or topic.