Python & Data Science
Statistics Under review

A Self Evaluation Rubric For Code Design Five Dime

You just finished decomposing a problem into functions and classes. Feels good, right? You stared at the blank screen, worked through the nouns and verbs, sketched the data flow, and finally typed out something that runs without errors. But now you’re staring at the output and wondering: is it actually good? How do you tell?

Without a rubric, you’re left with vague feelings. Either you think it’s perfect (and miss the flaws), or you’re lost (and don’t know what to fix next). You wait for someone else to review your code, but in self-study, that someone else is… you. It’s a trap that every self-taught programmer falls into.

Here’s the promise: by the end of this article, you’ll have a five-dimension rubric you can apply in five minutes after every exercise. You’ll know exactly what score you earned, which dimension to practice next, and what “good enough” actually looks like. This is the tool that ties together everything from this series — decomposition, interfaces, state management, error handling, and testability — into one actionable checklist.

Research backs this up. A study on self-assessment in STE(A)M education found that students who used structured rubrics improved their technical skills significantly faster than those who relied on vague self-judgment. The key wasn’t the rubric itself — it was the honest reflection it forced. As the SCORE.org mentoring model puts it, “Self-assessment is most powerful when it’s grounded in clear, observable criteria.” That’s exactly what we’re building.

What This Actually Means: A Rubric Is a Mirror, Not a Report Card

Let’s get the definition out of the way in plain English: a rubric is just a list of things to look for, with clear descriptions of what “bad,” “okay,” and “great” look like. It’s not a pass/fail judgment. It’s a mirror that shows you where you are right now, so you can decide where to go next.

Think of it like a cooking rubric. When you cook a dish, you might judge it on taste, texture, and presentation. A score of 1 on presentation (it looks like a mess) doesn’t mean the dish is bad — it means you should focus on plating next time. A score of 5 on taste means you nailed the seasoning. The rubric helps you see the full picture, not just one aspect.

Here are the five dimensions we’ll use for code design. Each one gets a score from 1 to 5:

Dimension1 (Needs work)3 (Getting there)5 (Solid)
DecompositionOne blob does everythingSome separation, but unclear responsibilitiesClear single-responsibility units
InterfaceWrong or missing typesType hints exist but are loosePrecise types, good defaults, informative errors
State ManagementState where it’s not neededSometimes right, sometimes notClass only when genuinely needed
Error HandlingNo validationBasic checks, uninformative errorsInformative errors at the right level
TestabilityGlobal state or hardcoded dependenciesPartially testableEach unit works in isolation

Target: 4+ on all five. If you’re scoring lower on one dimension, that’s your practice focus for the next exercise. This is the hardest part of self-teaching — knowing what to practice next. The rubric solves that.

Dimension 1: Decomposition — Did You Chop the Problem Into the Right Pieces?

This dimension ties directly back to the noun/verb + data-flow method we’ve been using throughout the series. The question is simple: did you split the problem into pieces that each do one thing?

Score 1: One blob does everything. A single function with 50 lines, or a class that does I/O, logic, and presentation all at once. You can’t reuse any part of it.

Score 3: Some separation exists. You have functions, but they share too much state or have unclear responsibilities. You might have a process_data function that both cleans and analyzes — two things.

Score 5: Clear single-responsibility units. Each function or class does one thing, and you can name that thing in a sentence without using the word “and.”

Here’s a before-and-after example. Imagine a data-cleaning script that reads, cleans, and saves in one function:

# Score 1: One blob does everything
def clean_and_save(path: str) -> None:
    df = pd.read_csv(path)  # Reading
    df = df.dropna()        # Cleaning
    df['age'] = pd.to_numeric(df['age'], errors='coerce')  # More cleaning
    df.to_csv('cleaned.csv', index=False)  # Saving
    print("Done!")  # Logging

This function does four things: read, clean, save, and log. If you want to just clean without saving, you’re stuck. Now the score-5 version:

# Score 5: Clear single-responsibility units
def load_data(path: str) -> pd.DataFrame:
    """Load CSV data from path."""
    return pd.read_csv(path)

def clean_data(df: pd.DataFrame) -> pd.DataFrame:
    """Remove nulls and fix data types."""
    df = df.dropna()
    df['age'] = pd.to_numeric(df['age'], errors='coerce')
    return df

def save_data(df: pd.DataFrame, path: str) -> None:
    """Save DataFrame to CSV."""
    df.to_csv(path, index=False)

def run_pipeline(input_path: str, output_path: str) -> None:
    """Orchestrate the full pipeline."""
    df = load_data(input_path)
    df = clean_data(df)
    save_data(df, output_path)
    print("Pipeline complete.")

Each function has one job. You can import clean_data in a notebook without dragging in the file I/O. The orchestrator is just four lines. If you can’t describe what a function does without using “and,” you’re probably at a 3 or below.

Dimension 2: Interface — Are Your Functions Easy to Use and Hard to Misuse?

This dimension checks your function signatures. Are they clear? Do they help the caller use them correctly?

Score 1: Wrong or missing types. No type hints, arguments are ambiguous (like a **kwargs that could be anything), defaults are missing or harmful (like a mutable default argument).

Score 3: Mostly correct. Type hints exist but are loose — Any or object instead of specific types. Defaults are okay but not always helpful.

Score 5: Precise types, good defaults, informative errors. Type hints are specific (like pd.DataFrame instead of Any), defaults make common cases easy, and errors tell you exactly what went wrong.

Here’s a poorly-typed function and its improved version:

# Score 1: Wrong or missing types
def clean_data(df, strategy):
    """Clean the data."""
    if strategy == 'dropna':
        return df.dropna()
    elif strategy == 'fill':
        return df.fillna(0)
    else:
        return df

What’s wrong? No type hints. What’s strategy supposed to be? A string? A list? The docstring doesn’t help. If someone passes 'drop_na' (with an underscore), they get no error — just the original data back silently. That’s a bug waiting to happen.

# Score 5: Precise types, good defaults, informative errors
from typing import Literal
import pandas as pd

def clean_data(
    df: pd.DataFrame,
    strategy: Literal['dropna', 'fill'] = 'dropna'
) -> pd.DataFrame:
    """Clean missing values from a DataFrame.
    
    Args:
        df: Input DataFrame with potential missing values.
        strategy: How to handle missing values.
            - 'dropna': Remove rows with any missing value.
            - 'fill': Fill missing values with 0.
    
    Returns:
        Cleaned DataFrame.
    
    Raises:
        ValueError: If strategy is not 'dropna' or 'fill'.
    """
    if strategy == 'dropna':
        return df.dropna()
    elif strategy == 'fill':
        return df.fillna(0)
    else:
        raise ValueError(f"Unknown strategy: '{strategy}'. Choose 'dropna' or 'fill'.")

Now the type hints tell you exactly what to pass. The Literal type restricts strategy to two values. The docstring explains each option. And if you pass something wrong, you get a clear error message. If someone else could use your function without reading its code, you’re probably at a 4 or 5.

Dimension 3: State Management — Are You Using State Only When You Genuinely Need It?

This dimension ties back to the STATE question from Part 10: “Does this thing need to remember something across calls?”

Score 1: State where it’s not needed. Using a class when a function would do, or global variables for no reason. Every self. attribute is a question mark.

Score 3: Sometimes right. Some functions are pure (no side effects, no state), but others carry unnecessary state. You might have a class with one method that doesn’t use self at all.

Score 5: Class only when genuinely needed. You can justify every self. attribute. Most logic is in pure functions that take inputs and return outputs.

Here’s an over-engineered class and its function-based refactor:

# Score 1: State where it's not needed
class ConfigProcessor:
    """Process configuration files."""
    
    def __init__(self, path: str):
        self.path = path
        self.config = None
    
    def load(self) -> dict:
        with open(self.path, 'r') as f:
            self.config = json.load(f)
        return self.config
    
    def get_value(self, key: str):
        if self.config is None:
            raise ValueError("Config not loaded yet. Call load() first.")
        return self.config.get(key)

Why is this over-engineered? The class stores path and config as state, but the only reason is to avoid passing them as arguments. The get_value method depends on load() being called first — a hidden dependency. If you want to use just the get_value logic, you have to instantiate the class and remember to call load().

# Score 5: Function-based, no unnecessary state
def load_config(path: str) -> dict:
    """Load JSON config from file."""
    with open(path, 'r') as f:
        return json.load(f)

def get_config_value(config: dict, key: str):
    """Get a value from a config dictionary."""
    return config.get(key)

# Usage
config = load_config('config.json')
value = get_config_value(config, 'database_url')

Now there’s no state. Each function is pure — it takes inputs and returns outputs. You can test get_config_value without touching any file. If you can’t explain why a piece of data needs to be stored as state, it probably shouldn’t be.

Dimension 4: Error Handling — What Happens When Things Go Wrong?

This dimension checks your code’s resilience. What happens when the input is bad, the file doesn’t exist, or the data has unexpected values?

Score 1: No validation. The function crashes with a cryptic error (like KeyError: 'age') or silently returns garbage (like None when you expected a DataFrame).

Score 3: Some checks exist. Basic validation is there, but errors are uninformative or at the wrong level. For example, catching everything with a generic except: that hides the real problem.

Score 5: Informative errors at the right level. Specific exceptions with clear messages. Validation happens early — before any processing. Errors are not caught silently unless intentionally (like logging a warning).

Here’s a function with no error handling and its improved version:

# Score 1: No validation
def divide_numbers(a, b):
    return a / b

If b is 0, you get ZeroDivisionError: division by zero. That’s not helpful. If a or b is a string, you get TypeError: unsupported operand type(s). The user has to guess what went wrong.

# Score 5: Informative errors at the right level
def divide_numbers(a: float, b: float) -> float:
    """Divide a by b.
    
    Args:
        a: Numerator.
        b: Denominator (must be non-zero).
    
    Returns:
        Result of a / b.
    
    Raises:
        ValueError: If b is zero.
        TypeError: If a or b are not numeric.
    """
    if not isinstance(a, (int, float)) or not isinstance(b, (int, float)):
        raise TypeError(f"Both arguments must be numeric. Got {type(a).__name__} and {type(b).__name__}.")
    if b == 0:
        raise ValueError("Denominator cannot be zero. Division by zero is undefined.")
    return a / b

Now the error messages tell you exactly what to fix. If a user gets “Denominator cannot be zero,” they know immediately what’s wrong. No guessing. If a user of your function gets an error message that tells them exactly what to fix, you’re at a 5.

Dimension 5: Testability — Can You Test Each Unit in Isolation?

This dimension checks if your code is actually testable. Can you write a unit test without touching a database, a file system, or the network?

Score 1: Global state or hardcoded dependencies. The function reads from a file or database directly, or uses a global variable that’s hard to control in tests.

Score 3: Partially testable. Some units are testable, but others require complex setup or have hidden dependencies. You might have to create real files or mock too many things.

Score 5: Each unit works in isolation. Functions are pure or accept dependencies as arguments. Mocking is straightforward — you can pass a fake object instead of the real one.

Here’s a hardcoded function and its refactored version:

# Score 1: Hardcoded dependency — reads from file directly
def count_lines() -> int:
    with open('data.txt', 'r') as f:
        return len(f.readlines())

To test this, you need a real file named data.txt on disk. That’s fragile — the test depends on the file system state. If the file doesn’t exist, the test fails for the wrong reason.

# Score 5: Accepts dependency as argument — testable with StringIO
from io import StringIO

def count_lines(file_obj) -> int:
    """Count lines in a file-like object.
    
    Args:
        file_obj: A file-like object (e.g., open file, StringIO).
    
    Returns:
        Number of lines.
    """
    return len(file_obj.readlines())

# In production:
with open('data.txt', 'r') as f:
    line_count = count_lines(f)

# In a test:
def test_count_lines():
    fake_file = StringIO("line1\nline2\nline3\n")
    assert count_lines(fake_file) == 3

Now the function works with any file-like object. In production, you pass a real file. In tests, you pass a StringIO — no disk access needed. If you can write a unit test without touching a database or a file system, you’re at a 5.

Putting It All Together: How to Score Your Next Exercise

Let’s apply the rubric to a real piece of code. Here’s a simple data pipeline — read a CSV, clean it, and save the result:

# Sample code to score
import pandas as pd

def process_data(path):
    df = pd.read_csv(path)
    df = df.dropna()
    df['age'] = pd.to_numeric(df['age'], errors='coerce')
    df.to_csv('output.csv', index=False)
    print("Done")

Let’s score it dimension by dimension:

  • Decomposition = 1: One blob does everything — read, clean, save, and log. You can’t reuse just the cleaning logic.
  • Interface = 1: No type hints. The parameter path could be anything. No docstring. No error messages.
  • State Management = 3: No unnecessary state (it’s a function, not a class), but it’s not pure — it writes to a file as a side effect.
  • Error Handling = 1: No validation. If the file doesn’t exist, you get a cryptic FileNotFoundError. If age column is missing, you get KeyError.
  • Testability = 1: Hardcoded file paths. To test, you need real files on disk.

Your lowest dimensions are Decomposition, Interface, Error Handling, and Testability — focus on those next.

Now here’s the improved version that scores 4+ on all dimensions:

# Improved version — scores 4+ on all dimensions
from typing import Optional
import pandas as pd

def load_data(path: str) -> pd.DataFrame:
    """Load CSV data from a file path.
    
    Args:
        path: Path to the CSV file.
    
    Returns:
        Loaded DataFrame.
    
    Raises:
        FileNotFoundError: If the file does not exist.
    """
    return pd.read_csv(path)

def clean_data(df: pd.DataFrame) -> pd.DataFrame:
    """Remove nulls and fix data types.
    
    Args:
        df: Input DataFrame.
    
    Returns:
        Cleaned DataFrame.
    """
    df = df.dropna()
    if 'age' in df.columns:
        df['age'] = pd.to_numeric(df['age'], errors='coerce')
    return df

def save_data(df: pd.DataFrame, path: str) -> None:
    """Save DataFrame to CSV.
    
    Args:
        df: DataFrame to save.
        path: Output file path.
    """
    df.to_csv(path, index=False)

def run_pipeline(input_path: str, output_path: Optional[str] = None) -> None:
    """Run the full data pipeline.
    
    Args:
        input_path: Path to input CSV.
        output_path: Path to output CSV. Defaults to 'output.csv'.
    """
    if output_path is None:
        output_path = 'output.csv'
    
    df = load_data(input_path)
    df = clean_data(df)
    save_data(df, output_path)
    print(f"Pipeline complete. Output saved to {output_path}")

Now let’s score it:

  • Decomposition = 5: Four functions, each with one job. The orchestrator is just a few lines.
  • Interface = 5: Type hints everywhere. Docstrings explain each parameter. A sensible default for output_path.
  • State Management = 5: No classes, no state. Functions are mostly pure (except save_data which has a side effect, but that’s its job).
  • Error Handling = 4: Basic validation exists (checking for column existence), but could be more explicit about file errors. Still, the errors from pd.read_csv are reasonably informative.
  • Testability = 5: Each function can be tested in isolation. clean_data takes a DataFrame and returns a DataFrame — no file I/O needed.

Target 4+ on all five. If you’re at a 3 on Decomposition, don’t practice Interface — practice Decomposition. The rubric tells you exactly where to focus.

The Catch: Self-Assessment Is Hard — Here’s How to Make It Honest

It’s easy to give yourself a 5 when you’re tired of looking at the code. Self-assessment is inherently biased — we all want to believe our code is good. Here’s how to stay honest.

Strategy 1: Score immediately after finishing. Do it before you forget the pain points. Did you struggle with a particular function? That’s a clue it might be a 3, not a 5.

Strategy 2: Wait a day and re-score. Fresh eyes catch more. You’ll notice things you missed — like a missing type hint or a function that does two things.

Strategy 3: Swap rubrics with a peer. Even if they’re not an expert, they can spot obvious issues. The SCORE.org mentoring model emphasizes that external calibration improves self-assessment accuracy. A mentor or peer can say, “Hey, your error handling is actually a 2, not a 4 — look at this silent exception.”

Strategy 4: Use multiple data points. The CEL 5D+ teacher evaluation rubric recommends combining self-assessment with observer feedback. For self-study, that means scoring your code, then running it through a linter or type checker, then asking a friend to review.

StrategyWhen to useWhy it helps
Score immediatelyRight after finishingCaptures pain points while fresh
Wait and re-scoreNext dayFresh eyes catch missed issues
Peer swapAfter self-scoringExternal calibration reduces bias
Multiple data pointsAlwaysCombines self + tool + peer feedback

The rubric is a tool for growth, not a trophy. A low score is a gift — it tells you exactly what to practice next. If you score a 2 on Error Handling, your next exercise should focus on adding validation and informative error messages. That’s how you improve.

What You’ve Learned and What’s Next

Let’s recap what you now have:

  1. A five-dimension rubric — Decomposition, Interface, State Management, Error Handling, and Testability — with clear 1/3/5 descriptions.
  2. A scoring system — Target 4+ on all five. Focus your next practice session on your lowest dimension.
  3. A self-assessment strategy — Score immediately, wait and re-score, swap with a peer, and use multiple data points.
  4. A concrete example — The before-and-after data pipeline shows exactly what a 1 looks like vs. a 5.

Your action plan: After every decomposition exercise, score yourself on all five dimensions. Write down the scores. If you’re below 4 on any dimension, that’s your focus for the next exercise. Repeat until all five are consistently 4+.

In the next part of this series, we’ll apply this rubric to a real-world messy codebase and refactor it together. You’ll see the rubric in action on code that’s not a toy example — it’s the kind of code you’d find in a production notebook or a legacy script. We’ll score it, identify the weak dimensions, and systematically fix each one.

Explore the earlier parts of this series if you need a refresher on any dimension. Part 10 covers the STATE question for state management. Part 11 shows how to refactor a god class into single-responsibility units. Each part builds on the last, and this rubric is the tool that ties them all together.

Check Your Understanding

Remember: List the five dimensions of the code design rubric.

Understand: In your own words, explain why a score of 1 on Decomposition is problematic, even if the code runs correctly.

Apply: Take a piece of code you wrote recently (or use the process_data function from this article). Score it on all five dimensions using the 1/3/5 scale. Write down your scores and identify your lowest dimension.

Analyze: Compare the before and after versions of the data pipeline in this article. For each dimension, identify the specific changes that moved the score from 1 to 5.

Evaluate: A peer tells you their code scores 5 on all dimensions. Based on what you learned about self-assessment bias, what questions would you ask to verify their scores?

Create: Design a new dimension for the rubric (e.g., “Performance” or “Readability”). Write the 1/3/5 descriptions for it, and explain why it would be useful for code design.

  • Part 10: When Does a Problem Actually Need a Class? The STATE Question, Revisited — This rubric’s State Management dimension directly applies the STATE question from that article.
  • Part 11: Refactoring a God Class Into Four Single-Responsibility Ones — The Decomposition dimension in this rubric is exactly what we practiced in that refactoring exercise.

Apply What You Learned is for Supporter and Insider subscribers.

Subscribe to unlock the exercises on this post.

See plans

Looking for something else?

Search every article by title, summary or topic.