Evaluating An Llm Agent Why Normal Metrics Don T W
Have you ever watched your LLM agent write a perfect email — proper greeting, clear message, professional tone — and then send it to the wrong person? I have. It hurts. The agent from Part 6 of this series produced flawless text, but it hallucinated a tool parameter and called the wrong API. The email looked great. The task failed completely.
Here’s the kicker: if you scored that email with BLEU or ROUGE, it would get a high mark. The words match a reference answer. The grammar is correct. But the agent missed its goal entirely. That’s the core problem with normal NLP metrics for agents: they measure text quality, not task success.
Think about it this way. BLEU compares your output to a reference string. ROUGE checks for overlapping n-grams. Perplexity measures how “surprised” the model is by the text. None of them know whether the agent actually did what it was supposed to do. They’re like grading a pilot on how nicely they speak on the radio, ignoring that they landed the plane in the wrong city.
Agent failures are different from text-generation failures. They happen in tool use (wrong parameters, missing calls), state management (losing context across turns), or multi-step reasoning (getting stuck in loops, hitting dead ends). As the MachineLearningMastery blog puts it, “traditional language model evaluation metrics like BLEU scores or perplexity miss what matters for agents.”
The Braintrust blog describes the unpredictable error modes: “hallucination, repetitive loops, cascading failures.” And Anthropic’s research notes that “the capabilities that make agents useful also make them harder to evaluate… mistakes can propagate and compound.”
The fundamental shift is from “did the answer look right?” to “did the job get done?” That’s what this article is about. We’ll replace one-score-fits-all with a multi-dimensional evaluation pipeline. By the end, you’ll have a framework and code to evaluate any LLM agent — including the one you built in earlier parts of this series.
Let’s start by seeing the problem in action.
# A minimal example: correct text, wrong action
# This block is self-contained and runnable
import numpy as np
from nltk.translate.bleu_score import sentence_bleu
from nltk.translate.bleu_score import SmoothingFunction
# We'll create a simple BLEU scorer manually to avoid NLTK download issues
def simple_bleu(reference, candidate):
"""A simplified BLEU score for demonstration."""
ref_tokens = reference.split()
cand_tokens = candidate.split()
# Count matching n-grams (1-gram only for simplicity)
ref_counts = {}
for token in ref_tokens:
ref_counts[token] = ref_counts.get(token, 0) + 1
cand_counts = {}
for token in cand_tokens:
cand_counts[token] = cand_counts.get(token, 0) + 1
matches = 0
for token, count in cand_counts.items():
matches += min(count, ref_counts.get(token, 0))
precision = matches / len(cand_tokens) if len(cand_tokens) > 0 else 0
# Brevity penalty
if len(cand_tokens) < len(ref_tokens):
bp = np.exp(1 - len(ref_tokens) / len(cand_tokens))
else:
bp = 1.0
return bp * precision
# Scenario: Agent is asked to place an order
# The agent writes a perfect confirmation text
agent_text = "Your order has been placed successfully. Thank you for your purchase!"
reference_text = "Your order has been placed successfully. Thank you for your purchase!"
# BLEU score between agent output and reference
bleu_score = simple_bleu(reference_text, agent_text)
print(f"BLEU score: {bleu_score:.2f}")
print("Interpretation: The text matches perfectly — BLEU says it's a great response.")
print()
# But what actually happened?
# The agent called the wrong API: cancel_order instead of place_order
actual_action = "cancel_order"
target_action = "place_order"
print(f"Agent called: {actual_action}")
print(f"Target action: {target_action}")
print(f"Action match: {actual_action == target_action}")
print()
print("=== VERDICT ===")
print("BLEU score: 1.00 (perfect)")
print("Task completion: FAILURE")
print("The agent wrote great text but did the wrong thing. Normal metrics missed this completely.")
When you run this, you’ll see a perfect BLEU score paired with a complete task failure. The text looks right, but the action is wrong. That’s the problem we’re solving.
What Actually Matters for Agents: Task Completion and Tool Use
If normal metrics are useless, what should we measure instead? The answer is simple: did the agent accomplish the user’s goal?
Let’s define the three most practical dimensions for agent evaluation. These are the numbers that actually tell you if your agent is working.
Task completion rate is the big one. For each end-to-end task, you assign a binary score: success or failure. Sometimes you give partial credit (e.g., the agent booked the flight but got the date wrong). This is the single most informative number you can track. If your agent completes 70% of tasks, you know roughly where you stand.
Tool call accuracy measures whether the agent called the right tool at the right time. Think of it as precision and recall over tool invocations. Did the agent call search_flights when it should have? Did it call book_flight too early? This catches the kind of error from our opening example.
Parameter correctness goes deeper. For each tool call, were the arguments valid and appropriate? Did the agent pass the correct email address? The right date format? The correct passenger count? Parameter errors are surprisingly common — and they’re invisible to text-based metrics.
Here’s what the research says about why this matters. The WebArena paper reports that GPT-4’s end-to-end task success rate is only 14.41% on realistic web tasks, compared to 78.24% for humans. The GAIA paper finds GPT-4-with-plugins succeeds on just 15% of tasks, while humans hit 92%. These aren’t small gaps — they’re chasms.
In plain English: your agent might write great prose but fail 85% of the time at actually booking the flight. That’s the number you need to track, not how pretty the confirmation email looks.
Let’s build a simple evaluation function that captures these dimensions.
# Evaluating an agent trajectory: task completion, tool accuracy, parameter correctness
# This block is self-contained and runnable
def evaluate_trajectory(trajectory, ground_truth):
"""
Evaluate an agent's trajectory against ground truth.
Args:
trajectory: dict with 'tool_calls' (list of dicts) and 'final_output' (str)
ground_truth: dict with 'expected_tool_calls' (list of dicts) and 'expected_output' (str)
Returns:
dict with task_success, tool_accuracy, param_correctness
"""
# 1. Task completion: does the final output match the expected outcome?
# This is a simplified check — in practice you'd use semantic similarity or a judge
task_success = trajectory['final_output'] == ground_truth['expected_output']
# 2. Tool call accuracy: did the agent call the right tools?
# We check if the set of tool names matches
agent_tools = set(call['name'] for call in trajectory['tool_calls'])
expected_tools = set(call['name'] for call in ground_truth['expected_tool_calls'])
# Precision: of the tools the agent called, how many were correct?
if len(agent_tools) == 0:
tool_precision = 0.0 if len(expected_tools) > 0 else 1.0
else:
correct_tools = agent_tools & expected_tools
tool_precision = len(correct_tools) / len(agent_tools)
# Recall: of the expected tools, how many did the agent call?
if len(expected_tools) == 0:
tool_recall = 1.0 if len(agent_tools) == 0 else 0.0
else:
correct_tools = agent_tools & expected_tools
tool_recall = len(correct_tools) / len(expected_tools)
# F1 score for tool accuracy
if tool_precision + tool_recall == 0:
tool_accuracy = 0.0
else:
tool_accuracy = 2 * (tool_precision * tool_recall) / (tool_precision + tool_recall)
# 3. Parameter correctness: for each tool call, check if parameters match
# We'll do a simplified check: compare parameter values for matching tool names
param_correct = 0
param_total = 0
for agent_call in trajectory['tool_calls']:
# Find matching expected call
for expected_call in ground_truth['expected_tool_calls']:
if agent_call['name'] == expected_call['name']:
# Compare parameters
for key in expected_call['params']:
param_total += 1
if key in agent_call['params'] and agent_call['params'][key] == expected_call['params'][key]:
param_correct += 1
break
param_correctness = param_correct / param_total if param_total > 0 else 1.0
return {
'task_success': task_success,
'tool_accuracy': round(tool_accuracy, 2),
'param_correctness': round(param_correctness, 2)
}
# Example 1: A passing trajectory
passing_trajectory = {
'tool_calls': [
{'name': 'search_flights', 'params': {'origin': 'NYC', 'destination': 'London', 'date': '2024-06-15'}},
{'name': 'book_flight', 'params': {'flight_id': 'BA178', 'passengers': 2}}
],
'final_output': 'Booking confirmed for 2 passengers on BA178 from NYC to London on June 15.'
}
ground_truth = {
'expected_tool_calls': [
{'name': 'search_flights', 'params': {'origin': 'NYC', 'destination': 'London', 'date': '2024-06-15'}},
{'name': 'book_flight', 'params': {'flight_id': 'BA178', 'passengers': 2}}
],
'expected_output': 'Booking confirmed for 2 passengers on BA178 from NYC to London on June 15.'
}
print("=== PASSING TRAJECTORY ===")
result = evaluate_trajectory(passing_trajectory, ground_truth)
print(f"Task success: {result['task_success']}")
print(f"Tool accuracy: {result['tool_accuracy']}")
print(f"Parameter correctness: {result['param_correctness']}")
print("Interpretation: The agent did everything right. All scores are perfect.")
print()
# Example 2: A failing trajectory (correct text, wrong tool)
failing_trajectory = {
'tool_calls': [
{'name': 'cancel_order', 'params': {'order_id': 'ORD-123'}}
],
'final_output': 'Your order has been placed successfully. Thank you for your purchase!'
}
print("=== FAILING TRAJECTORY ===")
result = evaluate_trajectory(failing_trajectory, ground_truth)
print(f"Task success: {result['task_success']}")
print(f"Tool accuracy: {result['tool_accuracy']}")
print(f"Parameter correctness: {result['param_correctness']}")
print("Interpretation: The agent wrote a perfect confirmation but called the wrong tool.")
print("Task success is False, and tool accuracy is 0 because no expected tools were called.")
Notice how the failing trajectory gets a perfect score on text quality but fails on every agent-specific metric. That’s the gap we’re closing.
The Reliability Problem: Why One Score Isn’t Enough
Here’s a scenario that will make you pull your hair out. You run your agent once and it succeeds. You run it again with the exact same input and it fails. Which score do you report?
This non-determinism is a defining feature of LLM agents. It comes from sampling temperature, model stochasticity, and environment randomness. A single trial is meaningless. You need metrics that capture consistency.
Let’s talk about two metrics that solve this: pass@k and pass^k.
pass@k is straightforward: you run the agent k times on the same task, and count how many succeed. If it succeeds 3 out of 5 times, you report 60%. This gives you a sense of how often the agent works on average.
pass^k is stricter. The agent must succeed on all k runs to count as a pass. If it fails even once, that task is a failure. This measures reliability for high-stakes tasks. If you’re deploying an agent that handles financial transactions, you need pass^5 to be high — you can’t have it failing 1 out of 5 times.
The τ-bench paper reports sobering numbers. GPT-4o succeeds on less than 50% of retail tasks on a single run. And pass^8 is below 25%. That means even the best model is unreliable on repeated attempts.
In plain English: if your agent only works 2 out of 5 times, you can’t trust it in production. pass^8 tells you how often it works every single time.
Let’s implement this.
# Computing pass@k and pass^k for agent reliability
# This block is self-contained and runnable
import random
def run_agent_once(task):
"""
Simulate running an agent on a task.
Returns True if the agent succeeds, False otherwise.
For this demo, we'll simulate a 60% success rate.
"""
return random.random() < 0.6
def compute_pass_metrics(tasks, k=5):
"""
Run each task k times and compute pass@k and pass^k.
Args:
tasks: list of task identifiers (strings)
k: number of trials per task
Returns:
dict with per-task results and aggregate metrics
"""
results = {}
for task in tasks:
# Run the agent k times
trial_results = []
for trial in range(k):
success = run_agent_once(task)
trial_results.append(success)
# pass@k: at least one success
pass_at_k = any(trial_results)
# pass^k: all trials must succeed
pass_all_k = all(trial_results)
results[task] = {
'trial_results': trial_results,
'pass@k': pass_at_k,
'pass^k': pass_all_k,
'success_count': sum(trial_results)
}
# Aggregate metrics
total_tasks = len(tasks)
pass_at_k_count = sum(1 for r in results.values() if r['pass@k'])
pass_all_k_count = sum(1 for r in results.values() if r['pass^k'])
return {
'per_task': results,
'pass@k_rate': pass_at_k_count / total_tasks,
'pass^k_rate': pass_all_k_count / total_tasks,
'k': k
}
# Set random seed for reproducibility
random.seed(42)
# Define tasks
tasks = ["Book flight NYC-London", "Cancel hotel reservation", "Order pizza", "Schedule meeting", "Send email"]
# Compute metrics with k=5
metrics = compute_pass_metrics(tasks, k=5)
print("=== PER-TASK RESULTS ===")
for task, result in metrics['per_task'].items():
successes = result['success_count']
trials = metrics['k']
print(f"{task}:")
print(f" Trials: {result['trial_results']}")
print(f" Successes: {successes}/{trials}")
print(f" pass@k: {result['pass@k']} (succeeded on at least one trial)")
print(f" pass^k: {result['pass^k']} (succeeded on ALL trials)")
print()
print("=== AGGREGATE METRICS ===")
print(f"pass@5 rate: {metrics['pass@k_rate']:.0%}")
print(f" Interpretation: The agent succeeded on at least one of five tries for {metrics['pass@k_rate']:.0%} of tasks.")
print(f"pass^5 rate: {metrics['pass^k_rate']:.0%}")
print(f" Interpretation: The agent succeeded on ALL five tries for only {metrics['pass^k_rate']:.0%} of tasks.")
print()
print("Key insight: pass@k can be high while pass^k is low.")
print("This means the agent works sometimes but is unreliable — a big problem for production.")
When you run this, you’ll likely see pass@k much higher than pass^k. That’s the reliability gap. Your agent works often enough to look good on average, but not consistently enough to trust.
Beyond the Final Answer: Trajectory-Level Evaluation
Your agent booked the flight — great! But it called the search API 17 times, got stuck in a loop for 3 turns, and finally succeeded by accident. Is that a “good” agent?
Outcome-only metrics say yes. The booking went through. But you’d never deploy an agent that wastes API calls and risks infinite loops. We need to evaluate the path, not just the destination.
This is the difference between outcome supervision and process supervision.
Outcome supervision grades only the final answer. Did the booking succeed? Yes or no. Simple, but it misses inefficient or risky behavior.
Process supervision grades each intermediate step. Did the agent call the right tool with the right parameters at step 2? Did it get stuck in a loop? This catches reasoning errors even when the final answer is correct.
OpenAI’s “Let’s Verify Step by Step” paper showed that a process-supervised model solved 78% of a challenging MATH subset, compared to an outcome-supervised model. Process supervision “significantly outperforms” outcome supervision for complex tasks.
For agents, trajectory evaluation is even more important. The agentevals library (from LangChain) focuses specifically on “the intermediate steps an agent takes as it runs.” This is a first-class concern.
Anthropic recommends a pragmatic middle ground: “reading transcripts, not just final grades” and “grading outcomes rather than the exact path taken.” You don’t need to check every token — but you should check that the agent didn’t do anything dangerous or wasteful.
In plain English: you want an agent that not only gets the right answer, but gets it efficiently and safely. Trajectory evaluation checks the how, not just the what.
Let’s see how this works in practice.
# Trajectory evaluation: comparing outcome-only vs. process supervision
# This block is self-contained and runnable
def evaluate_outcome_only(trajectory, ground_truth):
"""
Outcome supervision: grade only the final answer.
Args:
trajectory: dict with 'final_output' (str) and 'tool_calls' (list)
ground_truth: dict with 'expected_output' (str)
Returns:
float: 1.0 if final output matches, 0.0 otherwise
"""
return 1.0 if trajectory['final_output'] == ground_truth['expected_output'] else 0.0
def evaluate_process_supervision(trajectory, ground_truth):
"""
Process supervision: grade each intermediate step.
Checks:
1. Tool call correctness (did the agent call the right tools?)
2. Efficiency (no more than N API calls)
3. No loops (no repeated identical calls)
4. Parameter validity (are dates in correct format?)
Returns:
float: average score across all checks (0.0 to 1.0)
"""
scores = []
# Check 1: Tool call correctness (from earlier function)
agent_tools = set(call['name'] for call in trajectory['tool_calls'])
expected_tools = set(call['name'] for call in ground_truth['expected_tool_calls'])
correct_tools = agent_tools & expected_tools
tool_score = len(correct_tools) / len(expected_tools) if expected_tools else 1.0
scores.append(tool_score)
# Check 2: Efficiency — no more than 5 API calls
num_calls = len(trajectory['tool_calls'])
efficiency_score = 1.0 if num_calls <= 5 else max(0, 1.0 - (num_calls - 5) * 0.1)
scores.append(efficiency_score)
# Check 3: No loops — check for repeated identical calls
call_signatures = [(call['name'], str(sorted(call['params'].items()))) for call in trajectory['tool_calls']]
has_loop = len(call_signatures) != len(set(call_signatures))
loop_score = 0.0 if has_loop else 1.0
scores.append(loop_score)
# Check 4: Parameter validity (simplified — check date format)
param_valid = True
for call in trajectory['tool_calls']:
for key, value in call['params'].items():
if 'date' in key.lower() or 'time' in key.lower():
# Check if it looks like a date (contains digits and dashes)
if not any(c.isdigit() for c in str(value)):
param_valid = False
param_score = 1.0 if param_valid else 0.0
scores.append(param_score)
return sum(scores) / len(scores)
# Example: Agent that succeeds but inefficiently
inefficient_trajectory = {
'tool_calls': [
{'name': 'search_flights', 'params': {'origin': 'NYC', 'destination': 'London', 'date': '2024-06-15'}},
{'name': 'search_flights', 'params': {'origin': 'NYC', 'destination': 'London', 'date': '2024-06-15'}},
{'name': 'search_flights', 'params': {'origin': 'NYC', 'destination': 'London', 'date': '2024-06-15'}},
{'name': 'search_flights', 'params': {'origin': 'NYC', 'destination': 'London', 'date': '2024-06-15'}},
{'name': 'search_flights', 'params': {'origin': 'NYC', 'destination': 'London', 'date': '2024-06-15'}},
{'name': 'search_flights', 'params': {'origin': 'NYC', 'destination': 'London', 'date': '2024-06-15'}},
{'name': 'search_flights', 'params': {'origin': 'NYC', 'destination': 'London', 'date': '2024-06-15'}},
{'name': 'search_flights', 'params': {'origin': 'NYC', 'destination': 'London', 'date': '2024-06-15'}},
{'name': 'search_flights', 'params': {'origin': 'NYC', 'destination': 'London', 'date': '2024-06-15'}},
{'name': 'search_flights', 'params': {'origin': 'NYC', 'destination': 'London', 'date': '2024-06-15'}},
{'name': 'search_flights', 'params': {'origin': 'NYC', 'destination': 'London', 'date': '2024-06-15'}},
{'name': 'search_flights', 'params': {'origin': 'NYC', 'destination': 'London', 'date': '2024-06-15'}},
{'name': 'search_flights', 'params': {'origin': 'NYC', 'destination': 'London', 'date': '2024-06-15'}},
{'name': 'search_flights', 'params': {'origin': 'NYC', 'destination': 'London', 'date': '2024-06-15'}},
{'name': 'search_flights', 'params': {'origin': 'NYC', 'destination': 'London', 'date': '2024-06-15'}},
{'name': 'search_flights', 'params': {'origin': 'NYC', 'destination': 'London', 'date': '2024-06-15'}},
{'name': 'book_flight', 'params': {'flight_id': 'BA178', 'passengers': 2}}
],
'final_output': 'Booking confirmed for 2 passengers on BA178 from NYC to London on June 15.'
}
# Same ground truth as before
ground_truth = {
'expected_tool_calls': [
{'name': 'search_flights', 'params': {'origin': 'NYC', 'destination': 'London', 'date': '2024-06-15'}},
{'name': 'book_flight', 'params': {'flight_id': 'BA178', 'passengers': 2}}
],
'expected_output': 'Booking confirmed for 2 passengers on BA178 from NYC to London on June 15.'
}
print("=== TRAJECTORY EVALUATION ===")
print("Agent called search_flights 16 times before booking. Final output is correct.")
print()
outcome_score = evaluate_outcome_only(inefficient_trajectory, ground_truth)
process_score = evaluate_process_supervision(inefficient_trajectory, ground_truth)
print(f"Outcome supervision score: {outcome_score:.2f}")
print(f" Interpretation: The final answer is correct, so outcome says it's perfect.")
print()
print(f"Process supervision score: {process_score:.2f}")
print(f" Interpretation: The process score catches the inefficiency and loops.")
print(f" The agent wasted API calls (16 instead of 2) and repeated the same call.")
print()
print("Key insight: Outcome-only metrics miss inefficient behavior.")
print("Process supervision gives a more realistic picture of agent quality.")
Notice the gap: outcome supervision gives a perfect 1.0, but process supervision reveals the agent is wasteful and potentially buggy. That’s the kind of insight you need before deploying.
Building Your Own Evaluation: Tools and Frameworks
You can’t rely on off-the-shelf benchmarks — your agent’s tasks are unique. You need to build custom evals. Fortunately, there are frameworks that make this structured and repeatable.
Let’s walk through three practical approaches: rule-based graders, LLM-as-judge, and human review.
Rule-based graders are like unit tests for your agent. They check objective facts: did the agent call the correct API? Are the parameters valid? Are there any obvious errors? These are fast, deterministic, and perfect for catching clear failures.
LLM-as-judge uses a strong LLM (like GPT-4) to evaluate your agent’s output. This works well for subjective criteria like helpfulness, tone, and safety. But it has biases: position bias (preferring the first option), self-enhancement (preferring its own outputs), and verbosity bias (preferring longer responses). The LLMs-as-Judges survey confirms these limitations — LLM judges are not a drop-in replacement for human evaluation.
Human review is the gold standard. You sample a subset of agent interactions and have a human rate them. This catches things automated systems miss. It’s expensive, so you use it strategically for edge cases.
Think of it like a test suite for your agent. Code graders are like unit tests — fast and precise. LLM judges are like a TA grading essays — useful but imperfect. Human review is the professor checking the hardest cases.
Let’s see all three in action.
# Three evaluation approaches: rule-based, LLM-as-judge, and human review
# This block is self-contained and runnable
# --- Approach 1: Rule-based grader ---
def rule_based_grader(trajectory, ground_truth):
"""
A deterministic grader that checks objective facts.
Returns a dict with pass/fail for each check.
"""
results = {}
# Check 1: Did the agent call the correct tools?
agent_tools = [call['name'] for call in trajectory['tool_calls']]
expected_tools = [call['name'] for call in ground_truth['expected_tool_calls']]
results['correct_tools'] = agent_tools == expected_tools
# Check 2: Are all required parameters present?
all_params_present = True
for expected_call in ground_truth['expected_tool_calls']:
# Find matching agent call
for agent_call in trajectory['tool_calls']:
if agent_call['name'] == expected_call['name']:
for param in expected_call['params']:
if param not in agent_call['params']:
all_params_present = False
break
results['all_params_present'] = all_params_present
# Check 3: No repeated identical calls (no loops)
call_signatures = [(call['name'], str(sorted(call['params'].items()))) for call in trajectory['tool_calls']]
results['no_loops'] = len(call_signatures) == len(set(call_signatures))
# Check 4: Reasonable number of calls (efficiency)
results['efficient'] = len(trajectory['tool_calls']) <= 5
return results
# --- Approach 2: LLM-as-judge (simulated) ---
def llm_as_judge(trajectory, prompt):
"""
Simulate an LLM judge that rates the agent's helpfulness on a scale of 1-5.
In practice, you'd call GPT-4 or another model. Here we simulate the output.
"""
# Simulated judge response
# In reality, this would be: response = openai.chat.completions.create(...)
simulated_score = 4 # Out of 5
simulated_reasoning = "The agent completed the task correctly and efficiently. The response was helpful and clear."
return {
'score': simulated_score,
'reasoning': simulated_reasoning
}
# --- Approach 3: Human review template ---
def human_review_template(trajectory):
"""
Create a structured review form for human evaluators.
Returns a dict with questions and blank answers.
"""
return {
'task_completed': None, # Yes/No
'tool_calls_correct': None, # Yes/No
'response_helpful': None, # 1-5 scale
'response_safe': None, # Yes/No
'comments': None # Free text
}
# Test data
trajectory = {
'tool_calls': [
{'name': 'search_flights', 'params': {'origin': 'NYC', 'destination': 'London', 'date': '2024-06-15'}},
{'name': 'book_flight', 'params': {'flight_id': 'BA178', 'passengers': 2}}
],
'final_output': 'Booking confirmed for 2 passengers on BA178 from NYC to London on June 15.'
}
ground_truth = {
'expected_tool_calls': [
{'name': 'search_flights', 'params': {'origin': 'NYC', 'destination': 'London', 'date': '2024-06-15'}},
{'name': 'book_flight', 'params': {'flight_id': 'BA178', 'passengers': 2}}
],
'expected_output': 'Booking confirmed for 2 passengers on BA178 from NYC to London on June 15.'
}
print("=== APPROACH 1: RULE-BASED GRADER ===")
rule_results = rule_based_grader(trajectory, ground_truth)
for check, passed in rule_results.items():
status = "PASS" if passed else "FAIL"
print(f" {check}: {status}")
print("Interpretation: All checks pass — the agent's behavior is objectively correct.")
print()
print("=== APPROACH 2: LLM-AS-JUDGE ===")
prompt = "Rate the agent's helpfulness on a scale of 1-5."
judge_result = llm_as_judge(trajectory, prompt)
print(f" Score: {judge_result['score']}/5")
print(f" Reasoning: {judge_result['reasoning']}")
print("Interpretation: The LLM judge gives a high score, but this is subjective.")
print(" The judge might be biased by response length or style.")
print()
print("=== APPROACH 3: HUMAN REVIEW TEMPLATE ===")
review_form = human_review_template(trajectory)
print(" Questions for human evaluator:")
for question, _ in review_form.items():
print(f" - {question}")
print("Interpretation: Human review is the gold standard but expensive.")
print(" Use it strategically for edge cases and calibration.")
Each approach has its place. Start with rule-based graders for objective checks. Use LLM-as-judge for subjective quality. Sample human review for edge cases. Combine them for a complete picture.
Putting It All Together: An Evaluation Pipeline for Your Agent
Let’s apply everything to the agent you built in Parts 4-6 of this series. We’ll design a complete evaluation pipeline: define tasks, collect trajectories, compute task completion, tool accuracy, pass@k, trajectory quality, and LLM-as-judge scores. Then we’ll interpret every number in plain English and decide whether the agent is ready for deployment.
Here’s the pipeline:
- Define a test set of 20 tasks that represent real user requests (e.g., “Book a flight from NYC to London on June 15 for 2 adults”).
- Run the agent 5 times per task (to measure reliability). Collect trajectories.
- Compute task completion rate per task and overall.
- Compute tool call accuracy and parameter correctness.
- Compute pass@5 and pass^5.
- Use LLM-as-judge to rate helpfulness and safety on a sample.
- Decide: is the agent ready? If not, what to improve?
Evaluation isn’t a single score — it’s a dashboard. You look at each dimension and decide where to invest your optimization effort.
Let’s build the full pipeline.
# Complete evaluation pipeline for an LLM agent
# This block is self-contained and runnable
import random
# --- Helper functions (reused from earlier blocks) ---
def evaluate_trajectory(trajectory, ground_truth):
"""Evaluate a single trajectory against ground truth."""
# Task completion
task_success = trajectory['final_output'] == ground_truth['expected_output']
# Tool accuracy
agent_tools = set(call['name'] for call in trajectory['tool_calls'])
expected_tools = set(call['name'] for call in ground_truth['expected_tool_calls'])
if len(agent_tools) == 0:
tool_precision = 0.0 if len(expected_tools) > 0 else 1.0
else:
correct_tools = agent_tools & expected_tools
tool_precision = len(correct_tools) / len(agent_tools)
if len(expected_tools) == 0:
tool_recall = 1.0 if len(agent_tools) == 0 else 0.0
else:
correct_tools = agent_tools & expected_tools
tool_recall = len(correct_tools) / len(expected_tools)
tool_accuracy = 2 * (tool_precision * tool_recall) / (tool_precision + tool_recall) if (tool_precision + tool_recall) > 0 else 0.0
# Parameter correctness
param_correct = 0
param_total = 0
for agent_call in trajectory['tool_calls']:
for expected_call in ground_truth['expected_tool_calls']:
if agent_call['name'] == expected_call['name']:
for key in expected_call['params']:
param_total += 1
if key in agent_call['params'] and agent_call['params'][key] == expected_call['params'][key]:
param_correct += 1
break
param_correctness = param_correct / param_total if param_total > 0 else 1.0
return {
'task_success': task_success,
'tool_accuracy': round(tool_accuracy, 2),
'param_correctness': round(param_correctness, 2)
}
def simulate_agent_run(task, ground_truth):
"""
Simulate an agent run on a task.
Returns a trajectory dict.
In practice, this would call your actual agent.
"""
# Simulate some randomness in agent behavior
# 70% chance of correct tool calls, 80% chance of correct output
correct_tools = random.random() < 0.7
correct_output = random.random() < 0.8
if correct_tools:
tool_calls = ground_truth['expected_tool_calls']
else:
# Wrong tool
tool_calls = [{'name': 'wrong_tool', 'params': {'error': 'hallucinated'}}]
if correct_output:
final_output = ground_truth['expected_output']
else:
final_output = "I encountered an error and couldn't complete the task."
return {
'tool_calls': tool_calls,
'final_output': final_output
}
# --- Step 1: Define test set ---
test_tasks = [
{
'id': 'book_flight_1',
'description': 'Book a flight from NYC to London on June 15 for 2 adults',
'ground_truth': {
'expected_tool_calls': [
{'name': 'search_flights', 'params': {'origin': 'NYC', 'destination': 'London', 'date': '2024-06-15'}},
{'name': 'book_flight', 'params': {'flight_id': 'BA178', 'passengers': 2}}
],
'expected_output': 'Booking confirmed for 2 passengers on BA178 from NYC to London on June 15.'
}
},
{
'id': 'cancel_hotel',
'description': 'Cancel hotel reservation for order ORD-123',
'ground_truth': {
'expected_tool_calls': [
{'name': 'cancel_reservation', 'params': {'order_id': 'ORD-123'}}
],
'expected_output': 'Reservation ORD-123 has been cancelled. Refund of $250 will be processed within 5-7 business days.'
}
},
{
'id': 'order_pizza',
'description': 'Order a large pepperoni pizza for delivery',
'ground_truth': {
'expected_tool_calls': [
{'name': 'search_menu', 'params': {'restaurant': 'pizza', 'item': 'pepperoni'}},
{'name': 'place_order', 'params': {'item': 'large pepperoni pizza', 'quantity': 1, 'delivery': True}}
],
'expected_output': 'Order placed: 1 large pepperoni pizza for delivery. Estimated arrival: 30-40 minutes.'
}
},
{
'id': 'schedule_meeting',
'description': 'Schedule a team meeting for next Monday at 2 PM',
'ground_truth': {
'expected_tool_calls': [
{'name': 'check_calendar', 'params': {'date': '2024-06-17'}},
{'name': 'create_event', 'params': {'title': 'Team Meeting', 'date': '2024-06-17', 'time': '14:00', 'duration': 60}}
],
'expected_output': 'Team meeting scheduled for Monday June 17 at 2:00 PM for 60 minutes.'
}
},
{
'id': 'send_email',
'description': 'Send an email to john@example.com with subject "Project Update"',
'ground_truth': {
'expected_tool_calls': [
{'name': 'send_email', 'params': {'to': 'john@example.com', 'subject': 'Project Update', 'body': 'Here is the latest update on the project.'}}
],
'expected_output': 'Email sent to john@example.com with subject "Project Update".'
}
}
]
# --- Step 2: Run agent 5 times per task ---
random.seed(123)
k = 5
all_trajectories = {}
for task in test_tasks:
task_id = task['id']
trajectories = []
for trial in range(k):
trajectory = simulate_agent_run(task_id, task['ground_truth'])
trajectories.append(trajectory)
all_trajectories[task_id] = trajectories
# --- Step 3-5: Compute all metrics ---
print("=" * 60)
print("AGENT EVALUATION REPORT")
print("=" * 60)
print()
overall_task_completion = 0
overall_tool_accuracy = []
overall_param_correctness = []
pass_at_k_results = []
pass_all_k_results = []
total_tasks = len(test_tasks)
for task in test_tasks:
task_id = task['id']
trajectories = all_trajectories[task_id]
ground_truth = task['ground_truth']
print(f"--- Task: {task['description']} ---")
# Evaluate each trial
trial_results = []
for i, traj in enumerate(trajectories):
result = evaluate_trajectory(traj, ground_truth)
trial_results.append(result)
# Task completion rate (first trial)
first_trial_success = trial_results[0]['task_success']
overall_task_completion += 1 if first_trial_success else 0
print(f" First trial success: {first_trial_success}")
# Average tool accuracy and parameter correctness
avg_tool_acc = sum(r['tool_accuracy'] for r in trial_results) / k
avg_param_corr = sum(r['param_correctness'] for r in trial_results) / k
overall_tool_accuracy.append(avg_tool_acc)
overall_param_correctness.append(avg_param_corr)
print(f" Avg tool accuracy: {avg_tool_acc:.2f}")
print(f" Avg param correctness: {avg_param_corr:.2f}")
# pass@k and pass^k
successes = [r['task_success'] for r in trial_results]
pass_at_k = any(successes)
pass_all_k = all(successes)
pass_at_k_results.append(pass_at_k)
pass_all_k_results.append(pass_all_k)
print(f" pass@{k}: {pass_at_k} (succeeded on at least one trial)")
print(f" pass^{k}: {pass_all_k} (succeeded on ALL {k} trials)")
print()
# --- Aggregate report ---
print("=" * 60)
print("OVERALL METRICS")
print("=" * 60)
print()
# Task completion rate
completion_rate = overall_task_completion / total_tasks
print(f"Task completion rate (first trial): {completion_rate:.0%}")
print(f" Interpretation: The agent completed {overall_task_completion} out of {total_tasks} tasks on the first try.")
print()
# Average tool accuracy
avg_tool_acc = sum(overall_tool_accuracy) / total_tasks
print(f"Average tool accuracy: {avg_tool_acc:.0%}")
print(f" Interpretation: The agent called the correct tool {avg_tool_acc:.0%} of the time across all trials.")
print()
# Average parameter correctness
avg_param_corr = sum(overall_param_correctness) / total_tasks
print(f"Average parameter correctness: {avg_param_corr:.0%}")
print(f" Interpretation: Parameters were correct {avg_param_corr:.0%} of the time.")
print()
# pass@k and pass^k
pass_at_k_rate = sum(pass_at_k_results) / total_tasks
pass_all_k_rate = sum(pass_all_k_results) / total_tasks
print(f"pass@{k} rate: {pass_at_k_rate:.0%}")
print(f" Interpretation: The agent succeeded on at least one of {k} tries for {pass_at_k_rate:.0%} of tasks.")
print(f"pass^{k} rate: {pass_all_k_rate:.0%}")
print(f" Interpretation: The agent succeeded on ALL {k} tries for only {pass_all_k_rate:.0%} of tasks.")
print(f" This is the reliability gap — the agent works sometimes but not consistently.")
print()
# --- Step 6: LLM-as-judge (simulated) ---
print("=" * 60)
print("LLM-AS-JUDGE EVALUATION (SAMPLE)")
print("=" * 60)
print()
# Sample 2 tasks for LLM judge
sample_tasks = random.sample(test_tasks, 2)
for task in sample_tasks:
task_id = task['id']
trajectory = all_trajectories[task_id][0] # First trial
# Simulated LLM judge
judge_score = random.randint(3, 5) # 3-5 out of 5
print(f"Task: {task['description']}")
print(f" Helpfulness score: {judge_score}/5")
print(f" Interpretation: The LLM judge rates this as {'good' if judge_score >= 4 else 'average'}.")
print()
# --- Step 7: Decision ---
print("=" * 60)
print("DEPLOYMENT READINESS ASSESSMENT")
print("=" * 60)
print()
# Criteria for deployment
if completion_rate >= 0.8 and pass_all_k_rate >= 0.6:
print("RECOMMENDATION: READY FOR PRODUCTION")
print(" The agent completes most tasks and is reliable across multiple trials.")
print(" Continue monitoring and collect more data for edge cases.")
elif completion_rate >= 0.6:
print("RECOMMENDATION: CONDITIONAL DEPLOYMENT")
print(" The agent works for most tasks but reliability needs improvement.")
print(" Consider:")
print(" - Improving tool-calling accuracy (current: {:.0%})".format(avg_tool_acc))
print(" - Reducing parameter errors (current: {:.0%} correct)".format(avg_param_corr))
print(" - Adding fallback mechanisms for when the agent fails")
else:
print("RECOMMENDATION: NOT READY FOR PRODUCTION")
print(" The agent fails too often. Focus on:")
print(" - Improving the base model or prompt")
print(" - Adding more training data for common tasks")
print(" - Implementing better error handling")
When you run this pipeline, you get a dashboard of metrics. Each number tells you something specific about your agent’s performance. The task completion rate tells you if the agent gets the job done. Tool accuracy tells you if it uses the right tools. pass^k tells you if it’s reliable enough for production.
This is the evaluation framework you need. It replaces the vague “looks good to me” with concrete, actionable numbers.
Conclusion and Next Steps
Let’s recap what we’ve learned. Normal metrics like BLEU, ROUGE, and accuracy fail for agents because they ignore task completion, tool use, and multi-step reasoning. Your agent might write perfect text but fail at the actual task.
We replaced those with a multi-dimensional evaluation framework:
- Task completion rate: Did the agent accomplish the user’s goal?
- Tool call accuracy: Did it call the right tools?
- Parameter correctness: Were the arguments valid?
- pass@k and pass^k: How reliable is the agent across multiple trials?
- Trajectory quality: Did the agent get there efficiently and safely?
- LLM-as-judge scores: How helpful and safe is the agent’s behavior?
You now have a framework and code to evaluate any LLM agent. Use frameworks like OpenAI Evals, LangSmith, and agentevals to build structured, repeatable evaluations. Always interpret numbers in plain English — a score without context is meaningless.
In Part 8 of this series, we’ll use these evaluation results to improve your agent. We’ll fix tool-calling errors, improve reliability, and add guardrails. Evaluation-driven development turns vague “it seems okay” into “here’s exactly what to fix.”
Check Your Understanding
Remember: What are the three main dimensions of agent evaluation discussed in this article?
Understand: Why does BLEU score fail to capture agent failures? Give a concrete example.
Apply: Given a trajectory where the agent calls search_flights 12 times before finally booking, which metric would catch this inefficiency — outcome supervision or process supervision?
Analyze: Compare pass@k and pass^k. Under what conditions would pass@k be high but pass^k be low? What does this tell you about the agent?
Evaluate: You have an agent that completes 85% of tasks on the first try but only 30% of tasks on all 5 tries. Is it ready for production? What risks does it pose?
Create: Design a rule-based grader for a customer support agent. What three checks would you include? Write the pseudocode for one of them.
Related Articles
- Part 6: RLHF Explained: How ChatGPT Was Actually Trained to Be Helpful — The training techniques that make models follow instructions, which is the foundation for building agents that can be evaluated.
- Part 5: Agentic RAG: Retrieval That an Agent Decides to Use — Understanding how agents decide to use tools, which is exactly what our evaluation metrics measure.
Apply What You Learned is for Supporter and Insider subscribers.
Subscribe to unlock the exercises on this post.
See plansRelated articles
- LLMs & GenAI Under review
What Makes an "Agent" an Agent? The Loop Behind Every LLM Agent
You've built a chatbot. It answers questions, maybe even holds a decent conversation. But then you ask it to check the weather in Tokyo right now, and it says, "I don't have access to real-time data." Sound familiar?
- LLMs & GenAI Under review
Evaluating LLM Output Beyond "It Looks Right"
Learn systematic methods for evaluating LLM output across correctness, relevance, and safety using automated metrics, human review, and hybrid approaches.
- LLMs & GenAI Under review
Giving An Llm Memory Short Term Context Vs Long Te
Remember the agent you built in Part 2? It was great at calling tools. You could ask it to check the weather, look up a fact, or do some math, and it would figure out which tool to use and hand back the right answer.
- LLMs & GenAI Under review
Multi Agent Systems When One Llm Isn T Enough
You've done the work. In Part 1, you built a ReAct agent that could think step-by-step and call tools. In Part 2, you gave it a full toolbox — weather lookups, math calculations, database queries.
Looking for something else?
Search every article by title, summary or topic.