Building A Simple Research Or Trading Agent As A L
The Problem: You’ve Built a Loop — Now What?
You’ve done the hard work. In Part 7, you built a working agent loop: LLM call → tool execution → append result → repeat. It runs. It calls tools. It responds. But now you’re staring at the terminal, and a familiar question creeps in: Now what?
It’s like building a car engine and then realizing you have nowhere to drive. The loop is powerful, but without a concrete task, it’s just a spinning gear. That’s the paralysis this article solves.
We’re going to build two real, useful capstone agents. A research agent that answers complex questions by searching the web and reading articles. And a trading agent that looks at stock prices and news sentiment to make a buy/sell/hold call. Both use the exact same loop you already built — they just get different tools and a different system prompt.
Here’s the promise: by the end of this article, you’ll have a working, extendable agent for either task, plus the pattern to build your own. You’ll also know the hard parts — designing good tools, handling real-world API failures, and knowing when to stop the loop. As Leonie Monigatti puts it in her excellent tutorial, the goal is to “understand the loop before adopting a framework.” We’re doing exactly that.
Let’s start with the research agent.
What We’re Building: The Research Agent
The research agent’s job is simple: given a complex question (e.g., “What are the latest advances in quantum error correction?”), it should search the web, read relevant articles, and synthesize a concise answer. This is the “deep research” pattern that tools like Perplexity and ChatGPT’s browsing mode use.
To do this, the agent needs two tools:
web_search: Takes a query string, returns a list of URLs and snippets.read_page: Takes a URL, returns the page’s text content.
Here’s the catch: the search tool returns URLs, not answers. The agent must decide which URLs to read, then call read_page on each one. That’s a multi-step tool-use pattern — the loop runs until the agent decides it has enough information.
Let’s define the tool schemas first. We’ll use JSON schemas that describe each tool’s name, description, and parameters. This is the same pattern from Part 3 of the series.
# Tool schemas for the research agent
# Self-contained: defines schemas and prints them
web_search_schema = {
"name": "web_search",
"description": "Search the web for a given query. Returns a list of result URLs and snippets.",
"parameters": {
"type": "object",
"properties": {
"query": {
"type": "string",
"description": "The search query, e.g., 'quantum error correction 2025'"
}
},
"required": ["query"]
}
}
read_page_schema = {
"name": "read_page",
"description": "Fetch the full text content of a web page given its URL.",
"parameters": {
"type": "object",
"properties": {
"url": {
"type": "string",
"description": "The full URL of the page to read"
}
},
"required": ["url"]
}
}
print("Research agent tool schemas:")
print("1. web_search:", web_search_schema["name"])
print(" Description:", web_search_schema["description"])
print(" Parameters:", list(web_search_schema["parameters"]["properties"].keys()))
print()
print("2. read_page:", read_page_schema["name"])
print(" Description:", read_page_schema["description"])
print(" Parameters:", list(read_page_schema["parameters"]["properties"].keys()))
Now let’s implement the actual tool functions. In a real project, web_search would call a search API (like SerpAPI or Bing), and read_page would fetch HTML and extract text. For this demonstration, we’ll mock them — the important part is the pattern.
# Mock tool implementations for the research agent
# Self-contained: defines functions and runs a simulated agent loop
import json
# Mock search results database
MOCK_SEARCH_RESULTS = {
"quantum error correction 2025": [
{"url": "https://example.com/qec1", "snippet": "New surface code achieves 99.9% fidelity..."},
{"url": "https://example.com/qec2", "snippet": "Google's Willow chip demonstrates error correction at scale..."},
{"url": "https://example.com/qec3", "snippet": "Topological qubits show promise for fault-tolerant computing..."}
]
}
MOCK_PAGE_CONTENT = {
"https://example.com/qec1": "Surface codes have been a leading approach for quantum error correction. Recent work by Google Quantum AI demonstrates a surface code with 99.9% fidelity on a 53-qubit processor. This is a significant step toward fault-tolerant quantum computing.",
"https://example.com/qec2": "Google's Willow chip, announced in 2024, uses a new architecture that reduces error rates exponentially as more qubits are added. This is a key milestone for scalable quantum computing.",
"https://example.com/qec3": "Topological qubits, based on anyons, offer inherent protection against decoherence. Microsoft recently reported progress in creating topological qubits, though practical devices remain years away."
}
def web_search(query):
"""Mock search: returns results from a hardcoded database."""
return MOCK_SEARCH_RESULTS.get(query.lower(), [])
def read_page(url):
"""Mock page reader: returns content from a hardcoded database."""
return MOCK_PAGE_CONTENT.get(url, "Page not found.")
# Simulate the agent loop manually (without an LLM) to show the multi-step pattern
# In a real agent, the LLM decides which tool to call next.
# Here we hardcode the decisions to illustrate the flow.
print("=== Research Agent Simulation ===")
print("User query: 'What are the latest advances in quantum error correction?'")
print()
# Step 1: Agent decides to search
print("Step 1: Agent calls web_search('quantum error correction 2025')")
results = web_search("quantum error correction 2025")
print(f"Search returned {len(results)} results.")
for r in results:
print(f" - {r['url']}: {r['snippet'][:50]}...")
print()
# Step 2: Agent decides to read the first result
print("Step 2: Agent calls read_page('https://example.com/qec1')")
content = read_page("https://example.com/qec1")
print(f"Page content (first 100 chars): {content[:100]}...")
print()
# Step 3: Agent decides to read the second result
print("Step 3: Agent calls read_page('https://example.com/qec2')")
content2 = read_page("https://example.com/qec2")
print(f"Page content (first 100 chars): {content2[:100]}...")
print()
# Step 4: Agent synthesizes answer
print("Step 4: Agent synthesizes final answer")
print("Final answer:")
print("Recent advances in quantum error correction include:")
print("- Surface codes achieving 99.9% fidelity (Google Quantum AI)")
print("- Google's Willow chip reducing error rates exponentially")
print("- Progress on topological qubits (Microsoft)")
Notice the flow: the agent called web_search once, then read_page twice, then answered. In a real agent, the LLM decides each step based on the conversation history. The loop runs until the LLM outputs a final response instead of a tool call. This is exactly the pattern from Part 7, now applied to a real task.
As the DEV Community article on building a research agent notes, the key insight is that “the agent must be able to decide when it has enough information to answer.” Our loop handles that naturally: the LLM keeps calling tools until it’s ready to respond.
The Trading Agent: From Research to Decision
Now let’s pivot to the trading agent. Same loop, different tools. The trading agent’s job: given a stock ticker (e.g., “AAPL”), fetch recent price data and news sentiment, then output a buy/sell/hold recommendation with a confidence score.
This agent needs three tools:
get_stock_price: Takes a ticker symbol, returns the current price and recent price history.get_news_sentiment: Takes a ticker symbol, returns recent news headlines and a sentiment score (e.g., -1 to +1).calculator: Takes a mathematical expression, returns the result. This is needed because LLMs are notoriously bad at arithmetic.
Why the calculator tool? Because the agent might want to compute a moving average or price change percentage. Offloading math to a tool avoids arithmetic errors.
Let’s define the schemas and mock implementations.
# Tool schemas and mock implementations for the trading agent
# Self-contained: defines everything needed for the simulation
import json
# Schemas
get_stock_price_schema = {
"name": "get_stock_price",
"description": "Get the current stock price and recent price history for a given ticker symbol.",
"parameters": {
"type": "object",
"properties": {
"ticker": {
"type": "string",
"description": "Stock ticker symbol, e.g., 'AAPL'"
}
},
"required": ["ticker"]
}
}
get_news_sentiment_schema = {
"name": "get_news_sentiment",
"description": "Get recent news headlines and a sentiment score (-1 to +1) for a given ticker.",
"parameters": {
"type": "object",
"properties": {
"ticker": {
"type": "string",
"description": "Stock ticker symbol, e.g., 'AAPL'"
}
},
"required": ["ticker"]
}
}
calculator_schema = {
"name": "calculator",
"description": "Evaluate a mathematical expression and return the result. Use this for any arithmetic.",
"parameters": {
"type": "object",
"properties": {
"expression": {
"type": "string",
"description": "A mathematical expression, e.g., '(150 - 140) / 140 * 100'"
}
},
"required": ["expression"]
}
}
# Mock data
MOCK_PRICES = {
"AAPL": {"current": 175.30, "history": [172.10, 173.45, 174.20, 175.00, 175.30]},
"GOOGL": {"current": 142.80, "history": [140.50, 141.20, 142.00, 142.60, 142.80]},
"MSFT": {"current": 378.90, "history": [375.00, 376.40, 377.10, 378.20, 378.90]}
}
MOCK_SENTIMENT = {
"AAPL": {"headlines": ["Apple announces new AI features", "iPhone sales beat expectations"], "score": 0.8},
"GOOGL": {"headlines": ["Google faces antitrust ruling", "Alphabet earnings mixed"], "score": -0.2},
"MSFT": {"headlines": ["Microsoft Azure growth slows", "Copilot adoption rises"], "score": 0.3}
}
def get_stock_price(ticker):
"""Mock price fetcher."""
return MOCK_PRICES.get(ticker.upper(), {"current": 0, "history": []})
def get_news_sentiment(ticker):
"""Mock sentiment fetcher."""
return MOCK_SENTIMENT.get(ticker.upper(), {"headlines": [], "score": 0})
def calculator(expression):
"""Safe calculator using Python's eval with limited scope."""
try:
# Only allow basic arithmetic
result = eval(expression, {"__builtins__": {}}, {})
return result
except Exception as e:
return f"Error: {e}"
# Simulate the trading agent loop
print("=== Trading Agent Simulation ===")
print("User request: 'Analyze AAPL and give a buy/sell/hold recommendation.'")
print()
# Step 1: Get price
print("Step 1: Agent calls get_stock_price('AAPL')")
price_data = get_stock_price("AAPL")
print(f"Current price: ${price_data['current']}")
print(f"Recent prices: {price_data['history']}")
print()
# Step 2: Get sentiment
print("Step 2: Agent calls get_news_sentiment('AAPL')")
sentiment_data = get_news_sentiment("AAPL")
print(f"Headlines: {sentiment_data['headlines']}")
print(f"Sentiment score: {sentiment_data['score']}")
print()
# Step 3: Compute price change percentage using calculator
print("Step 3: Agent calls calculator to compute price change over 5 days")
prices = price_data['history']
change_pct = calculator(f"({prices[-1]} - {prices[0]}) / {prices[0]} * 100")
print(f"Price change over 5 days: {change_pct:.2f}%")
print()
# Step 4: Agent synthesizes decision
print("Step 4: Agent synthesizes recommendation")
print("Reasoning:")
print(f"- Price trend: Up {change_pct:.2f}% over 5 days (bullish)")
print(f"- News sentiment: {sentiment_data['score']} (positive)")
print("Decision: BUY with confidence 0.7")
print("Interpretation: The agent is 70% confident in a buy recommendation based on positive price momentum and favorable news. This is a learning project — not financial advice!")
The trading agent combines quantitative data (price) with qualitative data (sentiment) into a single decision. The calculator tool ensures arithmetic is correct. This is a simplified version of what real trading agents do, as seen in frameworks like TradingAgents (a multi-agent system for financial trading). But as the Prompt Engineering Guide notes, the core pattern is the same: query → tool-decision → execution → observation → response.
The Hardest Part: When Agents Go Wrong
Let’s be honest: agents fail. A lot. Research from Openlayer shows that tool-calling failures happen 3-15% of the time in production, and some workflows see rates near 41%. Three common failure modes are:
- Hallucinated tool invocations: The LLM calls a tool that doesn’t exist or passes wrong arguments.
- Infinite retry loops: The agent keeps retrying a failed tool call without learning from the failure.
- Agent paralysis: The agent gets contradictory information and can’t decide.
Here’s the scary part: if the tool call fails 10% of the time, and the agent makes 5 tool calls per task, there’s about a 40% chance at least one call fails. That’s why error handling is critical.
Let’s add a max_iterations guard and try/except blocks to handle failures gracefully. This is the pattern used by Anthropic’s “Ring 2” agentic loop, which runs while response.stop_reason == 'tool_use' with a safety limit.
# Agent loop with max_iterations guard and error handling
# Self-contained: demonstrates failure recovery
import time
# Mock tools that sometimes fail
import random
random.seed(42)
def unreliable_tool(input_data):
"""Simulates a tool that fails 30% of the time."""
if random.random() < 0.3:
raise ConnectionError("API timeout: could not reach external service.")
return f"Successfully processed: {input_data}"
def run_agent_with_guard(max_iterations=5):
"""Runs a simulated agent loop with max iterations and error handling."""
iteration = 0
messages = []
while iteration < max_iterations:
print(f"\nIteration {iteration + 1}")
# Simulate LLM deciding to call a tool (hardcoded for demo)
if iteration == 0:
tool_name = "unreliable_tool"
tool_input = "data_1"
else:
# After first iteration, LLM decides to stop
print("LLM decides to respond directly.")
break
print(f"LLM calls tool: {tool_name}({tool_input})")
try:
result = unreliable_tool(tool_input)
print(f"Tool succeeded: {result}")
messages.append({"role": "tool", "content": result})
except Exception as e:
print(f"Tool FAILED: {e}")
# Return a structured error message so the LLM can try a different approach
error_msg = f"Tool call failed with error: {e}. Try a different approach or retry later."
messages.append({"role": "tool", "content": error_msg})
iteration += 1
if iteration >= max_iterations:
print(f"\nReached max iterations ({max_iterations}). Agent stopped.")
print(f"\nFinal message history length: {len(messages)}")
return messages
print("=== Agent Loop with Error Handling ===")
print("Max iterations: 5")
print("Tool failure rate: ~30%")
print()
messages = run_agent_with_guard(max_iterations=5)
print("\n=== Summary ===")
print("The agent handled the failure gracefully and continued. Without the guard, it could have retried forever.")
When you run this, you’ll see the tool fail (about 30% of the time) and the agent logs the error. The loop continues because we caught the exception. The max_iterations guard ensures the agent stops even if it keeps trying to call tools. As the arXiv paper by Hou et al. documents, mainstream frameworks all ship such safeguards — and now yours does too.
Putting It All Together: The Capstone Agent Class
Now let’s combine everything into a single Agent class that can be configured for either research or trading. This is the graduation moment — you’ve built a reusable agent framework.
The class will have:
- A tools registry (dict of tool name → tool function)
- A message history list
- A max-iteration guard
- A
run()method that executes the loop
# The Agent class: reusable for research or trading
# Self-contained: defines the class and runs both agents
import json
class Agent:
def __init__(self, tools, system_prompt, max_iterations=10):
"""
Initialize the agent with a list of tool functions and a system prompt.
Args:
tools: list of dicts with 'name', 'description', 'parameters', and 'function'
system_prompt: str, the system prompt that defines the agent's behavior
max_iterations: int, maximum number of tool calls before forced stop
"""
self.tools = {tool['name']: tool for tool in tools}
self.system_prompt = system_prompt
self.max_iterations = max_iterations
self.messages = [{"role": "system", "content": system_prompt}]
def run(self, user_input):
"""
Run the agent loop on a user input.
For this demo, we simulate the LLM's decisions with a hardcoded sequence.
"""
self.messages.append({"role": "user", "content": user_input})
print(f"User: {user_input}\n")
iteration = 0
while iteration < self.max_iterations:
# Simulate LLM deciding to call a tool or respond
# In a real agent, this would be an LLM API call.
# Here we use a simple rule: call the first tool if available, then stop.
if iteration == 0:
# Pick the first tool in the registry
tool_name = list(self.tools.keys())[0]
tool_def = self.tools[tool_name]
print(f"Agent decides to call: {tool_name}")
# Generate a mock argument based on the tool's parameters
# For simplicity, we pass a hardcoded value
if tool_name == "web_search":
args = {"query": user_input}
elif tool_name == "get_stock_price":
args = {"ticker": "AAPL"}
else:
args = {}
print(f"Arguments: {args}")
try:
result = tool_def['function'](**args)
print(f"Tool result: {str(result)[:100]}...")
self.messages.append({"role": "tool", "content": str(result)})
except Exception as e:
print(f"Tool failed: {e}")
self.messages.append({"role": "tool", "content": f"Error: {e}"})
iteration += 1
else:
# After first tool, agent decides to respond directly
print("Agent decides to respond directly.")
final_answer = f"Based on the information gathered, here is my answer to '{user_input}'."
print(f"\nAgent: {final_answer}")
self.messages.append({"role": "assistant", "content": final_answer})
return final_answer
print(f"Reached max iterations ({self.max_iterations}). Forcing stop.")
return "Agent stopped due to iteration limit."
# --- Instantiate Research Agent ---
print("=" * 60)
print("RESEARCH AGENT")
print("=" * 60)
# Define research tools (using mocks from earlier)
def web_search(query):
return f"Search results for '{query}': [url1, url2, url3]"
def read_page(url):
return f"Content of {url}: ... (page text)"
research_tools = [
{"name": "web_search", "description": "Search the web", "parameters": {"query": "string"}, "function": web_search},
{"name": "read_page", "description": "Read a web page", "parameters": {"url": "string"}, "function": read_page}
]
research_prompt = "You are a research assistant. Use web_search and read_page to answer questions."
research_agent = Agent(tools=research_tools, system_prompt=research_prompt, max_iterations=5)
research_agent.run("What are the latest advances in quantum error correction?")
print("\n" + "=" * 60)
print("TRADING AGENT")
print("=" * 60)
# Define trading tools (using mocks)
def get_stock_price(ticker):
return f"Current price of {ticker}: $175.30"
def get_news_sentiment(ticker):
return f"Sentiment for {ticker}: positive (score 0.8)"
def calculator(expression):
try:
return eval(expression, {"__builtins__": {}}, {})
except:
return "Error"
trading_tools = [
{"name": "get_stock_price", "description": "Get stock price", "parameters": {"ticker": "string"}, "function": get_stock_price},
{"name": "get_news_sentiment", "description": "Get news sentiment", "parameters": {"ticker": "string"}, "function": get_news_sentiment},
{"name": "calculator", "description": "Evaluate math expression", "parameters": {"expression": "string"}, "function": calculator}
]
trading_prompt = "You are a trading assistant. Use tools to analyze stocks and give buy/sell/hold recommendations."
trading_agent = Agent(tools=trading_tools, system_prompt=trading_prompt, max_iterations=5)
trading_agent.run("Analyze AAPL and give a recommendation.")
Notice how the same Agent class handles both tasks. The only differences are the tools and the system prompt. This is the power of the pattern: once you have the loop, you can build any agent by defining its tools.
What’s Next: From DIY to Frameworks
You’ve built a working agent from scratch. Now you can appreciate what frameworks like LangGraph, CrewAI, and the OpenAI Agents SDK provide: state management, parallel tool calls, persistence, and error recovery. But the core pattern is the same.
Let’s look at a minimal LangGraph version of the research agent. LangGraph uses a StateGraph where nodes are functions and edges define the flow. The conditional edge checks if the agent should call a tool or respond.
# Minimal LangGraph-style research agent (conceptual, not runnable without langgraph library)
# This shows the pattern, not a full implementation
"""
from langgraph.graph import StateGraph, END
from typing import TypedDict, List
class AgentState(TypedDict):
messages: List[dict]
next_step: str
def call_model(state):
# Call LLM to decide next action
response = llm.invoke(state['messages'])
return {"messages": state['messages'] + [response], "next_step": "tools" if response.tool_calls else "respond"}
def execute_tools(state):
# Execute tool calls
for tool_call in state['messages'][-1].tool_calls:
result = tool_registry[tool_call.name](**tool_call.args)
state['messages'].append({"role": "tool", "content": result})
return state
def should_continue(state):
return state['next_step']
# Build graph
graph = StateGraph(AgentState)
graph.add_node("agent", call_model)
graph.add_node("tools", execute_tools)
graph.set_entry_point("agent")
graph.add_conditional_edges("agent", should_continue, {"tools": "tools", "respond": END})
graph.add_edge("tools", "agent")
app = graph.compile()
"""
print("LangGraph pattern: nodes (agent, tools) connected by conditional edges.")
print("The same loop: agent decides → tools execute → agent decides again.")
print("Frameworks add state management, parallel calls, and persistence.")
As the LangGraph documentation shows, the framework formalizes the loop you already built. The DataCamp tutorial on LangGraph explains that it adds “State, Nodes, Edges, reducers, and persistence” — all useful for production, but the core idea is the same.
Recap: What You Built Today
Let’s celebrate what you accomplished:
- You built a reusable
Agentclass with a tools registry and max-iteration guard. - You built a research agent that searches the web and reads pages to answer complex questions.
- You built a trading agent that fetches price data and news sentiment to make a buy/sell/hold decision.
- You learned how to handle tool call failures and infinite loops.
- You saw how the same pattern scales to frameworks like LangGraph.
You now have a working, extendable agent for either task. The next step is to try it with real APIs — connect web_search to SerpAPI, get_stock_price to Yahoo Finance, and swap the simulated LLM decisions for actual GPT-4 or Claude calls.
Check Your Understanding
Remember: What are the three common failure modes of tool-using agents?
Understand: Explain in your own words why a max-iteration guard is essential for an agent loop.
Apply: Given a new tool (e.g., a database query tool), write the JSON schema and integrate it into the agent class.
Analyze: Compare the DIY agent loop with LangGraph’s StateGraph. What does the framework add? What does it hide?
Evaluate: When would you choose to build a DIY agent vs. using a framework? Justify your answer.
Create: Design a new agent for a task of your choice (e.g., a travel planner, a recipe finder). List the tools it needs and sketch the expected tool-call sequence.
Related Articles
- Part 1: What Is an LLM Agent? — The intuition behind tool-using agents.
- Part 3: Tool Schemas and Function Calling — How to define tools in JSON.
- Part 7: The Agent Loop — The core request/response cycle this capstone builds on.
- LangGraph 101: Let’s Build A Deep Research Agent (Towards Data Science) — A framework-based version of the research agent.
- TradingAgents — Multi-Agents LLM Financial Trading Framework (GitHub) — A production-scale multi-agent trading system.
Apply What You Learned is for Supporter and Insider subscribers.
Subscribe to unlock the exercises on this post.
See plansRelated articles
- LLMs & GenAI Under review
Multi Agent Systems When One Llm Isn T Enough
You've done the work. In Part 1, you built a ReAct agent that could think step-by-step and call tools. In Part 2, you gave it a full toolbox — weather lookups, math calculations, database queries.
- Statistics Under review
Google Cloud Automl Vs Azure Automl Vs Aws Sagemak
You've finally decided to let AutoML handle the grunt work. You've read about what it automates and what it doesn't. You're sold on the idea. Now your boss comes by your desk and says, "Great, we're using AutoML.
- Statistics Under review
Pearson vs Spearman vs Kendall: Picking the Right Correlation for Your Data
Learn when to use Pearson, Spearman, or Kendall correlation in Python — and why a weak score may mean you simply chose the wrong tool for your data.
- LLMs & GenAI Under review
Giving An Llm Memory Short Term Context Vs Long Te
Remember the agent you built in Part 2? It was great at calling tools. You could ask it to check the weather, look up a fact, or do some math, and it would figure out which tool to use and hand back the right answer.
Looking for something else?
Search every article by title, summary or topic.