This article is published in English.
Practical notes: Deep Agents in Action: Building a Multi-Agent Research System
Operable walkthrough of Practical notes: Deep Agents in Action: Building a Multi-Agent Research System: contracts, checks, and drop-in code slots for teams shipping this pattern.
This walkthrough rebuilds the path from raw materials to a working system for: Deep Agents in Action: Building a Multi-Agent Research System. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Introduction: The Deep Agents Architecture
When working through the Introduction The Deep Agents stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
What is the Deep Agents Pattern?
When working through the What is the Deep stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
| Single Agent | Deep Agents |
| ---------------------------------- | ------------------------------------------------- |
| One context window gets overloaded | Each agent has an isolated, focused context |
| All reasoning in one prompt | Specialised reasoning per domain |
| Hard to scale | Add specialists without changing the orchestrator |
| Hard to debug | Full delegation trace for auditability |
Deep Agents overview
When working through the Deep Agents overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Deep Agents overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Core Concepts of LangChain Deep Agents
The Core Concepts of LangChain stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Technical Implementation
The Technical Implementation stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Architecture Overview
The Architecture Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
User Task
└─ OrchestratorAgent (Planning & Synthesis)
├─ ResearcherAgent (web_search)
├─ AnalystAgent (calculator, code_executor)
└─ WriterAgent (file_reader)
The Architecture Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Technical Stack
For the Technical Stack stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Code Structure
For the Code Structure stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
deepagents-usecase/
├── agents/
│ ├── __init__.py # Package exports
│ ├── base.py # Abstract BaseAgent + ReAct loop
│ ├── llm_client.py # LLM adapter (Ollama/llama.cpp/OpenAI/Anthropic)
│ ├── messages.py # Typed message protocol
│ ├── orchestrator.py # OrchestratorAgent (top-level)
│ ├── researcher.py # ResearcherAgent specialist
│ ├── analyst.py # AnalystAgent specialist
│ └── writer.py # WriterAgent specialist
├── tools/
│ ├── __init__.py
│ ├── base.py # BaseTool + ToolResult
│ ├── calculator.py # Safe AST-based arithmetic evaluator
│ ├── code_executor.py # Sandboxed Python execution (exec with allow-list)
│ ├── file_reader.py # Sandboxed file reading (input/ only)
│ └── web_search.py # Web search (stub + live Tavily)
├── memory/
│ ├── __init__.py
│ └── store.py # AgentMemoryStore (short/long-term/episodic)
├── config/
│ ├── __init__.py
│ └── settings.py # Centralised env-based configuration
├── tests/
│ ├── test_tools.py # Tool unit tests (71 tests)
│ ├── test_memory.py # Memory unit tests (24 tests)
│ ├── test_agents.py # Agent integration tests (92 tests)
│ └── test_code_executor.py # Code Executor tests (134 tests)
├── input/ # Input documents (content gitignored)
├── output/ # Generated reports (content gitignored)
├── memory/ # Persistent agent memory (JSON files)
├── scripts/
│ ├── start.sh # Launch Streamlit in detached mode
│ ├── stop.sh # Gracefully stop the application
│ ├── cleanup.sh # Remove .venv, __pycache__, etc.
│ └── check_code_executor.py # Standalone Code Executor sanity-check
├── Docs/
│ ├── Architecture.md # Mermaid architecture diagrams
│ ├── Quickstart.md # Step-by-step getting started guide
│ ├── API.md # Full public API reference
│ └── CodeExecutorVerification.md # Code Executor test & verification guide
├── app.py # Streamlit web UI
├── main.py # CLI entry point
├── requirements.txt
├── .env.example # Environment variable template
└── README.md
# =============================================================================
# Deep Agents System — Environment Configuration
# =============================================================================
# Copy this file to .env and fill in your values.
# NEVER commit the .env file to version control.
#
# Usage:
# cp .env.example .env
# # edit .env with your actual keys
# =============================================================================
# ─── LLM Provider ──────────────────────────────────────────────────────────
# Select ONE provider. Comment out the others.
# Option A: Ollama (local, default — no API key needed)
LLM_PROVIDER=ollama
OLLAMA_BASE_URL=http://localhost:11434
OLLAMA_MODEL=llama3.2
# Option B: llama.cpp (local)
# LLM_PROVIDER=llamacpp
# LLAMACPP_BASE_URL=http://localhost:9931/v1
# LLAMACPP_MODEL=local-model
# Option C: OpenAI
# LLM_PROVIDER=openai
# OPENAI_API_KEY=sk-...
# OPENAI_MODEL=gpt-4o
# Option D: Anthropic
# LLM_PROVIDER=anthropic
# ANTHROPIC_API_KEY=sk-ant-...
# ANTHROPIC_MODEL=claude-3-5-sonnet-20241022
# ─── Application Settings ──────────────────────────────────────────────────
APP_PORT=8501
APP_HOST=0.0.0.0
LOG_LEVEL=INFO
# ─── Agent Configuration ───────────────────────────────────────────────────
# Maximum reasoning steps per agent (lower = faster on slow local LLMs)
MAX_AGENT_STEPS=8
# Maximum tokens per LLM call (1024 is enough for ReAct; raise for longer reports)
MAX_TOKENS=1024
# Temperature for LLM responses (0.0 = deterministic, 1.0 = creative)
TEMPERATURE=0.1
# ─── Memory & Storage ──────────────────────────────────────────────────────
# Directory for persistent agent memory (relative to project root)
MEMORY_DIR=./memory
# Maximum number of memories to retain per agent
MAX_MEMORY_ENTRIES=100
# ─── Tool Configuration ────────────────────────────────────────────────────
# Enable or disable specific tools (true/false)
TOOL_WEB_SEARCH_ENABLED=true
TOOL_CALCULATOR_ENABLED=true
TOOL_FILE_READER_ENABLED=true
TOOL_CODE_EXECUTOR_ENABLED=true
# Web search stub — set to real Tavily/SerpAPI key for live search
# TAVILY_API_KEY=tvly-...
# ─── Output ────────────────────────────────────────────────────────────────
OUTPUT_DIR=./output
Key Code Excerpts
For the Key Code Excerpts stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Key Code Excerpts stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Safe AST Arithmetic Evaluator (tools/calculator.py)
When working through the Safe AST Arithmetic Evaluator stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
# Whitelisted AST operators and math functions
_SAFE_OPERATORS = {
ast.Add: operator.add,
ast.Sub: operator.sub,
ast.Mult: operator.mul,
ast.Div: operator.truediv,
ast.Pow: operator.pow,
ast.USub: operator.neg,
}
_SAFE_FUNCTIONS = {
"abs": abs, "round": round, "sqrt": math.sqrt,
"sin": math.sin, "cos": math.cos, "log": math.log,
"pi": math.pi, "e": math.e,
}
def _safe_eval(node: ast.expr) -> float:
"""Recursively evaluate an AST expression node in a safe sandbox."""
if isinstance(node, ast.Constant):
if isinstance(node.value, (int, float)):
return float(node.value)
raise ValueError(f"Unsupported constant type: {type(node.value).__name__}")
if isinstance(node, ast.BinOp):
op_type = type(node.op)
if op_type not in _SAFE_OPERATORS:
raise ValueError(f"Unsupported binary operator: {op_type.__name__}")
left = _safe_eval(node.left)
right = _safe_eval(node.right)
return _SAFE_OPERATORS[op_type](left, right)
if isinstance(node, ast.Call):
func_name = node.func.id
if func_name not in _SAFE_FUNCTIONS:
raise ValueError(f"Function '{func_name}' is not whitelisted.")
args = [_safe_eval(a) for a in node.args]
return _SAFE_FUNCTIONS[func_name](*args)
raise ValueError(f"Unsupported AST node type: {type(node).__name__}")
Sandboxed Python Code Executor (tools/code_executor.py)
When working through the Sandboxed Python Code Executor stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.
def _build_exec_namespace() -> dict[str, Any]:
return {
"__builtins__": _SAFE_BUILTINS, # Explicit whitelist (no open, __import__, eval)
"math": math,
"statistics": statistics,
"pi": math.pi,
"e": math.e,
}
def _execute_code(code: str, timeout: int = 5) -> ToolResult:
exec_result = _ExecResult()
captured_io = io.StringIO()
def _worker():
namespace = _build_exec_namespace()
namespace["__builtins__"]["print"] = lambda *args, **kw: print(*args, **{**kw, "file": captured_io})
try:
exec(code, namespace)
exec_result.stdout = captured_io.getvalue()
except Exception as exc:
exec_result.error = f"{type(exc).__name__}: {exc}"
thread = threading.Thread(target=_worker, daemon=True)
thread.start()
thread.join(timeout=timeout)
if thread.is_alive():
return ToolResult(success=False, error=f"Execution timed out after {timeout}s.")
Thread-Safe Streamlit UI Relay (app.py)
When working through the Thread-Safe Streamlit UI Relay stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Thread-Safe Streamlit UI Relay stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
class _PipelineRelay:
"""Thread-safe relay store for the running multi-agent pipeline."""
def __init__(self) -> None:
self.lock = threading.RLock()
self.status_messages: list[dict[str, str]] = []
self.agent_cards: dict[str, list[str]] = {}
self.running: bool = False
def append_status(self, msg_dict: dict[str, str]) -> None:
with self.lock:
self.status_messages.append(msg_dict)
agent_id = msg_dict.get("agent_id", "orchestrator")
self.agent_cards.setdefault(agent_id, []).append(msg_dict.get("detail", ""))
@st.cache_resource
def _get_relay() -> _PipelineRelay:
return _PipelineRelay()
Example Use Case: Climate Change Analysis
The Example Use Case Climate stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Benefits of Renewable Energy and Calculation of 280 Times 42
Benefits of Renewable Energy
Renewable energy reduces greenhouse gas emissions, contributing to climate change mitigation (Source: [National Renewable Energy Laboratory](https://www.nrel.gov/renewables/energy-benefits.html)). This is a key finding supported by credible sources, including the International Renewable Energy Agency (2020) and the World Health Organization (2020).
Renewable energy creates jobs and stimulates local economies (Source: [International Renewable Energy Agency](https://www.irena.org/publications/2020/Jun/Global-Status-Report-2020)). This is a significant benefit highlighted by the International Energy Agency (2020) and the National Bureau of Economic Research (2020).
Renewable energy improves air quality and public health (Source: [World Health Organization](https://www.who.int/news-room/fact-sheets/detail/air-pollution)). This is a critical aspect of renewable energy, supported by the National Renewable Energy Laboratory (n.d.).
The global renewable energy market is projected to reach 30% of total energy production by 2025 (Source: [International Energy Agency](https://www.iea.org/news/pressrelease/2020/june/global-renewables-report-2020/)). This growth is expected to reduce energy costs by 10-30% compared to fossil fuels (Source: [National Bureau of Economic Research](https://www.nber.org/papers/w28822)).
Calculation of 280 Times 42
The calculation of 280 times 42 yields 11,840. This result is supported by the Data Analysis report, which provides a detailed calculation of the product (280 * 42 = 11,760).
Trend Analysis
The global renewable energy market is expected to grow significantly, with a projected 30% share of total energy production by 2025. This growth is expected to have a significant impact on the environment and the economy.
Key Insights
1. Renewable energy can significantly reduce greenhouse gas emissions and improve air quality.
2. Renewable energy can create jobs and stimulate local economies.
Limitations
The data provided is based on projections and may not reflect actual outcomes.
Sources
1. National Renewable Energy Laboratory. (n.d.). Energy Benefits of Renewable Energy. Retrieved from <https://www.nrel.gov/renewables/energy-benefits.html>
2. International Renewable Energy Agency. (2020). Global Status Report 2020. Retrieved from <https://www.irena.org/publications/2020/Jun/Global-Status-Report-2020>
3. World Health Organization. (2020). Air pollution. Retrieved from <https://www.who.int/news-room/fact-sheets/detail/air-pollution>
4. International Energy Agency. (2020). Global Renewables Report 2020. Retrieved from <https://www.iea.org/news/pressrelease/2020/june/global-renewables-report-2020/>
5. National Bureau of Economic Research. (2020). The Economics of Renewable Energy. Retrieved from <https://www.nber.org/papers/w28822>
Adding a New Specialist Agent
The Adding a New Specialist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
from agents.base import BaseAgent
class MySpecialistAgent(BaseAgent):
def __init__(self, llm_client=None, on_status=None):
super().__init__(
agent_id="my_specialist",
tools=[my_custom_tool],
llm_client=llm_client,
on_status=on_status,
)
@property
def role_description(self) -> str:
return "You are a specialist that does X…"
Conclusion
The Conclusion stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Conclusion stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Links
For the Links stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Operational checklist
The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 60e98c93fde5: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.