This article is published in English.
Practical notes: Your AI Agent Isn’t Smart. Here’s How to Build One That
Operable walkthrough of Practical notes: Your AI Agent Isn’t Smart. Here’s How to Build One That: contracts, checks, and drop-in code slots for teams shipping this pattern.
Use this as an operator-facing rebuild of the ideas in “Your AI Agent Isn’t Smart. Here’s How to Build One That Actually Thinks”: clear stages, ordered code slots, and recovery notes that survive a handoff.
Table of Contents
The Table of Contents stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Why Most AI Agents Are Just Fancy Prompt Chains
The Why Most AI Agents stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
The Problem With “Just Use ReAct”
The The Problem With Just stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
The Architecture: Four Nodes, One Loop
The The Architecture Four Nodes stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
START --> Planner --> Executor <--> Replanner --> Reporter --> END
State Management: The Backbone of Everything
The State Management The Backbone stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The State Management The Backbone stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
import operator
from typing import Annotated, TypedDict
from pydantic import BaseModel, Field
class StrategyState(TypedDict, total=False):
"""Global state that flows through the LangGraph nodes."""
query: str
plan: list[dict]
scratchpad: Annotated[list[dict], operator.add]
current_step: int
final_report: str
replan_count: int
Structured Output Schemas: PlanStep and Plan
For the Structured Output Schemas PlanStep stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
AVAILABLE_TOOLS_TEXT = """
- get_metrics(ticker, metric?): Return stock metrics. 'metric' is optional
(P/E, EPS, Revenue, Market Cap, Sector).
- search_news(ticker): Return recent news headlines for a ticker.
- compare_metrics(tickers: list, metric): Compare one metric across
multiple tickers.
"""
class PlanStep(BaseModel):
"""A single executable step inside an analysis plan."""
step_id: int = Field(description="Sequential step number")
tool: str = Field(
description=f"Tool to use. Must be one of:\n{AVAILABLE_TOOLS_TEXT}"
)
args: dict = Field(description="Arguments for the tool call")
purpose: str = Field(description="Why this step is needed")
class Plan(BaseModel):
"""The full plan generated by the planner node."""
goal: str = Field(description="The overall analysis goal")
steps: list[PlanStep] = Field(
description="Ordered list of steps to execute"
)
class ReplanDecision(BaseModel):
"""Output of the replanner node."""
reasoning: str = Field(
description="Analysis of current progress and findings"
)
should_replan: bool = Field(
description="Whether the plan needs modification"
)
updated_steps: list[PlanStep] = Field(
default_factory=list,
description="Remaining steps if replan is needed. Empty if no changes.",
)
The Tool System: Three Tools, One Registry
For the The Tool System Three stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.
The ToolRegistry
For the The ToolRegistry stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the The ToolRegistry stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
from langchain_core.tools import BaseTool
from typing import Iterable, Mapping, Any
class ToolRegistry:
"""Namespace-aware container for LangChain tools."""
def __init__(self) -> None:
self._tools_by_toolset: dict[str, dict[str, BaseTool]] = {}
def add_tools(self, toolset: str, tools: Iterable[BaseTool]) -> None:
bucket = self._tools_by_toolset.setdefault(toolset, {})
bucket.update({t.name: t for t in tools})
def get_tools(self, toolset: str) -> tuple[BaseTool, ...]:
return tuple(self._tools_by_toolset.get(toolset, {}).values())
def invoke(
self, toolset: str, tool_name: str, tool_args: Mapping[str, Any]
) -> Any:
t = self._tools_by_toolset.get(toolset, {}).get(tool_name)
if t is None:
raise ValueError(
f"Unknown tool '{tool_name}' in toolset '{toolset}'"
)
return t.invoke(dict(tool_args))
The Planner Node: Think Before You Act
When working through the The Planner Node Think stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
def planner_node(state: StrategyState) -> dict:
"""Create a step-by-step research plan using structured output."""
planner = model.with_structured_output(Plan)
prompt = PLAN_PROMPT.format(
available_tools=AVAILABLE_TOOLS_TEXT,
ticker_choices=ticker_choices_text(),
metric_choices=metric_choices_text(),
query=state["query"],
)
plan: Plan = planner.invoke(prompt)
steps = [s.model_dump() for s in plan.steps]
return {"plan": steps, "current_step": 0}
PLAN_PROMPT: Where the Guardrails Live
When working through the PLANPROMPT Where the Guardrails stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
PLAN_PROMPT = """\
You are a financial research planner. Given a user's analysis request,
create a step-by-step research plan using the available tools.
Available tools:
{available_tools}
Rules:
- Use only the tools listed above.
- Every plan step must be executable with one of those tools.
- When a tool accepts 'ticker' or 'tickers', use only these exact values:
{ticker_choices}
- When a tool accepts 'metric', use one of these exact values:
{metric_choices}
- There are no other tools available. Final synthesis is handled separately.
Create an efficient plan. Group related lookups. Aim for 4-8 steps.
User request: {query}"""
The Executor Node: One Step at a Time
When working through the The Executor Node One stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the The Executor Node One stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
MAX_STEPS = 12 # Safety limit on total steps
def executor_node(state: StrategyState) -> dict:
"""Execute the next pending step from the plan."""
plan = state.get("plan", [])
current_step = state.get("current_step", 0)
if current_step >= len(plan):
return {}
if current_step >= MAX_STEPS:
return {"current_step": len(plan)}
step = plan[current_step]
tool_name = step["tool"]
tool_args = step["args"]
try:
result = str(
TOOL_REGISTRY.invoke(
AgentName.EXECUTOR.value, tool_name, tool_args
)
)
status = "Error" if result.startswith("Error:") else "Success"
except Exception as exc:
result = f"Error: {exc}"
status = "Error"
entry = {
"step": current_step + 1,
"tool": tool_name,
"args": tool_args,
"result": result,
"status": status,
}
return {
"scratchpad": [entry],
"current_step": current_step + 1,
}
The Replanner Node: Where Self-Correction Happens
The The Replanner Node Where stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
MAX_REPLANS = 2 # Prevent infinite replanning
def replanner_node(state: StrategyState) -> dict:
"""Review progress and optionally modify the remaining plan."""
plan = state.get("plan", [])
current_step = state.get("current_step", 0)
replan_count = state.get("replan_count", 0)
scratchpad = state.get("scratchpad", [])
remaining = plan[current_step:]
if len(remaining) = MAX_REPLANS:
return {}
scratchpad_text = "\n".join(
f"Step {e['step']}: {format_tool_call(e['tool'], e['args'])} "
f"-> [{e['status']}] {e['result'][:150]}..."
for e in scratchpad
)
remaining_text = "\n".join(
f"Step {s['step_id']}: {format_tool_call(s['tool'], s['args'])} "
f"- {s['purpose']}"
for s in remaining
)
replanner = model.with_structured_output(ReplanDecision)
prompt = REPLAN_PROMPT.format(
goal=state["query"],
scratchpad=scratchpad_text,
remaining_steps=remaining_text,
)
decision: ReplanDecision = replanner.invoke(prompt)
if decision.should_replan and decision.updated_steps:
new_steps = plan[:current_step] + [
s.model_dump() for s in decision.updated_steps
]
return {"plan": new_steps, "replan_count": replan_count + 1}
return {"replan_count": replan_count + 1}
REPLAN_PROMPT
The REPLANPROMPT stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
REPLAN_PROMPT = """\
You are a financial research planner reviewing progress on a research task.
Original goal: {goal}
Completed steps and findings so far:
{scratchpad}
Remaining steps in the plan:
{remaining_steps}
Based on the findings so far, should the remaining plan change?
If an expected tool failed or revealed something unexpected, add a step
to investigate.
If a step is now redundant, remove it.
Use only the available executable tools already shown in the plan.
Do not add recommendation, summary, or report-writing steps."""
Wiring the Graph: LangGraph Assembly
The Wiring the Graph LangGraph stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Wiring the Graph LangGraph stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
from langgraph.graph import StateGraph, END
from enum import Enum
class AgentName(Enum):
PLANNER = "planner"
EXECUTOR = "executor"
REPLANNER = "replanner"
REPORT = "report"
def build_graph():
"""Build and compile the LangGraph planning-agent workflow."""
workflow = StateGraph(StrategyState)
workflow.add_node(AgentName.PLANNER.value, planner_node)
workflow.add_node(AgentName.EXECUTOR.value, executor_node)
workflow.add_node(AgentName.REPLANNER.value, replanner_node)
workflow.add_node(AgentName.REPORT.value, report_node)
workflow.set_entry_point(AgentName.PLANNER.value)
workflow.add_edge(AgentName.PLANNER.value, AgentName.EXECUTOR.value)
workflow.add_edge(AgentName.EXECUTOR.value, AgentName.REPLANNER.value)
workflow.add_conditional_edges(
AgentName.REPLANNER.value,
should_continue_execution,
{
AgentName.EXECUTOR.value: AgentName.EXECUTOR.value,
AgentName.REPORT.value: AgentName.REPORT.value,
},
)
workflow.add_edge(AgentName.REPORT.value, END)
return workflow.compile()
def should_continue_execution(state: StrategyState) -> str:
"""Return the next node name after re-planning."""
if state.get("current_step", 0) >= len(state.get("plan", [])):
return AgentName.REPORT.value
return AgentName.EXECUTOR.value
Real Execution Example: Following One Query
For the Real Execution Example Following stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
{
"metric": "P/E",
"values": {
"NVDA": 58.3,
"AMD": 102.5
}
}
Where This Goes Next
For the Where This Goes Next stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Final Thoughts
For the Final Thoughts stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Final Thoughts stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Let’s Keep Learning Together
When working through the Let s Keep Learning stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
A message from our Founder
When working through the A message from our stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Operational checklist
When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.
Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.
Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for fea74fe7fb83: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.