Home / Articles / ReAct Agents in LangGraph: Thought, Action, Observation Step by Step

This article is published in English.

ReAct Agents in LangGraph: Thought, Action, Observation Step by Step

Implement the ReAct loop as explicit graph nodes with typed state, tool calls, and stop conditions you can test.

4143 words

This walkthrough rebuilds an operable path for: ReAct Agents Explained: A Step-by-Step Implementation Using LangGraph. Focus on contracts, checks, and code you can drop into a repo without guessing intent. For Overview, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Introduction

For Introduction, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate planning from tool execution. The planner proposes; the executor mutates; the verifier checks outcomes against the goal.

What Is a ReAct Agent?

For What Is a ReAct Agent?, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and cost next to functional results. Visibility early prevents surprise bills when the path moves from demo to shared environments. Separate planning from tool execution. The planner proposes; the executor mutates; the verifier checks outcomes against the goal.

Why ReAct Is Better Than Pure Chain-of-Thought

For Why ReAct Is Better Than Pure Chain-of-Thought, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate planning from tool execution. The planner proposes; the executor mutates; the verifier checks outcomes against the goal. For Why ReAct Is Better Than Pure Chain-of-Thought, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Thought: I don’t know the answer yet. I should search.
Action: Search("Paris weather this week")
Observation: It will rain on Thursday.
Thought: I should suggest indoor activities.
Final Answer: ...

ReAct Prompting

For ReAct Prompting, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Bound tool schemas tightly. Wide free-text args invite injection and make audits expensive.

Purpose of ReAct Prompting

For Purpose of ReAct Prompting, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and cost next to functional results. Visibility early prevents surprise bills when the path moves from demo to shared environments. Bound tool schemas tightly. Wide free-text args invite injection and make audits expensive.

Key Elements of ReAct Prompting

For Key Elements of ReAct Prompting, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Bound tool schemas tightly. Wide free-text args invite injection and make audits expensive. For Key Elements of ReAct Prompting, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

1. Chain-of-Thought Reasoning

For 1. Chain-of-Thought Reasoning, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive model calls so a retry does not re-bill the same work.

2. Explicit Action Space

For 2. Explicit Action Space, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and cost next to functional results. Visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive model calls so a retry does not re-bill the same work.

3. Observation Integration

For 3. Observation Integration, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive model calls so a retry does not re-bill the same work. For 3. Observation Integration, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

4. Iterative Looping

For 4. Iterative Looping, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate planning from tool execution. The planner proposes; the executor mutates; the verifier checks outcomes against the goal.

5. Final Answer Generation

For 5. Final Answer Generation, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and cost next to functional results. Visibility early prevents surprise bills when the path moves from demo to shared environments. Separate planning from tool execution. The planner proposes; the executor mutates; the verifier checks outcomes against the goal.

Canonical ReAct Prompt Structure

For Canonical ReAct Prompt Structure, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate planning from tool execution. The planner proposes; the executor mutates; the verifier checks outcomes against the goal. For Canonical ReAct Prompt Structure, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Question: <user question>Thought: <reason about what to do next>

Action: <selected tool>
Action Input: <tool input>
Observation: <tool output>
... (repeat as needed)
Thought: I now know the final answer
Final Answer: <answer to the user>

Zero-Shot ReAct Prompting

For Zero-Shot ReAct Prompting, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Bound tool schemas tightly. Wide free-text args invite injection and make audits expensive.

ReAct Prompting vs ReAct Agents

For ReAct Prompting vs ReAct Agents, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and cost next to functional results. Visibility early prevents surprise bills when the path moves from demo to shared environments. Bound tool schemas tightly. Wide free-text args invite injection and make audits expensive.

Why LangGraph for ReAct Agents?

For Why LangGraph for ReAct Agents?, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Bound tool schemas tightly. Wide free-text args invite injection and make audits expensive. For Why LangGraph for ReAct Agents?, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

The Core Problem: ReAct Is a State Machine, Not a Prompt

For The Core Problem: ReAct Is a State Machine, Not a Prompt, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive model calls so a retry does not re-bill the same work.

What Breaks Without LangGraph

For What Breaks Without LangGraph, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and cost next to functional results. Visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive model calls so a retry does not re-bill the same work.

1. Implicit Control Flow

For 1. Implicit Control Flow, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive model calls so a retry does not re-bill the same work. For 1. Implicit Control Flow, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

while True:
    llm_output = llm(prompt)
    if "Action:" in llm_output:
        tool_result = call_tool(...)
    else:
        break

2. Fragile State Management

For 2. Fragile State Management, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate planning from tool execution. The planner proposes; the executor mutates; the verifier checks outcomes against the goal.

3. No First-Class Loop Semantics

For 3. No First-Class Loop Semantics, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and cost next to functional results. Visibility early prevents surprise bills when the path moves from demo to shared environments. Separate planning from tool execution. The planner proposes; the executor mutates; the verifier checks outcomes against the goal.

4. Poor Production Readiness

For 4. Poor Production Readiness, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate planning from tool execution. The planner proposes; the executor mutates; the verifier checks outcomes against the goal. For 4. Poor Production Readiness, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Key Concepts in LangGraph (Agent-First View)

For Key Concepts in LangGraph (Agent-First View), define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Bound tool schemas tightly. Wide free-text args invite injection and make audits expensive.

1. State: The Agent’s Memory

For 1. State: The Agent’s Memory, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and cost next to functional results. Visibility early prevents surprise bills when the path moves from demo to shared environments. Bound tool schemas tightly. Wide free-text args invite injection and make audits expensive.

2. Nodes: Cognitive & Operational Units

For 2. Nodes: Cognitive & Operational Units, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Bound tool schemas tightly. Wide free-text args invite injection and make audits expensive. For 2. Nodes: Cognitive & Operational Units, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

3. Edges: Explicit Control Flow

For 3. Edges: Explicit Control Flow, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive model calls so a retry does not re-bill the same work.

4. Deterministic Execution with Flexibility

For 4. Deterministic Execution with Flexibility, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and cost next to functional results. Visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive model calls so a retry does not re-bill the same work.

ReAct + LangGraph: A Natural Fit

For ReAct + LangGraph: A Natural Fit, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive model calls so a retry does not re-bill the same work. For ReAct + LangGraph: A Natural Fit, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Use case: Hotel cancellation assistant (policy + refund calculation)

For Use case: Hotel cancellation assistant (policy + refund calculation), define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate planning from tool execution. The planner proposes; the executor mutates; the verifier checks outcomes against the goal.

Problem Statement

For Problem Statement, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Separate planning from tool execution. The planner proposes; the executor mutates; the verifier checks outcomes against the goal.

Step 1: Install dependencies

Separate planning from tool execution. The planner proposes; the executor mutates; the verifier checks outcomes against the goal.

pip install -U langgraph langchain langchain-openai
export OPENAI_API_KEY="..."

Step 2: Define tools (your “Actions”)

Bound tool schemas tightly. Wide free-text args invite injection and make audits expensive.

from typing import TypedDict, Annotated
from datetime import datetime
import json

from pydantic import BaseModel
from langchain_openai import ChatOpenAI
from langchain_core.messages import (
    BaseMessage,
    HumanMessage,
    ToolMessage,
    SystemMessage
)
from langchain_core.tools import tool
from langgraph.graph import StateGraph, START, END
from langgraph.graph.message import add_messages
from langgraph.prebuilt import tools_condition

@tool
def get_cancellation_policy(rate_plan: str) -> str:
    """
    Returns cancellation policy text for a given rate plan.
    """
    policies = {
        "flexible": "Free cancellation until 24 hours before check-in. After that, first night is charged.",
        "semi-flex": "Free cancellation until 72 hours before check-in. After that, 50% of the stay is charged.",
        "non-refundable": "No refund after booking. Full stay amount is charged on cancellation."
    }
    key = rate_plan.strip().lower()
    return policies.get(key, "Policy not found. Supported: flexible, semi-flex, non-refundable.")
@tool
def calculate_refund(
    rate_plan: str,
    check_in: str,
    cancel_date: str,
    nightly_rate: float,
    nights: int
) -> str:
    """
    Calculates refund amount based on a simplified policy model.
    Dates format: YYYY-MM-DD
    """
    rp = rate_plan.strip().lower()
    ci = datetime.strptime(check_in, "%Y-%m-%d").date()
    cd = datetime.strptime(cancel_date, "%Y-%m-%d").date()
    total = nightly_rate * nights
    days_before = (ci - cd).days
    if rp == "non-refundable":
        refund = 0.0
        charged = total
        rule = "Non-refundable: no refund."
    elif rp == "flexible":
        if days_before >= 1:
            refund = total
            charged = 0.0
            rule = "Flexible: cancelled >= 24h before check-in, full refund."
        else:
            charged = nightly_rate  # 1 night penalty
            refund = max(total - charged, 0.0)
            rule = "Flexible: late cancel, 1 night charged."
    elif rp == "semi-flex":
        if days_before >= 3:
            refund = total
            charged = 0.0
            rule = "Semi-flex: cancelled >= 72h before check-in, full refund."
        else:
            charged = 0.5 * total
            refund = total - charged
            rule = "Semi-flex: late cancel, 50% charged."
    else:
        return "Unsupported rate plan. Use: flexible, semi-flex, non-refundable."
    return (
        f"Rule: {rule}\n"
        f"Days before check-in: {days_before}\n"
        f"Total: ${total:.2f}\n"
        f"Charged: ${charged:.2f}\n"
        f"Refund: ${refund:.2f}"
    )

Step 3: Build a ReAct loop in LangGraph (Reason → Tool → Reason)

Bound tool schemas tightly. Wide free-text args invite injection and make audits expensive.

from typing import TypedDict, Annotated
from langchain_core.messages import BaseMessage, HumanMessage
from langgraph.graph import StateGraph, START, END
from langgraph.graph.message import add_messages

from langchain_openai import ChatOpenAI
from langchain_core.tools import Tool
from langgraph.prebuilt import ToolNode, tools_condition
# 1) Define state
class AgentState(TypedDict):
    messages: Annotated[list[BaseMessage], add_messages]
    booking_id: str
# 2) Define structured Output schema
class RefundDecision(BaseModel):
    booking_id: str
    rate_plan: str
    total_amount: float
    charged_amount: float
    refund_amount: float
    policy_summary: str
    explanation: str
# 2) Choose model and System prompt
llm = ChatOpenAI(model="gpt-4o-mini", temperature=0)
SYSTEM_PROMPT = SystemMessage(
    content="""
You are a hotel cancellation assistant.
Rules:
- Use tools when needed.
- Never guess policy or refund.
- Final answer MUST be valid JSON with this schema:
{
  "booking_id": "...",
  "rate_plan": "...",
  "total_amount": number,
  "charged_amount": number,
  "refund_amount": number,
  "policy_summary": "...",
  "explanation": "..."
}
"""
)
# 3) Register tools
tools = [get_cancellation_policy, calculate_refund]
def safe_tool_node(state):
    last_msg = state["messages"][-1]
    if not hasattr(last_msg, "tool_calls") or not last_msg.tool_calls:
        return {}
    tool_call = last_msg.tool_calls[0]
    tool_name = tool_call["name"]
    if tool_name not in ALLOWED_TOOLS:
        return {
            "messages": [
                ToolMessage(
                    content=f"Tool '{tool_name}' is not allowed.",
                    tool_call_id=tool_call["id"]
                )
            ]
        }
    for tool in tools:
        if tool.name == tool_name:
            result = tool.invoke(tool_call["args"])
            return {
                "messages": [
                    ToolMessage(
                        content=result,
                        tool_call_id=tool_call["id"]
                    )
                ]
            }
# 4)Before reasoning, check if we already processed this booking.
REFUND_MEMORY = {}
def memory_lookup_node(state):
    booking_id = state["booking_id"]
    if booking_id in REFUND_MEMORY:
        return {
            "messages": [
                HumanMessage(
                    content=f"Cached decision found:\n{REFUND_MEMORY[booking_id]}"
                )
            ]
        }
    return {}
# 5) Reasoning node: LLM decides next action (tool call) or final answer
def agent_node(state: AgentState):
    # Bind tools so the model can produce tool calls
    llm_with_tools = llm.bind_tools(tools)
    response = llm_with_tools.invoke(state["messages"])
    return {"messages": [response]}
# 6) After final decision, store it.
def memory_write_node(state):
    booking_id = state["booking_id"]
    final_answer = state["messages"][-1].content
    REFUND_MEMORY[booking_id] = final_answer
    return {}

Build and Compile the graph

Bound tool schemas tightly. Wide free-text args invite injection and make audits expensive.

# 7) Build the graph
builder = StateGraph(AgentState)

# Nodes
builder.add_node("memory_lookup", memory_lookup_node)
builder.add_node("agent", agent_node)
builder.add_node("tools", safe_tool_node)
builder.add_node("memory_write", memory_write_node)
# Flow
builder.add_edge(START, "memory_lookup")
builder.add_edge("memory_lookup", "agent")
builder.add_conditional_edges(
    "agent",
    tools_condition,
    {
        "tools": "tools",   # model wants to act
        END: "memory_write" # model finished reasoning
    }
)
builder.add_edge("tools", "agent")
builder.add_edge("memory_write", END)
graph = builder.compile()

Step 4: Run the agent on the use case

Bound tool schemas tightly. Wide free-text args invite injection and make audits expensive.

query = """
Booking details:
Rate plan: Non-Refundable
Check-in: 2026-01-20
Nights: 2
Nightly rate: 120
Cancelled on: 2026-01-18
"""

result = graph.invoke({
    "booking_id": "BKG-12345",
    "messages": [
        SYSTEM_PROMPT,
        HumanMessage(content=query)
    ]
})
final_output = result["messages"][-1].content
print(final_output)

Output

Bound tool schemas tightly. Wide free-text args invite injection and make audits expensive.

{
  "booking_id": "BKG-12345",
  "rate_plan": "Non-Refundable",
  "total_amount": 240.0,
  "charged_amount": 240.0,
  "refund_amount": 0.0,
  "policy_summary": "Non-refundable bookings do not allow refunds after confirmation.",
  "explanation": "The booking was made under a non-refundable rate plan, which charges the full stay amount regardless of cancellation timing."
}

What happens internally (ReAct behavior)

Bound tool schemas tightly. Wide free-text args invite injection and make audits expensive.

1) Thought (Reason)

Bound tool schemas tightly. Wide free-text args invite injection and make audits expensive.

2) Action (Tool call)

Bound tool schemas tightly. Wide free-text args invite injection and make audits expensive.

3) Observation (Tool outputs)

Checkpoint after expensive model calls so a retry does not re-bill the same work.

4) Final Answer

Checkpoint after expensive model calls so a retry does not re-bill the same work.

Why this is “ReAct” (not just tools)

Checkpoint after expensive model calls so a retry does not re-bill the same work.

Conclusion

Separate planning from tool execution. The planner proposes; the executor mutates; the verifier checks outcomes against the goal.

Operational checklist

Bound tool schemas tightly. Wide free-text args invite injection and make audits expensive.

Conditional edges should encode business rules as named functions, not buried prompt text.

Colocate types with components and keep props narrow. Wide prop bags become the debt that TypeScript was meant to prevent.

Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last change.