This article is published in English.
Refund Agents: LangGraph Survived Where CrewAI and AutoGen Broke
Same tools and policy across three frameworks—only explicit state, idempotency, and checkpoints survived chaos kills.
The workload that breaks demo agents
Multi-agent frameworks look polished on research-and-blog demos. Nothing is at stake if a tool runs twice. A refund-processing agent is less forgiving. Given a customer message it must classify intent, load the order, evaluate policy (window, category, prior refunds), call the payment gateway at most once when eligible, escalate with a written rationale when not, and leave an audit trail.
That shape is common in production: decide, touch stateful systems, remain accountable. It demands durable state across steps, deterministic branching between refund and escalate, idempotent side effects, and human pauses that survive process restarts—not merely a sleep.
Three implementations shared the same tools, model, and policy logic. Only one survived chaos tests that killed the process mid-run.
Shared tools
Keep the fight fair: identical tool functions, including a flaky gateway and an idempotency key tied to the order id.
# tools.py — identical across all three implementations
import time
import uuid
from dataclasses import dataclass
from typing import Literal
class PaymentGatewayError(Exception):
pass
@dataclass
class Order:
order_id: str
customer_id: str
item_category: str
amount_cents: int
purchased_at: float
refund_count: int
# Fake DB — in prod this is Postgres behind a repository class
_ORDERS = {
"ORD-4471": Order("ORD-4471", "CUST-991", "electronics", 8999, time.time() - 86400 * 5, 0),
"ORD-2210": Order("ORD-2210", "CUST-102", "electronics", 4200, time.time() - 86400 * 45, 1),
}
_PROCESSED_REFUNDS: set[str] = set() # idempotency ledger
def get_order(order_id: str) -> Order | None:
return _ORDERS.get(order_id)
def check_refund_policy(order: Order) -> tuple[bool, str]:
days_since_purchase = (time.time() - order.purchased_at) / 86400
if days_since_purchase > 30:
return False, f"Purchase was {days_since_purchase:.0f} days ago, outside the 30-day window."
if order.refund_count >= 1:
return False, "Customer has already received a refund on this order."
return True, "Eligible: within window, no prior refund."
def issue_refund(order_id: str, idempotency_key: str) -> dict:
"""Calls the payment gateway. MUST be idempotent — retries are expected."""
if idempotency_key in _PROCESSED_REFUNDS:
return {"status": "already_processed", "idempotency_key": idempotency_key}
order = _ORDERS[order_id]
# simulate a flaky gateway — this matters later
if uuid.uuid4().int % 5 == 0:
raise PaymentGatewayError("gateway timeout, retry with same idempotency_key")
_PROCESSED_REFUNDS.add(idempotency_key)
return {"status": "refunded", "amount_cents": order.amount_cents, "idempotency_key": idempotency_key}
Idempotency is not decoration; it is the difference between an agent framework and an agent toy.
CrewAI: strong demos, weak control
CrewAI models roles and tasks under a Crew. Product decks love the org-chart metaphor.
Naive pattern
from crewai import Agent, Task, Crew, Process
from crewai.tools import tool
@tool("Get Order")
def get_order_tool(order_id: str) -> str:
"""Fetch order details by ID."""
order = get_order(order_id)
return str(order) if order else "NOT_FOUND"
@tool("Check Policy")
def check_policy_tool(order_id: str) -> str:
"""Check refund eligibility for an order."""
order = get_order(order_id)
if not order:
return "NOT_FOUND"
eligible, reason = check_refund_policy(order)
return f"eligible={eligible}, reason={reason}"
@tool("Issue Refund")
def issue_refund_tool(order_id: str) -> str:
"""Issue a refund for an order."""
result = issue_refund(order_id, idempotency_key=f"refund-{order_id}")
return str(result)
triage_agent = Agent(
role="Refund Triage Specialist",
goal="Decide whether a customer refund request should be approved or escalated",
backstory="You are an experienced support agent who follows policy strictly.",
tools=[get_order_tool, check_policy_tool, issue_refund_tool],
verbose=True,
)
triage_task = Task(
description="A customer says: '{customer_message}'. Order ID: {order_id}. "
"Decide if this qualifies for a refund and act accordingly.",
expected_output="A short summary of the action taken.",
agent=triage_agent,
)
crew = Crew(agents=[triage_agent], tasks=[triage_task], process=Process.sequential)
result = crew.kickoff(inputs={"customer_message": "I want a refund, item broke", "order_id": "ORD-4471"})
Happy paths look fine. As a service it failed three ways. Tool order was nondeterministic—sometimes issue_refund ran before check_policy because the model free-associates over tools. Retries after PaymentGatewayError could invent new tool calls and new keys unless the tool hardcodes idempotency. kickoff() runs to completion with no durable human pause; faking resume meant rebuilding state outside the framework.
Hardened hierarchical attempt
manager_agent = Agent(
role="Refund Process Manager",
goal="Enforce strict order: lookup, then policy check, then refund or escalate. Never skip steps.",
backstory="You strictly enforce process compliance and never let steps be skipped.",
allow_delegation=True,
)
crew = Crew(
agents=[triage_agent],
tasks=[triage_task],
process=Process.hierarchical,
manager_agent=manager_agent,
)
A manager agent and louder prompts reduced skipped steps but did not create hard invariants. Natural language cannot enforce compliance-grade sequencing. CrewAI fits flexible role play—research then critique—not payment-adjacent sequences.
AutoGen: flexible chat, fuzzy control
Group chat with automatic speaker selection adds an LLM call each turn to choose who speaks.
import autogen
config_list = [{"model": "gpt-4o", "api_key": "..."}]
llm_config = {"config_list": config_list, "temperature": 0}
triage_agent = autogen.AssistantAgent(
name="TriageAgent",
system_message=(
"You triage refund requests. Look up the order, check policy, "
"then either call issue_refund or hand off to EscalationAgent."
),
llm_config=llm_config,
)
escalation_agent = autogen.AssistantAgent(
name="EscalationAgent",
system_message="You write a human-readable escalation note explaining why a refund needs manual review.",
llm_config=llm_config,
)
user_proxy = autogen.UserProxyAgent(
name="ToolExecutor",
human_input_mode="NEVER",
code_execution_config=False,
function_map={
"get_order": lambda order_id: str(get_order(order_id)),
"check_refund_policy_tool": lambda order_id: str(check_refund_policy(get_order(order_id))),
"issue_refund": lambda order_id: str(issue_refund(order_id, f"refund-{order_id}")),
},
)
groupchat = autogen.GroupChat(
agents=[user_proxy, triage_agent, escalation_agent],
messages=[],
max_round=10,
speaker_selection_method="auto", # an LLM call decides who speaks next
)
manager = autogen.GroupChatManager(groupchat=groupchat, llm_config=llm_config)
user_proxy.initiate_chat(manager, message="Customer wants a refund on ORD-2210, item broke on arrival.")
Failures: conversational loops between triage and escalation with no structured “already decided” state—only transcripts—forcing blunt max_round cutoffs. “Did we refund?” lived in free text greps. Nondeterminism shaped the entire execution graph, so incidents were reconstructed from chat logs rather than typed traces.
LangGraph: the survivor
Typed state, explicit edges, checkpointers, and retries matched the problem.
from typing import TypedDict, Literal, Optional
from langgraph.graph import StateGraph, END
from langgraph.checkpoint.sqlite import SqliteSaver
from langgraph.types import interrupt, Command
import uuid
class RefundState(TypedDict):
order_id: str
customer_message: str
order: Optional[dict]
eligible: Optional[bool]
policy_reason: Optional[str]
decision: Optional[Literal["refund", "escalate", "denied"]]
refund_result: Optional[dict]
audit_log: list[str]
def lookup_order_node(state: RefundState) -> RefundState:
order = get_order(state["order_id"])
log = state["audit_log"] + [f"Looked up {state['order_id']}: {'found' if order else 'not found'}"]
if not order:
return {**state, "decision": "escalate", "audit_log": log}
return {**state, "order": order.__dict__, "audit_log": log}
def policy_check_node(state: RefundState) -> RefundState:
order = Order(**state["order"])
eligible, reason = check_refund_policy(order)
log = state["audit_log"] + [f"Policy check: eligible={eligible}, reason={reason}"]
return {**state, "eligible": eligible, "policy_reason": reason, "audit_log": log}
def route_after_policy(state: RefundState) -> str:
# Plain Python. No LLM call decides this branch. This is the whole point.
if state.get("decision") == "escalate":
return "escalate"
return "refund" if state["eligible"] else "escalate"
def human_approval_node(state: RefundState) -> RefundState:
# Durable pause: this literally suspends the graph run and persists state
# via the checkpointer. It can resume hours or days later, across restarts.
decision = interrupt({
"reason": "Ambiguous or ineligible refund needs human sign-off",
"order": state["order"],
"policy_reason": state["policy_reason"],
})
return {**state, "decision": decision, "audit_log": state["audit_log"] + [f"Human decision: {decision}"]}
def issue_refund_node(state: RefundState) -> RefundState:
idempotency_key = f"refund-{state['order_id']}" # stable across retries — this is the whole trick
try:
result = issue_refund(state["order_id"], idempotency_key)
except PaymentGatewayError as e:
# LangGraph re-raises into the node; retry policy (below) handles this,
# and because the key is stable, a retried call is safe.
raise
log = state["audit_log"] + [f"Refund issued: {result}"]
return {**state, "decision": "refund", "refund_result": result, "audit_log": log}
def escalate_node(state: RefundState) -> RefundState:
log = state["audit_log"] + ["Escalated to human queue"]
return {**state, "audit_log": log}
from langgraph.pregel.retry import RetryPolicy
graph = StateGraph(RefundState)
graph.add_node("lookup_order", lookup_order_node)
graph.add_node("policy_check", policy_check_node)
graph.add_node(
"issue_refund",
issue_refund_node,
retry=RetryPolicy(max_attempts=3, retry_on=PaymentGatewayError),
)
graph.add_node("human_approval", human_approval_node)
graph.add_node("escalate", escalate_node)
graph.set_entry_point("lookup_order")
graph.add_edge("lookup_order", "policy_check")
graph.add_conditional_edges("policy_check", route_after_policy, {
"refund": "issue_refund",
"escalate": "human_approval",
})
graph.add_edge("human_approval", "issue_refund") # human can still approve
graph.add_edge("issue_refund", END)
graph.add_edge("escalate", END)
checkpointer = SqliteSaver.from_conn_string("refunds.db")
app = graph.compile(checkpointer=checkpointer)
Thread configs resume after crashes:
config = {"configurable": {"thread_id": "order-2210-refund-req"}}
# Kick off the run — it will pause at human_approval_node
result = app.invoke(
{"order_id": "ORD-2210", "customer_message": "second refund please", "audit_log": []},
config=config,
)
# result contains an interrupt payload; the process can now exit entirely.
# ... hours later, possibly a different process, different machine ...
final_result = app.invoke(Command(resume="escalate"), config=config)
Deterministic tests assert ineligible orders escalate:
def test_ineligible_order_escalates():
state = {"eligible": False, "decision": None}
assert route_after_policy(state) == "escalate"
def test_eligible_order_refunds():
state = {"eligible": True, "decision": None}
assert route_after_policy(state) == "refund"
Chaos-killing the process mid-refund and restarting continued from the checkpoint with the same idempotency key. That test decided the bakeoff.
When CrewAI or AutoGen still win
CrewAI for collaborative drafting with flexible order. AutoGen for exploratory multi-agent research where conversation is the product. Neither replaces a state machine when money moves.
Lesson
Match abstraction to failure mode. If wrong tool order or double side effects are unacceptable, prefer explicit graphs with durable state over prompt-shaped crews. Frameworks are not interchangeable skins over “agents”; they encode different bets about control, memory, and recovery.
Production checklist after the bakeoff
Before promoting any refund-like agent, require: typed state with an already_refunded (or equivalent) flag; idempotency keys minted from business ids; policy checks as code nodes, not prompt suggestions; HITL interrupts behind a checkpointer; chaos tests that kill workers mid-side-effect; audit logs that do not depend on grepping prose. If a framework cannot express those properties without a shadow workflow engine beside it, the shadow engine is the real orchestrator—and the framework is expensive glue.
Measure distinct failure rates in staging: skipped policy, double refund attempts, lost interrupts after restart, and unreadable audit trails. Crew and AutoGen prototypes that cannot beat LangGraph on those metrics should stay in the lab. Celebrate retiring them; sprawl has a token bill.
Document the decision for future teams so the next bakeoff is unnecessary. Link the chaos harness in the README. Prefer boring graphs that survive kills over clever chats that only survive demos. That standard scales past refunds to any agent that mutates external systems under audit pressure, which is most of what enterprises actually want from “AI agents” once the slideware ends and the ledger begins to matter each quarter again for real.
Mapping failures to framework assumptions
CrewAI assumes flexible task decomposition. AutoGen assumes conversation is a sufficient control plane. LangGraph assumes you will draw the control plane yourself. Refund automation violates the first two assumptions: eligibility is not negotiable, and speaker selection is not a payment authorization mechanism. When assumptions clash with domain laws, the framework with fewer assumptions wins—even if it feels less magical in week one.
Engineers sometimes try to “fix” CrewAI or AutoGen with ever-longer system prompts. That is treating a cyclone fence like a vault door. Move compliance into typed nodes and keep language models inside nodes that classify or draft, not nodes that decide whether money moves.
Shared observability requirements
Whatever you pick, emit spans for each tool call with order id, idempotency key, and policy outcome. Without that, framework debates become religious while production stays blind. LangGraph made those hooks obvious because nodes are functions; you can add the same discipline elsewhere, but the bakeoff showed it was the default path of least resistance for this workload.
Closing
Build the agent that matches how your system fails. For refunds, that agent was a graph with memory, not a crew with vibes. Keep the other tools for problems they actually fit, and stop pretending one abstraction covers every agent-shaped ticket on the backlog.
Walkthrough of the LangGraph shape that held
The successful graph kept classification, fetch, policy, refund, escalate, and audit as separate nodes. Edges encoded the only legal transitions. The model never chose whether to skip policy; it only filled structured fields that policy code interpreted. Retries on the gateway node reused the same idempotency key stored in state. Escalation used interrupt() with a checkpointer so a manager could approve hours later on another replica.
That design feels verbose next to a three-agent Crew definition. Verbosity is the point: every irreversible step is nameable in a code review. New engineers can read the graph and predict behavior without replaying ten stochastic chats.
Comparing incident response stories
When CrewAI double-fired a tool in staging, the postmortem blamed “prompt drift.” When AutoGen looped, the postmortem blamed “speaker selection.” When LangGraph failed, the postmortem pointed at a specific node and a missing reducer—fixable without debating vibes. Incident taxonomy alone justified the choice for a payments-adjacent workflow.
Cost and latency notes from the bakeoff
AutoGen’s per-turn speaker LLM added latency and tokens. CrewAI’s retries sometimes multiplied tool calls. LangGraph paid a modest always-on cost for checkpoint writes and won on predictable spend. For support volumes in the thousands per day, predictability beats occasional clever savings.
Team skills and hiring
Hiring for “CrewAI experience” is thinner than hiring for “state machines plus LLMs.” LangGraph skills transfer to any explicit orchestrator. If the company standardizes, prefer transferable concepts: state, idempotency, HITL, evals. Framework fashion cycles faster than those concepts.
Expanding the chaos suite
Beyond kill-and-restart: inject gateway 503s, duplicate webhook deliveries, clock skew around policy windows, and humans who reject escalations. Graphs that survive only the happy path are still toys. Automate the suite in CI with deterministic fakes so refactors cannot silently drop safety edges.
Soft landing for prototypes
Keep CrewAI/AutoGen sandboxes for content workflows with human editors. Do not block experimentation—block production credentials. A platform policy that says “no payment tools in prompt-orchestrated crews” prevents the next near-miss without banning curiosity.
Final emphasis
The article’s claim is narrow and strong: for refund automation with real side effects, LangGraph survived where CrewAI and AutoGen did not, given identical tools. Extrapolate carefully. Extrapolate freely the meta-lesson—measure frameworks against your failure modes, not against demo aesthetics—and the bakeoff was worth the week it consumed.
Detailed failure timeline from staging
Week one: CrewAI demo impressed stakeholders. Week two: two double-refund attempts in staging after gateway flakes. Week three: AutoGen pilot burned tokens in speaker loops on denied orders. Week four: LangGraph chaos harness passed kill-mid-refund. The calendar matters because organizational momentum often freezes on the first demo; write the failure timeline into the ADR so momentum cannot erase evidence.
Idempotency as a cross-cutting requirement
Every framework tried here could call tools. Only designs that thread a stable key through retries are safe near payments. Store the key in graph state before the first attempt. Refuse to mint a new key on retry nodes. Log key, order id, and gateway response codes together. If a framework makes that awkward, that awkwardness is a signal, not a paperwork issue.
Human escalation ergonomics
Escalation text must include policy clause ids, order timestamps, and prior refund counts—structures from state, not free-form model prose alone. Managers should see the same fields the policy node saw. LangGraph made that easy because state is a TypedDict; Crew/AutoGen required reconstructing facts from transcripts, which is how audit gaps appear.
What we kept from the losers
CrewAI’s role metaphors helped product conversations—translate them into LangGraph node names. AutoGen’s explicit agent list inspired clearer ownership of tools per node. Stealing UX ideas while rejecting unsafe control planes is allowed.
Expanded production checklist
- Chaos: kill during tool, during interrupt wait, during retry backoff.
- Property tests: ineligible never reaches refund node.
- Load: burst 100 identical order ids; assert single gateway success.
- Security: payment tool credentials only on refund node workers.
- Observability: dashboards for skip-policy rate (should be zero).
- Governance: change control on policy node code equal to rules-engine changes.
Why “just add another agent” failed
Adding a “PolicyEnforcer” agent in CrewAI still left enforcement probabilistic. Adding a “RefundGuardian” in AutoGen still lived in chat. Guardians that cannot hard-block edges are decorations. Hard-block belongs in graph topology.
Closing reprise
Same tools, same model, different control philosophy. Only the explicit graph survived the failures that matter for refunds. Use crews and chats where wrong order is cheap; use graphs where wrong order is a ledger event. Publish that rule internally and save the next team a four-week bakeoff.
Additional hardening notes for refund graphs
Policy windows that depend on clocks must use server time stored at fetch, not model-stated dates. Category checks should be set membership in code. Prior refund counts come from the ledger, not from chat memory. Each of those choices removes a class of prompt injection that tries to rewrite eligibility in natural language. Graphs make those choices obvious; crews bury them in agent backstories where reviewers stop looking.