Home / Articles / Why teams migrate from LangChain chains to LangGraph workflows

This article is published in English.

Why teams migrate from LangChain chains to LangGraph workflows

Production agents need durable state, branches, and HITL. Keep LangChain tools inside nodes; move control flow to an explicit graph when operability demands it.

1928 words

Teams that outgrow linear LangChain chains often land on LangGraph—not because chains are “dead,” but because production agents need durable state, cycles, and explicit control flow. The migration is driven by operational pain: retries, human gates, branching, and observability of partial progress.

What LangChain got right

LangChain made LLMs programmable with composable prompts, tools, retrievers, and LCEL piping. It normalized the idea that an application is a graph of model calls and data transforms, and it shipped batteries for RAG demos that still teach the ecosystem. For single-pass or lightly branching flows, it remains a productive toolkit.

Where it starts to hurt in production

Linear or ad-hoc agent executors blur when you need:

  • Loops that revisit retrieval after a failed grade
  • Pause/resume across process restarts
  • Per-user checkpoints
  • Conditional branches that are code, not prompt suggestions
  • Clear answers to “which step failed?”

At that point, stuffing more memory into a chain hides topology inside prompts. Failures become narrative instead of typed state transitions.

What LangGraph actually changes

LangGraph first-classes state, nodes, edges, and checkpoints. The runtime knows the cursor in the workflow. Interrupts, replays, and streaming of node events become natural. You still use LangChain components inside nodes; the orchestration layer changes.

The shape of the code, concretely

A chain-shaped mental model:

from langchain.agents import AgentExecutor, create_tool_calling_agent

agent = create_tool_calling_agent(llm, tools, prompt)
executor = AgentExecutor(agent=agent, tools=tools, verbose=True)
result = executor.invoke({"input": "Find the latest invoice and flag anomalies"})

A graph-shaped mental model with explicit branches and persistence:

from langgraph.graph import StateGraph, END

def call_model(state: AgentState) -> AgentState:
    response = llm.invoke(state["messages"])
    return {"messages": [response]}

def route(state: AgentState) -> str:
    last = state["messages"][-1]
    return "tools" if last.tool_calls else END

graph = StateGraph(AgentState)
graph.add_node("agent", call_model)
graph.add_node("tools", tool_node)
graph.add_conditional_edges("agent", route, {"tools": "tools", END: END})
graph.add_edge("tools", "agent")
app = graph.compile(checkpointer=checkpointer)

The second form makes “grade then maybe rewrite” a visible edge, not a paragraph in a mega-prompt.

Where CrewAI and Pydantic AI fit

CrewAI optimizes role/task collaboration when the metaphor is a crew, not a state machine. Pydantic AI and similar typed agent kits emphasize schema-first tool calling. They can coexist with LangGraph or replace it when the control-flow needs are lighter. Migration pressure toward LangGraph is strongest when durability and branching dominate—not when a short crew task list suffices.

The real pain points behind the migration

  1. Hidden control flow in agent executors.
  2. No first-class wait for humans without hacking global state.
  3. Retry storms that redo irreversible tool calls.
  4. Weak differentiation between “model error” and “business step skipped.”
  5. Testing that cannot target a single node with fixture state.

LangGraph does not magically fix bad tools, but it makes those pains addressable in code review.

When LangChain still makes sense

  • ETL-like prompt pipelines
  • Simple RAG without loops
  • Glue code inside graph nodes
  • Teaching and prototyping speed

Do not rewrite a working LCEL batch job into a graph for fashion.

When it is worth the rewrite

  • Multi-step agents with reflection
  • Compliance HITL
  • Long-running research or ops workflows
  • Need for time-travel debugging of state

Migration playbook

  1. Inventory chains and mark those with retries/branches/HITL.
  2. Extract shared state schema (TypedDict / Pydantic).
  3. Lift each chain segment into a node with clear inputs/outputs.
  4. Replace prompt-described branches with conditional edges.
  5. Add checkpointers before enabling pause/resume in production.
  6. Keep LangChain retrievers/tools inside nodes to avoid a big-bang rewrite of lean integrations.

Organizational effects

Graphs create a shared dialect between ML and platform engineers: nodes map to ownership, edges to SLAs. Incident response improves when pages cite node names. That clarity is a large part of “why migrate,” beyond any microbenchmark.

Costs and tradeoffs

Graphs add boilerplate and a learning curve. Over-fragmenting every helper into a node creates noise. Start coarse: retrieve → generate → grade → rewrite loop, then split nodes when metrics say so.

Closing

LangChain taught the industry how to wire models. LangGraph teaches how to run them as systems. The migration is less a repudiation than an admission that production agents are workflows—and workflows deserve state machines, not only strings of pipes.

Field notes from teams that moved

Expect parallel runs: keep the chain path for low-risk traffic while a percentage of sessions hit the graph. Compare tool-error rates, median steps to completion, and human-interrupt frequency. If the graph wins on operability even with similar answer quality, cut over. If not, the pain was elsewhere—usually tool design or evaluation—not the orchestrator brand.

Document anti-goals: LangGraph will not fix an index that cannot answer, nor a tool without idempotency keys. Pair the migration with contracts for tools and with offline eval sets that exercise branches you care about.

Interest metrics show builders moving toward graph- and crew-style kits as production agents demand durable state and branching—not merely longer chains. The first wave of LLM apps rewarded simple call-and-retrieve pipelines; the current wave rewards explicit workflows.

Patterns that show up in postmortems

When a chain-based agent fails in production, the write-up often sounds the same: the model “decided” to skip a verification tool, or a retry duplicated a side effect, or nobody could tell whether retrieval had run. Graphs do not eliminate those bugs, but they change the evidence available afterward. Node-level logs and checkpoint diffs show the last good state. That shortens mean time to understanding even when mean time to repair still depends on tool quality.

Designing state so migrations pay off

A useful state schema names business milestones explicitly: retrieved, drafted, graded, approved, committed. Edges move those flags; prompts do not invent them. During migration, map each old chain segment to a milestone. If a milestone cannot be named, the segment may not need its own node yet.

Human-in-the-loop without global mutable hacks

Chains often park HITL in external queues glued on with callbacks. LangGraph interrupts keep the wait inside the runtime: the checkpoint freezes, a UI collects approval, and resume continues with the same thread id. That design removes an entire class of “lost approval” bugs that plague hand-rolled wait tables.

Streaming and UX expectations

Users of agent products expect token streams and also step streams (“searching,” “grading,” “waiting for approval”). Graph event streams map cleanly to those step chips. Chains can fake it with custom callbacks, but the graph model matches the UX vocabulary already forming in the market.

Cost control

More nodes can mean more model calls. Budget max revisit counts on reflection loops. Cache retrieval results in state for the lifetime of a thread. Prefer cheap classifiers for routing nodes and reserve large models for synthesis. Migration is a chance to insert those controls deliberately rather than discovering them on the cloud bill.

Interop with existing LangChain investments

Retrievers, tool wrappers, output parsers, and prompt templates rarely need a rewrite. Nodes import them. The sunk cost argument against LangGraph usually collapses once teams see that migration is an orchestration transplant, not a ground-up rebuild. Where CrewAI or other kits already own a subsystem, wrap them as a single node rather than forcing a monoculture.

Decision matrix (condensed)

Signal Lean chain Lean graph
Single pass RAG yes optional
Reflection loops painful natural
HITL mid-flow bolted on native
Multi-day jobs awkward checkpoints
Simple ETL prompts ideal overkill

Story of a typical rewrite week

Day 1–2: draw the current chain as a whiteboard graph; name states. Day 3: implement the happy path with two conditional edges. Day 4: add checkpointer and interrupt on the dangerous tool. Day 5: shadow traffic and compare traces. Teams that skip the whiteboard step recreate chain spaghetti inside nodes and wonder why nothing improved.

What “done” means for migration

Migration is done when operators can answer, from tooling alone, which node last ran, what state keys changed, and how to replay from the prior checkpoint—without reading Slack archaeology. Answer quality may be unchanged on day one; operability should not be.

Concrete differences in error surfaces

Chain errors often surface as a single exception wrapping a model failure deep inside a runnable sequence. Graph errors can be attributed to a node name and to the state keys present when it failed. Support engineers use that attribution to decide whether to fix retrieval, grading prompts, or tool adapters. Over months, that difference dominates the migration ROI narrative more than any microbenchmark of tokens per second.

Versioning graphs

Treat compiled graph definitions as versioned artifacts. When node contracts change, bump a graph_version in checkpoint metadata and reject incompatible resumes. Without that discipline, pause/resume becomes a footgun during rolling deploys. Chains rarely faced this because they seldom paused mid-flight; graphs make the problem visible—and solvable.

Local development experience

LangGraph’s ability to step through nodes with fixture state improves PR review. Reviewers can run a single node with recorded inputs instead of replaying an entire chain. That workflow encourages smaller, testable nodes—the same pressure good service architecture already applies to HTTP handlers.

When not to fragment

If two “nodes” always run together with no branch between them, keep them as one node containing sequential LangChain calls. Graphs should encode decisions, not every function boundary. Over-fragmentation is the failure mode of enthusiastic migrations.

Ecosystem trajectory

As checkpointers, studio debuggers, and deployment helpers mature, the cost of choosing graphs early falls. Still, the strategic reason remains control-flow honesty: agents are workflows, and workflows deserve explicit state machines when production reliability matters.

Appendix: conversation starters for architecture review

Ask whether the current agent can pause for legal review without losing state; whether duplicate tool calls are prevented on retry; whether a new engineer can name the steps from a trace alone; and whether evaluation covers branch paths, not only the happy retrieval path. Negative answers are migration signals. Positive answers may mean LangChain-plus-discipline is already enough—and that is a valid outcome too.

Appendix: conversation starters for architecture review

Ask whether the current agent can pause for legal review without losing state; whether duplicate tool calls are prevented on retry; whether a new engineer can name the steps from a trace alone; and whether evaluation covers branch paths, not only the happy retrieval path. Negative answers are migration signals. Positive answers may mean LangChain-plus-discipline is already enough—and that is a valid outcome too.

Operability is the migration KPI that finance eventually notices.

Document the control-flow assumptions beside the code so future edits cannot silently delete an edge or a filter. Prefer machine-checkable assertions over tribal knowledge shared only in chat threads. Rehearse failure drills whenever topology or identity rules change. Keep evaluation sets versioned with the graph so regressions surface before customers do.