This article is published in English.
Stop Tools Conditionally: Artifacts Beat return_direct in LangGraph
Per-call stop/continue from tool artifacts—hand-built ReAct and create_agent middleware—when static return_direct cannot decide.
When static return_direct is the wrong tool
Anyone who has shipped a LangGraph tool-calling agent has met return_direct=True: skip sending the tool result back through the model and end the loop. It looks perfect until the stop decision must depend on this invocation’s outcome, not on which tool was registered.
This walkthrough hits that wall, builds a minimal ReAct loop by hand, then rebuilds the same behavior on create_agent with middleware. The short answer: yes, it works—but the first middleware instinct can fail for subtle message-order reasons, not because the framework silently drops updates.
Version pins for clarity: “legacy” means langgraph==0.6.6 (last line before create_react_agent gave way to create_agent); “current” means langchain==1.4.2 pulling langgraph==1.2.11. Every ReAct agent is a model ↔ tools loop; the focus here is the edge from tools back to the model—and when that edge should disappear for a given call.
The requirement that broke the default loop
Integrating a search tool looked ordinary: the model calls search(query), reads results, answers or continues. Two properties made the stock loop a poor fit.
On a hit, the tool returned a large JSON page. Stuffing that blob back into context for another model pass is expensive and usually pointless—if search already answered the question, a second call mostly rephrases at full cost.
On a miss, failures split into opposites: a real dead end (nothing to match—retrying is waste) versus a transient timeout or 503 (retry makes sense). So the rule is per call: success → stop, retryable failure → continue, fatal failure → stop—decided from the payload, not from the tool’s static type.
Why return_direct cannot say that
In langgraph.prebuilt.chat_agent_executor on the pinned legacy version, routing looks like:
should_return_direct = {t.name for t in tool_classes if t.return_direct}
...
def route_tool_responses(state):
for m in reversed(_get_state_value(state, "messages")):
if not isinstance(m, ToolMessage):
break
if m.name in should_return_direct:
return END
...
return entrypoint
should_return_direct is computed once from the tool’s .return_direct attribute at graph build time. The flag means “this tool always ends the loop.” It has no per-invocation mode. That is a category mismatch, not a defect in the flag.
Forum threads echo the same pain: bulky tool results force useless follow-up model calls, and maintainers often suggest wiring tool_node → END by hand. Separate discussion around Command updates plus return_direct (including langgraph#5496) is supporting context; the core argument does not need a bug—static flags simply cannot carry dynamic outcomes.
The control-flow shape
Stated plainly:
Tool call
├── success or unfixable failure → stop, use the tool's result
└── fixable failure → let the model decide
Two destinations, chosen fresh every call. The rest of this piece implements that shape twice—once by hand, once with middleware.
Split channels: content and artifact
LangChain’s @tool already separates what the model sees from what application code receives via response_format="content_and_artifact". The tool returns (content, artifact). ToolMessage.content goes to the model; artifact stays on the message for orchestration and never enters the LLM path.
@tool(response_format="content_and_artifact")
def search(query: str):
return "the content the LLM sees", {"stop": True, "debug": "extra stuff"}
node = ToolNode([search])
result = node.invoke(state)
msg = result["messages"][0]
# msg.content -> "the content the LLM sees"
# msg.artifact -> {"stop": True, "debug": "extra stuff"}
The desired pattern:
tool result
│
┌──────────┴──────────┐
↓ ↓
content artifact
│ │
↓ ↓
model router
│
continue / stop
versus what return_direct collapses into one static answer:
return_direct content_and_artifact
│ │
└── tool content → model
definition artifact → routing metadata
→ routing
One fixed flag trying to answer both “what does the user see?” and “should the loop end?”—or two channels, each answering one question.
Legacy hand-built ReAct
Rather than invent a loop, trim create_react_agent down to the skeleton: keep node names and the cycle, drop prompt hooks, structured response formats, dynamic model resolution, remaining-steps bookkeeping, checkpointing, interrupts, and parallel Send dispatch.
Three nodes remain:
agent— call the model; iftool_callsexist, continue, else finish.tools— plainToolNode; appendToolMessageresults.finalize— no model call; wrap the tool’s chosen final text as anAIMessageverbatim.
Two routers:
should_continueafteragent: tool calls →tools, elseEND.route_after_toolsaftertools: inspect the artifact and either return toagentor go tofinalize(replacing the original staticreturn_directset check).
finalize is a deliberate trade: it skips an LLM call and shows exactly what the tool produced, but the tool must emit presentable text and the model cannot fuse this result with other evidence. For “tool output is already the answer,” that trade wins.
Search outcomes as metadata, not graph commands
The search tool sets artifact["stop"] from what happened this call. stop is application metadata, not a LangChain-reserved field. Critically, the tool reports an outcome; orchestration interprets it. That keeps routing composable with policies the tool never sees.
@tool(response_format="content_and_artifact")
def search(query: str) -> tuple[str, dict]:
"""Search a knowledge base for information about the query."""
outcome = force_outcome or rng.choices(
list(resolved_weights), weights=list(resolved_weights.values())
)[0]
if outcome == "retryable":
return rng.choice(_RETRYABLE_MESSAGES), {"stop": False} if outcome == "fatal":
return rng.choice(_FATAL_MESSAGES), {"stop": True} query_lower = query.lower()
for topic, page in _INDEX.items():
if topic in query_lower or query_lower in topic:
return page, {"stop": True}
return "Nothing in the index overlaps with this query.", {"stop": True}
Three cases:
- Success with real content →
stop=True(another model pass would only rephrase). - Retryable failure →
stop=False(give the model another turn). - Fatal miss →
stop=True(looping burns tokens on the same non-answer).
stop=False is not “retry now”—it only withholds a hard stop. The model still chooses to re-call search, try something else, or answer. The router collapses to a one-line artifact check:
def route_after_tools(self, state: AgentState) -> str:
last_message = state["messages"][-1]
if (
isinstance(last_message, ToolMessage)
and isinstance(last_message.artifact, dict)
and last_message.artifact.get("stop")
):
return "finalize"
return "agent"
A harness that forces "success" | "retryable" | "fatal" makes paths deterministic against a real Groq model: success and fatal go tools → finalize → END with no second model call; retryable returns to agent.
Boundary: when routing also depends on remaining steps, prior attempt counts, or auth flags, artifact alone is insufficient—the router must read wider graph state. content_and_artifact shines when this tool’s result decides the next hop.
Beyond search
Any tool whose outcome is richer than ok/fail fits: a write_record tool can set already_applied; a poller can set progress for a UI the model never narrates. Artifact is plain data—usable from a conditional edge, middleware, or a UI that never touches the graph. return_direct is a routing decision baked into the definition; it has no “carry info, decide later” mode.
The same idea on create_agent
Pins: Python 3.12, langchain==1.4.2 / langgraph==1.2.11, langchain-groq==1.1.3. create_agent replaces the hand graph with declarative wiring plus middleware.
First instinct: wrap_tool_call, return Command(goto=END) when stop is set.
class StopOnArtifact(AgentMiddleware):
def wrap_tool_call(self, request, handler):
result = handler(request)
if isinstance(result, ToolMessage):
stop = isinstance(result.artifact, dict) and result.artifact.get("stop")
if stop:
relay = AIMessage(content=str(result.content))
return Command(goto=END, update={"messages": [result, relay]})
return Command(goto="model", update={"messages": [result]})
return result
In the tested build, that path only short-circuits when END is already reachable the way return_direct wires it. Middleware can compute stop=True while the loop still returns to the model until the model eventually answers without tools. That looked like #5496 live on current stacks—until two scripted variants showed otherwise:
A: update={"messages": [result]} -> stops correctly
B: update={"messages": [result, relay]} -> loops back to the model
Version A works. Version B adds a relay AIMessage without tool_calls in the same update. The exit check walks backward for the last AIMessage to evaluate return_direct; it finds the relay, sees no tool calls, and keeps looping. The Command was applied—the message ordering hid the original tool-calling message from the exit check. Not a dropped update, and not #5496.
Even after fixing that, the shipped approach uses before_model instead: it does not require return_direct at all.
class StopOnArtifact(AgentMiddleware):
@hook_config(can_jump_to=["end"])
def before_model(self, state, runtime):
last = state["messages"][-1]
if isinstance(last, ToolMessage) and isinstance(last.artifact, dict) and last.artifact.get("stop"):
relay = AIMessage(content=str(last.content))
return {"jump_to": "end", "messages": [relay]}
return None
before_model runs just before each model call—on later iterations, that is right after tools. @hook_config(can_jump_to=["end"]) permits jumping to END independent of any tool flag. Returning {"jump_to": "end", ...} is a plain state update the graph’s edge reads. One hook both detects the artifact and builds the relay AIMessage—the job legacy split across route_after_tools and finalize.
Forced outcomes match the hand-built graph: success and fatal short-circuit with verbatim tool content; retryable reopens the model turn.
Takeaway
content_and_artifact was not designed as a routing primitive. It separates audiences—model-visible content versus app-only metadata—and that same separation cleanly carries “should we stop?” without asking the model to reason about control flow. return_direct conflates presentation and termination into one static flag and fails exactly when those answers must disagree per call.
If a use case needs a conditional stop, keep the tool result and the routing decision apart: expose metadata beside the answer, let orchestration decide. content_and_artifact already provides that channel.
Design notes teams forget after the first green test
Conditional stopping looks solved once the three forced outcomes pass. Production adds concurrency: two tool calls in one model turn, or a batch of searches where only one should finalize. Decide whether any stop artifact short-circuits the whole step, whether all must agree, or whether a priority order applies. Encode that policy in the router, not in tribal knowledge.
Observability should show the artifact beside the ToolMessage without logging secrets from content. When a stop fires, record which rule matched—success, fatal, or policy override—so support can explain why the assistant did not “think longer.” Pair that with token accounting: the entire point of finalize-on-success is fewer model calls; dashboards should prove the savings.
Be careful porting patterns across LangGraph minors. Middleware hook names, Command reachability, and return-direct exit checks have shifted across the 0.6 → 1.x line. Keep a characterization test that forces success/retryable/fatal on every upgrade. If a hook suddenly loops forever, suspect message-list shape before filing framework bugs—relay messages are a recurring footgun.
Finally, resist stuffing control flags into content “just this once.” The moment the model sees stop=true in prose, it may narrate control flow or echo internal codes to users. Artifact exists so orchestration can be decisive while the user-facing channel stays clean.
Mapping the pattern onto neighboring frameworks
The same content-versus-control split appears outside LangGraph. Any agent runtime that merges tool stdout into the only message channel eventually invents ad-hoc markers, JSON envelopes, or sideband metadata. Prefer an official side channel when the platform offers one; invent a documented envelope when it does not; never rely on the model to ignore control tokens buried in prose.
If a team must support both legacy create_react_agent graphs and new create_agent apps, keep the tool’s artifact contract identical and swap only the router implementation. That isolates version churn to orchestration tests. When middleware grows—auth checks, spend caps, PII redaction—run those hooks before interpreting stop, so a policy denial cannot be mistaken for a successful short-circuit. Order of hooks is part of the public behavior of the agent, even if it feels like plumbing.
Document for future readers why finalize (or the before_model jump) exists: it is an explicit product choice that tool text may be user-visible without a polish pass. If product later wants a spoken summary style, reintroduce a model node on the stop path rather than overloading the tool to write two tones at once. Separating “compute result” from “narrate result” keeps tools reusable across voice, chat, and API clients.
Worked intuition for stop versus continue
Imagine a checkout tool that sometimes returns a completed receipt, sometimes a “payment processor timeout,” and sometimes “card permanently declined.” Those three map cleanly onto success stop, retryable continue, and fatal stop. The receipt’s content can be customer-ready HTML; the artifact carries { "stop": true, "reason": "completed" }. The timeout leaves a short explanation in content for the model and { "stop": false, "reason": "transient" } in artifact. The permanent decline stops the loop with a user-safe message and { "stop": true, "reason": "fatal" } so the agent does not hammer the processor. The same schema then generalizes to search, ticket creation, or document export without rewriting the router—only the tool’s mapping from raw API errors to the small reason vocabulary changes. Keeping that vocabulary tiny (completed / transient / fatal / policy_block) is what prevents artifact chaos as more tools adopt the pattern. Reviewers should reject one-off boolean names per tool when a shared enum already exists in the orchestration package.
Companion verification habits
Keep the forced-outcome harness in CI with a fake chat model that emits predetermined tool calls. Real Groq runs are for occasional end-to-end confidence, not for every commit. Assert exact path sequences: which nodes ran, whether a second model call occurred, and that finalize content equals tool content on stop paths. When someone “simplifies” middleware and reintroduces return_direct, the harness should fail loudly. Store golden transcripts beside the harness so failures are diffable. Conditional stopping is a behavioral contract; tests are how that contract survives refactors across LangGraph releases and across engineers who only skim the original design notes.
If product later needs the model to blend tool output with prior turns even on success, add an optional polish node after finalize rather than deleting the short-circuit. Feature flags beat rewrites: stop_mode=hard|polish|never lets experiments proceed without losing the artifact contract. Measure token spend under each mode on the same query set before choosing a default.
Reader’s contract for adopting the pattern
Copy the artifact schema and the router tests before copying the prose. The essay’s value is the separation of concerns, not the search tool anecdote. If your domain uses different failure labels, map them onto the same three buckets and keep the router boring. Resist adding a fourth bucket until a real incident demands it. When in doubt, prefer continuing to the model over hard-stopping on ambiguous errors—silent short-circuits that hide partial failures are worse than an extra inexpensive model call that explains uncertainty to the user.
Ship the harness with the article so readers can prove the edge cases on their own stack before trusting the pattern in production traffic.