This article is published in English.
Practical notes: Build Your Own Claude Code Using Langchin: A Deepdive Into
Operable walkthrough of Practical notes: Build Your Own Claude Code Using Langchin: A Deepdive Into: contracts, checks, and drop-in code slots for teams shipping this pattern.
This walkthrough rebuilds the path from raw materials to a working system for: Build Your Own Claude Code Using Langchin: A Deepdive Into LangChain’s Deep Agents. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
The Architecture of a Coding Agent
When working through the The Architecture of a stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Part 0: The loop from scratch — no framework, no magic
When working through the Part 0 The loop stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
def run_agent_loop(client, user_message, tools, tool_functions, max_turns=20):
messages = [{"role": "user", "content": user_message}]
for _ in range(max_turns):
# 1. Ask the model what to do next.
response = client.messages.create(
model="claude-sonnet-4-6", max_tokens=1024, tools=tools, messages=messages
)
messages = [*messages, {"role": "assistant", "content": response.content}]
# 2. Plain text and no tool request? The job is done.
tool_uses = [b for b in response.content if b.type == "tool_use"]
if not tool_uses:
return "".join(b.text for b in response.content if b.type == "text")
# 3. Run each requested tool and hand the results back. 4. Repeat.
results = [
{"type": "tool_result", "tool_use_id": call.id,
"content": tool_functions[call.name](**call.input)}
for call in tool_uses
]
messages = [*messages, {"role": "user", "content": results}]
raise RuntimeError(f"agent loop did not finish within {max_turns} turns")
Part 1: The loop — the engine that drives everything
When working through the Part 1 The loop stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Part 1 The loop stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
pip install deepagents langchain-anthropic
from deepagents import create_deep_agent
def get_weather(city: str) -> str:
"""Get the weather for a given city."""
return f"It's always sunny in {city}!"
agent = create_deep_agent(
model="anthropic:claude-sonnet-4-6",
tools=[get_weather],
system_prompt="You are a helpful assistant.",
)
result = agent.invoke(
{"messages": [{"role": "user", "content": "What's the weather in San Francisco?"}]}
)
Part 2: The tools — giving the agent hands
The Part 2 The tools stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
Why not just let the model run any command?
The Why not just let stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
Building it with Deep Agents
The Building it with Deep stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Building it with Deep stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
import subprocess
from langchain_core.tools import tool
MAX_OUTPUT_CHARS = 20_000 # roughly 5k tokens
@tool
def run_tests(path: str = ".") -> str:
"""Run the project's pytest suite and return its output.
Output is truncated to the last 20k characters (failures appear at the
end). The run is killed after 5 minutes.
"""
try:
result = subprocess.run(["pytest", path], capture_output=True,
text=True, check=False, timeout=300)
except subprocess.TimeoutExpired:
return "pytest timed out after 300s"
output = result.stdout + result.stderr
if len(output) > MAX_OUTPUT_CHARS:
return ("[... output truncated to the last 20,000 characters ...]\n"
+ output[-MAX_OUTPUT_CHARS:])
return output
agent = create_deep_agent(
model="anthropic:claude-sonnet-4-6",
tools=[run_tests], # merged in alongside the built-ins
system_prompt="You are a coding assistant. Always run tests after editing.",
)
Part 3: Planning — thinking before doing
For the Part 3 Planning thinking stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Part 4: Context management — beating the memory limit
For the Part 4 Context management stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Building it with Deep Agents
For the Building it with Deep stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Building it with Deep stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Part 5: Subagents — divide and conquer
When working through the Part 5 Subagents divide stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
code_searcher = {
"name": "code-searcher",
"description": "Searches the codebase to find where specific logic lives. "
"Use this for any open-ended 'where is X?' question.",
# ⚠️ NOTE: the key is `system_prompt`, NOT `prompt` (see correction below).
"system_prompt": "You are an expert at navigating codebases. Use the grep "
"and glob tools to locate relevant files, then report a "
"concise summary of what you found and where. Do not make "
"any edits.",
}
agent = create_deep_agent(
model="anthropic:claude-sonnet-4-6",
tools=[run_tests],
system_prompt="You are a coding assistant.",
subagents=[code_searcher],
)
Part 6: Safety and human-in-the-loop — the brakes
When working through the Part 6 Safety and stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Building it with Deep Agents
When working through the Building it with Deep stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Building it with Deep stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
from deepagents import create_deep_agent
# Import path for backends can vary by version - check the current
# "Backends" page in the Deep Agents docs.
from deepagents.backends import LocalShellBackend
agent = create_deep_agent(
model="anthropic:claude-sonnet-4-6",
system_prompt="You are a coding assistant working inside this project.",
backend=LocalShellBackend(), # enables the `execute` shell tool
)
from langgraph.types import Command
result = agent.invoke({"messages": [...]}, config)
result["__interrupt__"] # the pending action + allowed decisions — show your user
# You decide; the loop picks up exactly where it paused:
agent.invoke(Command(resume={"decisions": [{"type": "approve"}]}), config)
# ... or {"type": "reject"} — the tool is never run, and the model is told so.
Part 7: Memory and persistence — remembering across sessions
The Part 7 Memory and stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Building it with Deep Agents
The Building it with Deep stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
from deepagents import create_deep_agent
from langgraph.checkpoint.memory import InMemorySaver
agent = create_deep_agent(
model="anthropic:claude-sonnet-4-6",
tools=[run_tests],
system_prompt="You are a coding assistant.",
checkpointer=InMemorySaver(), # remembers state within a session
)
# A "thread_id" ties messages together into one ongoing conversation.
config = {"configurable": {"thread_id": "project-alpha"}}
agent.invoke(
{"messages": [{"role": "user", "content": "Start refactoring the auth module."}]},
config=config,
)
# Later, same thread_id - the agent remembers the earlier turn:
agent.invoke(
{"messages": [{"role": "user", "content": "Now update the tests too."}]},
config=config,
)
Putting it all together
The Putting it all together stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Putting it all together stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
from deepagents import create_deep_agent
from deepagents.backends import LocalShellBackend
from langchain_core.tools import tool
from langgraph.checkpoint.memory import InMemorySaver
# --- A custom tool (Part 2) — token-budgeted, see Part 2 for the full body ---
@tool
def run_tests(path: str = ".") -> str:
"""Run the project's pytest suite and return its (truncated) output."""
import subprocess
try:
result = subprocess.run(["pytest", path], capture_output=True,
text=True, check=False, timeout=300)
except subprocess.TimeoutExpired:
return "pytest timed out after 300s"
output = result.stdout + result.stderr
return output if len(output) <= 20_000 else "[... truncated ...]\n" + output[-20_000:]
# --- A specialized subagent (Part 5) — note the `system_prompt` key ---
code_searcher = {
"name": "code-searcher",
"description": "Finds where specific logic lives in the codebase. "
"Use for open-ended 'where is X?' questions.",
"system_prompt": "You navigate codebases using grep and glob, then report a "
"concise summary of what you found. You never make edits.",
}
# --- A system prompt that teaches good behavior (Parts 2 & 3) ---
SYSTEM_PROMPT = """You are a careful coding assistant.
Workflow:
1. Plan the task as a to-do list before doing anything.
2. Use your built-in read, grep, and glob tools to explore - never the raw shell
equivalents like cat or grep.
3. Make focused edits.
4. ALWAYS run the tests after editing, and fix anything that breaks.
5. Delegate broad codebase searches to the code-searcher subagent.
"""
# --- Assemble the agent (Parts 1, 4, 6, 7 handled by the harness) ---
agent = create_deep_agent(
model="anthropic:claude-sonnet-4-6", # Part 1: the loop, model-agnostic
tools=[run_tests], # Part 2: custom hands
system_prompt=SYSTEM_PROMPT, # Parts 2 & 3: behavior + planning
subagents=[code_searcher], # Part 5: delegation
backend=LocalShellBackend(root_dir=".", virtual_mode=False), # Part 6: shell access
interrupt_on={"execute": True, # Part 6: the brakes —
"write_file": True, "edit_file": True}, # approval before anything destructive
checkpointer=InMemorySaver(), # Part 7: memory across turns
)
# Part 4 (context management) and built-in planning come on automatically.
config = {"configurable": {"thread_id": "my-project"}}
result = agent.invoke(
{"messages": [{"role": "user", "content": "The login tests are failing. Fix them."}]},
config=config,
)
print(result["messages"][-1].content)
The honest part: what’s easy and what’s hard
For the The honest part what stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Wrapping up
For the Wrapping up stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Operational checklist
The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.
Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 9ef98d98a69a: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.
For the hardening note 0 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 0/891: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 1 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 1/891: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 2 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 2/891: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.