This article is published in English.
Practical notes: Context Engineering in Practice: Building a Production AI
Operable walkthrough of Practical notes: Context Engineering in Practice: Building a Production AI: contracts, checks, and drop-in code slots for teams shipping this pattern.
The following notes reconstruct a practical path around “Context Engineering in Practice: Building a Production AI Agent with the Claude Agent SDK”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Table of Contents:
The Table of Contents stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Want to Go Deeper into Context Engineering?
The Want to Go Deeper stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
1. What We Are Building
The 1 What We Are stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The 1 What We Are stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
# terminal
python3 -m venv .venv
source .venv/bin/activate
python -m pip install claude-agent-sdk==0.2.139
export ANTHROPIC_API_KEY="your-api-key"
# code/
research_agent/
config.py # naive and engineered ClaudeAgentOptions
hooks.py # pre-compaction checkpoint
metrics.py # message-stream and context measurements
runner.py # repeated runs and comparison
tools.py # in-process MCP tools
workspace.py # scratchpad and bounded retrieval
knowledge/
memory_approaches.json
tests/
CLAUDE.md
2. Establish the Naive Baseline
For the 2 Establish the Naive stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
# research_agent/minimal.py
import asyncio
from claude_agent_sdk import ClaudeAgentOptions, ResultMessage, query
QUESTION = "Compare approaches for long-term memory in production AI agents."
async def main() -> None:
options = ClaudeAgentOptions(
model="sonnet",
allowed_tools=["WebSearch", "WebFetch"],
permission_mode="dontAsk",
max_turns=20,
)
async for message in query(prompt=QUESTION, options=options):
if isinstance(message, ResultMessage):
print(message.result or "")
asyncio.run(main())
# research_agent/config.py
def naive_options(*, run_root, server, model, max_budget_usd):
tools = ["Read", "Glob", "Grep", "WebSearch", "WebFetch"]
return ClaudeAgentOptions(
cwd=run_root,
model=model,
tools=tools,
allowed_tools=[*tools, "mcp__research__*"],
permission_mode="dontAsk",
mcp_servers={"research": server},
strict_mcp_config=True,
setting_sources=[],
system_prompt={
"type": "preset",
"preset": "claude_code",
"append": NAIVE_PROMPT,
},
env={"ENABLE_TOOL_SEARCH": "false"},
max_turns=20,
max_budget_usd=max_budget_usd,
)
# research_agent/tools.py
@tool(
"load_knowledge_corpus",
"Return the entire local memory-research collection. Intended only for the naive baseline.",
{},
annotations=ToolAnnotations(readOnlyHint=True, openWorldHint=False),
)
async def load_knowledge(_: dict[str, Any]) -> dict[str, Any]:
return _text_result(load_corpus(corpus_path))
UserMessage(question)
AssistantMessage(ToolUseBlock: load_knowledge_corpus)
UserMessage(ToolResultBlock: entire corpus)
AssistantMessage(ToolUseBlock: WebSearch + WebFetch)
UserMessage(ToolResultBlock: raw search and page content)
AssistantMessage(final report)
# Captured output: python scripts/smoke_test.py
subtype=success
result=OPENROUTER_OK
models=claude-sonnet-5
# Captured output: python scripts/show_naive_trial.py ../measurements/comparison.json
Naive trial 1
SDK result: success
Assistant steps: 4
Tool calls: 6
WebFetch: 3
WebSearch: 1
mcp__research__load_knowledge_corpus: 1
mcp__research__search_knowledge: 1
Final active context: 16,478 tokens
Final tool-result payload: 6,851 tokens
Cumulative tree input: 138,182 tokens
Estimated cost: $0.754
3. Write: Move Working State Outside the Conversation
For the 3 Write Move Working stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
# research_agent/workspace.py
def initialize_workspace(root: Path, question: str) -> Path:
workspace = root / "workspace"
workspace.mkdir(parents=True, exist_ok=True)
(workspace / "artifacts").mkdir(exist_ok=True)
(workspace / "checkpoints").mkdir(exist_ok=True)
for name, template in WORKSPACE_FILES.items():
path = workspace / name
if not path.exists():
path.write_text(template.format(question=question), encoding="utf-8")
return workspace
4. Select: Assemble a Bounded Working Set
For the 4 Select Assemble a stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the 4 Select Assemble a stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
# research_agent/tools.py
@tool(
"search_knowledge",
"Search the local memory-research collection and return only the most relevant cited passages.",
{
"type": "object",
"properties": {
"query": {"type": "string", "minLength": 1},
"top_k": {"type": "integer", "minimum": 1, "maximum": 10},
},
"required": ["query", "top_k"],
"additionalProperties": False,
},
annotations=ToolAnnotations(readOnlyHint=True, openWorldHint=False),
)
async def search_knowledge(args: dict[str, Any]) -> dict[str, Any]:
try:
return _text_result(
search_corpus(corpus_path, args["query"], args["top_k"])
)
except (KeyError, TypeError, ValueError, sqlite3.Error) as error:
return _error_result(error)
5. Compress: Make Long Sessions Recoverable
When working through the 5 Compress Make Long stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
# research_agent/hooks.py
def build_precompact_hook(workspace: Path):
async def archive_before_compaction(
input_data: dict[str, Any],
tool_use_id: str | None,
context: Any,
) -> dict[str, Any]:
del tool_use_id, context
checkpoint_dir = workspace / "checkpoints"
checkpoint_dir.mkdir(parents=True, exist_ok=True)
timestamp = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%SZ")
session_id = input_data["session_id"]
trigger = input_data["trigger"]
stem = f"{timestamp}-{session_id}-{trigger}"
transcript = Path(input_data["transcript_path"])
metadata = {
"session_id": session_id,
"trigger": trigger,
"created_at": datetime.now(timezone.utc).isoformat(),
"source_transcript": str(transcript),
"custom_instructions": input_data.get("custom_instructions"),
"archived": transcript.is_file(),
}
if transcript.is_file():
shutil.copy2(transcript, checkpoint_dir / f"{stem}.jsonl")
(checkpoint_dir / f"{stem}.json").write_text(
json.dumps(metadata, indent=2) + "\n", encoding="utf-8"
)
return {}
return archive_before_compaction
6. Isolate: Delegate Focused Investigations
When working through the 6 Isolate Delegate Focused stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
# research_agent/config.py
"paper-researcher": AgentDefinition(
description="Analyzes primary research papers for memory mechanisms and trade-offs.",
prompt=(
"Investigate only the assigned paper question. Use primary sources. "
"Return at most five findings, each with a URL and an explicit limitation."
),
tools=["WebSearch", "WebFetch", "mcp__research__search_knowledge"],
model=model,
),
7. Isolate Heavy Tool Outputs in the Environment
When working through the 7 Isolate Heavy Tool stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours. When working through the 7 Isolate Heavy Tool stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
# research_agent/workspace.py (full function; use the Gist when publishing)
def materialize_source(corpus_path: Path, workspace: Path, source_id: str) -> dict:
document = next(
(item for item in load_corpus(corpus_path) if item["source_id"] == source_id),
None,
)
if document is None:
raise ValueError(f"unknown source_id: {source_id}")
path = safe_artifact_path(workspace, f"{source_id}.txt")
content = (
f"Title: {document['title']}\n"
f"URL: {document['url']}\n"
f"Published: {document['published']}\n\n"
f"{document['text']}\n"
)
path.write_text(content, encoding="utf-8")
return {
"path": str(path),
"characters": len(content),
"preview": content[:240],
}
# Captured output: python scripts/demonstrate_failure.py
Blocked artifact path: artifact name must use only letters, numbers, dots, underscores, or hyphens
8. Carry Context Across Sessions
The 8 Carry Context Across stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
# examples/session_modes.py
def session_options(session_id: str) -> dict[str, ClaudeAgentOptions]:
return {
"continue": ClaudeAgentOptions(continue_conversation=True),
"resume": ClaudeAgentOptions(resume=session_id),
"fork": ClaudeAgentOptions(resume=session_id, fork_session=True),
}
9. Assemble the Context-Engineered Agent
The 9 Assemble the Context-Engineered stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
# research_agent/config.py
tools = ["Read", "Write", "Edit", "Glob", "Grep", "WebSearch", "WebFetch", "Agent"]
return ClaudeAgentOptions(
cwd=run_root,
model=model,
tools=tools,
allowed_tools=[*tools, "mcp__research__*"],
permission_mode="dontAsk",
mcp_servers={"research": server},
strict_mcp_config=True,
setting_sources=["project"],
system_prompt={"type": "preset", "preset": "claude_code", "append": ENGINEERED_PROMPT},
env={"ENABLE_TOOL_SEARCH": "true"},
hooks={"PreCompact": [HookMatcher(hooks=[build_precompact_hook(workspace)])]},
agents=research_subagents(model),
max_turns=30,
max_budget_usd=max_budget_usd,
)
# research_agent/runner.py
while True:
async for message in client.receive_response():
metrics.observe(message)
report_path = workspace / "final_report.md"
if mode == "naive" or _report_meets_contract(report_path):
break
if metrics.completion_retries >= MAX_COMPLETION_RETRIES:
break
metrics.completion_retries += 1
await client.query(COMPLETION_REPAIR_PROMPT)
# Captured output: python -m unittest discover -s tests -q
----------------------------------------------------------------------
Ran 12 tests in 0.048s
OK
10. Compare the Two Architectures
The 10 Compare the Two stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The 10 Compare the Two stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
# research_agent/metrics.py
if isinstance(message, ResultMessage):
self.query_results += 1
self.session_id = message.session_id
self.result_subtype = message.subtype
result = message.result or ""
self.sdk_success = (
message.subtype == "success"
and bool(result.strip())
and "not logged in" not in result.lower()
)
self.estimated_cost_usd = (
(self.estimated_cost_usd or 0.0) + (message.total_cost_usd or 0.0)
)
self.result_usage = message.usage
self.model_usage = self._merge_model_usage(self.model_usage, message.model_usage)
self.total_tree_input_tokens = self._tree_input_tokens(self.model_usage)
if isinstance(message, SystemMessage) and message.subtype == "compact_boundary":
self.compactions += 1
# terminal
python scripts/show_comparison.py ../measurements/comparison.json
# Captured output: python scripts/show_comparison.py ../measurements/comparison.json
Measured comparison - 3 runs per architecture
Metric Naive Engineered
Artifact success 3/3 3/3
SDK success 3/3 1/3
Mean tree input 102,210 737,036
Mean peak context 17,184 32,668
Final tool-result tokens 7,351 6,202
Mean subagents 0 3
Mean estimated cost $0.663 $3.270
Compactions 0 0
Want to Go Deeper into Context Engineering?
For the Want to Go Deeper stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Operational checklist
For the Operational checklist stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.
Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 46aa5395a30a: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.
For the hardening note 0 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 0/807: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 1 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 1/807: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 2 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 2/807: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 3 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 3/807: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 4 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 4/807: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 5 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 5/807: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 6 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 6/807: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 7 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 7/807: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 8 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 8/807: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 9 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 9/807: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 10 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 10/807: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 11 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 11/807: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 12 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 12/807: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 0 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 0/826: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 1 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 1/826: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.