This article is published in English.
Practical notes: Building Multi-Agent System From Scratch — Part 6
Operable walkthrough of Practical notes: Building Multi-Agent System From Scratch — Part 6: contracts, checks, and drop-in code slots for teams shipping this pattern.
The following notes reconstruct a practical path around “Building Multi-Agent System From Scratch — Part 6: Observability and Debugging”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
What should one trace represent?
The What should one trace stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
blog-pipeline
├── research-agent
│ ├── search-web
│ └── research-model
├── writer-agent
│ └── writer-model
├── citation-check
│ └── citation-review-model
└── reviewer-agent
└── reviewer-model
Connect Langfuse to LangGraph
The Connect Langfuse to LangGraph stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
pip install -U langfuse
export LANGFUSE_PUBLIC_KEY="pk-lf-..."
export LANGFUSE_SECRET_KEY="sk-lf-..."
export LANGFUSE_BASE_URL="https://cloud.langfuse.com"
export LANGFUSE_TRACING_ENVIRONMENT="development"
from langfuse import get_client, propagate_attributes
from langfuse.langchain import CallbackHandler
langfuse = get_client()
langfuse_handler = CallbackHandler()
Trace one complete article run
The Trace one complete article stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Trace one complete article stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
def run_blog_pipeline(topic: str, blog_id: str, user_id: str):
initial_state = {
"topic": topic,
"audience": "developers new to agent systems",
"research_brief": "",
"sources": [],
"open_questions": [],
"article_draft": "",
"review_feedback": "",
"citation_issues": [],
"approved": False,
"revision_count": 0,
"status": "researching",
}
with langfuse.start_as_current_observation(
as_type="span",
name="blog-pipeline",
input={"topic": topic, "audience": initial_state["audience"]},
) as pipeline_span:
with propagate_attributes(
trace_name="blog-pipeline",
session_id=blog_id,
user_id=user_id,
tags=["blog-pipeline", "langgraph"],
version="1.0.0",
metadata={"workflow": "research-write-review"},
):
trace_id = langfuse.get_current_trace_id()
result = graph.invoke(
initial_state,
config={"callbacks": [langfuse_handler]},
)
pipeline_span.update(
output={
"status": result["status"],
"approved": result["approved"],
"revision_count": result["revision_count"],
}
)
return result, trace_id
Add observations that explain the handoff
For the Add observations that explain stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
from langchain_core.runnables import RunnableConfig
def research_node(state: BlogState, config: RunnableConfig) -> dict:
with langfuse.start_as_current_observation(
as_type="span",
name="research-agent",
input={"topic": state["topic"], "audience": state["audience"]},
) as span:
result = research_agent.invoke(
{"topic": state["topic"], "audience": state["audience"]},
config=config,
)
span.update(
output={
"source_count": len(result["sources"]),
"open_question_count": len(result["open_questions"]),
"brief": result["research_brief"],
}
)
return {
"research_brief": result["research_brief"],
"sources": result["sources"],
"open_questions": result["open_questions"],
"status": "writing",
}
Mark the signals that require attention
For the Mark the signals that stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
def record_source_assessment(assessment: SourceAssessment) -> None:
if assessment.suspicious_content:
langfuse.update_current_span(
level="WARNING",
status_message="Untrusted source contained agent-directed instructions.",
)
def record_pipeline_failure(error: Exception) -> None:
langfuse.update_current_span(
level="ERROR",
status_message=f"Pipeline failed: {type(error).__name__}",
)
Turn Reviewer decisions into scores
For the Turn Reviewer decisions into stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Turn Reviewer decisions into stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
result, trace_id = run_blog_pipeline(
topic="How AI agents use tools",
blog_id="blog-ai-tools-001",
user_id="philip",
)
if trace_id:
langfuse.create_score(
trace_id=trace_id,
name="review_approved",
value=1 if result["approved"] else 0,
data_type="BOOLEAN",
comment=result["status"],
)
langfuse.create_score(
trace_id=trace_id,
name="revision_count",
value=float(result["revision_count"]),
data_type="NUMERIC",
)
Debug a failed run in five questions
When working through the Debug a failed run stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Observability also has a privacy boundary
When working through the Observability also has a stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
What we built
When working through the What we built stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the What we built stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Operational checklist
For the Operational checklist stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.
Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for cf19385cb4a9: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.