This article is published in English.
Practical notes: A field guide to multi-agent architectures
Operable walkthrough of Practical notes: A field guide to multi-agent architectures: contracts, checks, and drop-in code slots for teams shipping this pattern.
This walkthrough rebuilds the path from raw materials to a working system for: A field guide to multi-agent architectures. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
1. Hierarchical Multi-Agent Systems
When working through the 1 Hierarchical Multi-Agent Systems stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
from fundamental_analysis_tools import evaluate_fundamentals
from langchain.agents import create_agent
from technical_analysis_tools import technical_analysis
agent = create_agent(
model=llm,
tools=[technical_analysis, evaluate_fundamentals]
)
Supervisor
├── Technical Analysis Tool
└── Fundamental Analysis Tool
Supervisor
├── Technical Analyst Agent
│ ├── Tool A
│ ├── Tool B
│ └── ...
└── Fundamental Analyst Agent
├── Tool C
├── Tool D
└── ...
from langchain.tools import tool
from langgraph_supervisor import create_supervisor
@tool
def get_weather(city: str) -> str:
"""Use this tool to get the weather of a city or location"""
return f"The weather is sunny in {city}"
weather_expert = create_agent(
model=llm,
tools=[get_weather],
name="weather_expert"
)
technical_analyst_agent = create_agent(
model=llm,
tools=[technical_analysis],
name="technical_analyst"
)
fundamental_analyst_agent = create_agent(
model=llm,
tools=[evaluate_fundamentals],
name="fundamental_analyst"
)
analysis_squad = create_supervisor(
[technical_analyst_agent, fundamental_analyst_agent],
model=llm,
supervisor_name="analysis_supervisor",
prompt=(
"You are a team supervisor managing a fundamental analyst and a "
"technical analyst..."
)
)
analysis_app = analysis_squad.compile(
name="fundamental_and_technical_analyst"
)
supervisor_graph = create_supervisor(
[analysis_app, weather_expert],
model=llm,
prompt=(
"You are a team supervisor managing a fundamental analyst, a "
"technical analyst and a weather expert..."
)
)
supervisor = supervisor_graph.compile()
When should you use a hierarchical architecture?
When working through the When should you use stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
2. Explicit Multi Agent Workflows
When working through the 2 Explicit Multi Agent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the 2 Explicit Multi Agent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
def execute_plan(
plan: Plan,
state: State,
) -> Command[
Literal[
"fundamental_analysis_agent",
"technical_analysis_agent",
"respond",
]
]:
gotos = [
Send(
step.action.agent_to_use,
{"query": step.action.query_to_send},
)
for step in plan.steps
]
if not gotos:
gotos.append(
Send("respond", {"messages": state["messages"]})
)
return Command(goto=gotos)
def router(state: State) -> Command[
Literal["fundamental_analysis_agent", "technical_analysis_agent", "respond"]
]:
# Get plan
query = state['messages'][-1].content
response = llm_planner.invoke(query) #invoke planner
return execute_plan(response, state)
When to use explicit multi agent workflows
The When to use explicit stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
3. Agent Swarm
The 3 Agent Swarm stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
from langgraph_swarm import (
create_handoff_tool,
create_swarm,
)
## We'll have to redefine our all agents to include the handoff tool
weather_expert = create_agent(
model=llm,
tools=[
get_weather,
create_handoff_tool(
agent_name="technical_analyst",
description="Transfer for technical analysis related questions"
),
create_handoff_tool(
agent_name="fundamental_analyst",
description="Transfer for fundamental analysis related questions"
),
],
name="weather_expert"
)
technical_analyst_agent = ...
fundamental_analyst_agent = ...
swarm_workflow = create_swarm(
[fundamental_analyst_agent, technical_analyst_agent, weather_expert],
default_active_agent="weather_expert" #a default agent must be specified.
)
swarm = swarm_workflow.compile()
When should you use a swarm?
The When should you use stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The When should you use stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
4. Blackboard Multi Agent Systems
For the 4 Blackboard Multi Agent stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
a. The blackboard
For the a The blackboard stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
class Blackboard(TypedDict, total=False):
query: str
# Problem frame
problem_framed: bool
ticker: str | None
location: str | None
wants_technical: bool
wants_fundamental: bool
# Specialist panels
technical: dict | None
fundamental: dict | None
environmental: dict | None
# Integrated solution
synthesis: str | None
# Control state
next_knowledge_source: str | None
cycles: int
b. Knowledge sources
For the b Knowledge sources stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the b Knowledge sources stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
def can_analyse_technicals(board: Blackboard) -> bool:
return (
board.get("ticker") is not None
and board.get("wants_technical", False)
and board.get("technical") is None
)
c. The control component
When working through the c The control component stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
def control_component(board: Blackboard) -> dict:
eligible = [
source
for source in KNOWLEDGE_SOURCES #these are subagents
if source.precondition(board)
]
if not eligible:
return {"next_knowledge_source": None}
chosen = max(
eligible,
key=lambda source: source.priority,
)
return {
"next_knowledge_source": chosen.name,
"cycles": board.get("cycles", 0) + 1,
}
inspect → identify eligible specialists → activate one
→ contribute → inspect again
Execution is deliberately sequential
When working through the Execution is deliberately sequential stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
When should you use a blackboard architecture?
When working through the When should you use stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the When should you use stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
5. Actor-Critic (Adversarial) Multi Agent Systems
The 5 Actor-Critic Adversarial Multi stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
actor → critic → judge
↑ |
└──── revise ─────┘
a. The Actor
The a The Actor stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
def actor(state: AdversarialState) -> dict:
if state.get("latest_critique") is None:
prompt = f"Produce a first draft.\n\nTASK: {state['task']}"
else:
prompt = (
"Revise the current draft.\n\n"
f"TASK:\n{state['task']}\n\n"
f"CURRENT DRAFT:\n{state['draft']}\n\n"
f"CRITIQUE:\n{state['latest_critique']}\n\n"
f"JUDGE'S PRIORITY:\n{state['latest_verdict']['focus']}"
)
draft = llm.invoke(prompt).content
return {
"draft": draft,
"round": state.get("round", 0) + 1,
}
b. The critic
The b The critic stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The b The critic stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
class Issue(BaseModel):
severity: Literal["blocking", "major", "minor"]
description: str
class Critique(BaseModel):
issues: list[Issue]
summary: str
def weighted_issue_score(issues: list[dict]) -> int:
weights = {
"blocking": 5,
"major": 2,
"minor": 1,
}
return sum(
weights[issue["severity"]]
for issue in issues
)
c. The judge
For the c The judge stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
class Verdict(BaseModel):
decision: Literal["accept", "revise"]
reasoning: str
focus: str
def judge(state: AdversarialState) -> dict:
verdict = llm.with_structured_output(Verdict).invoke(
[
HumanMessage(
content=(
"Evaluate the critic's findings on their merits. "
"Accept if only minor issues remain. "
"Request revision only for blocking or material problems.\n\n"
f"DRAFT:\n{state['draft']}\n\n"
f"CRITIQUE:\n{state['latest_critique']}"
)
)
]
)
return {
"latest_verdict": verdict.model_dump(),
}
Why a watchdog is still necessary
For the Why a watchdog is stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
round 1: 12
round 2: 8
round 3: 8
round 4: 9
def watchdog(state: AdversarialState) -> dict:
round_number = state.get("round", 0)
scores = state.get("issue_counts", [])
if round_number >= HARD_ROUND_CAP:
return {
"halt_reason": "hard round cap reached",
}
if len(scores) >= STALEMATE_WINDOW:
window = scores[-STALEMATE_WINDOW:]
if window[-1] >= window[0]:
return {
"halt_reason": (
f"stalemate detected: {window}"
),
}
return {}
Finalisation must distinguish success from exhaustion
For the Finalisation must distinguish success stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Finalisation must distinguish success stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
When should you use an adversarial architecture?
When working through the When should you use stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Choosing an architecture
When working through the Choosing an architecture stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Final Thoughts
When working through the Final Thoughts stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Final Thoughts stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Operational checklist
The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.
Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for f6f8c689c406: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.