Home / Articles / Practical notes: Multi-Agent Systems: When 2 Agents Beat 1 (and When They Don’t)

This article is published in English.

Practical notes: Multi-Agent Systems: When 2 Agents Beat 1 (and When They Don’t)

Operable walkthrough of Practical notes: Multi-Agent Systems: When 2 Agents Beat 1 (and When They Don’t): contracts, checks, and drop-in code slots for teams shipping this pattern.

3090 words

This walkthrough rebuilds the path from raw materials to a working system for: Multi-Agent Systems: When 2 Agents Beat 1 (and When They Don’t). The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

The Code Review Problem

When working through the The Code Review Problem stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

--- a/src/billing/invoice.py
+++ b/src/billing/invoice.py
@@ -42,7 +42,9 @@
 class InvoiceService:
-    def calculate_total(self, items):
-        return sum(i.price * i.qty for i in items)
+    def calculate_total(self, items, discount_pct=0):
+        subtotal = sum(i.price * i.qty for i in items)
+        return subtotal * (1 - discount_pct)

--- a/src/billing/api.py
+++ b/src/billing/api.py
@@ -18,6 +18,8 @@
 @router.post("/invoice")
 def create_invoice(req: InvoiceRequest):
+    discount = req.discount_pct   # NEW: from user input
     svc = InvoiceService()
-    total = svc.calculate_total(req.items)
+    total = svc.calculate_total(req.items, discount)
     return {"total": total}

--- a/config/feature_flags.yaml
+++ b/config/feature_flags.yaml
@@ -5,3 +5,4 @@
 flags:
   new_dashboard: true
+  discount_billing: true   # rollout: 100% immediately

Design 1: The Single-Agent Reviewer

When working through the Design 1 The Single-Agent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

from langchain_core.tools import tool
import textwrap

MOCK_DIFF = """...""" # The diff shown above
MOCK_FILES = {
    "src/billing/invoice.py": "class InvoiceService:\n    def calculate_total(self, items, discount_pct=0):\n        subtotal = sum(i.price * i.qty for i in items)\n        return subtotal * (1 - discount_pct)\n",
    "src/billing/api.py": '@router.post("/invoice")\ndef create_invoice(req: InvoiceRequest):\n    discount = req.discount_pct\n    svc = InvoiceService()\n    total = svc.calculate_total(req.items, discount)\n    return {"total": total}\n',
    "tests/test_billing.py": "def test_calculate_total():\n    # only tests no-discount path\n    assert svc.calculate_total(items) == 300\n",
}

@tool
def get_diff(pr_id: str) -> str:
    """Fetch the PR diff."""
    return MOCK_DIFF

@tool
def read_file(path: str) -> str:
    """Read a file from the repo."""
    return MOCK_FILES.get(path, f"FILE NOT FOUND: {path}")

@tool
def search_symbol(name: str) -> str:
    """Search for a symbol across the codebase."""
    if "discount" in name.lower():
        return "Found: InvoiceService.calculate_total(discount_pct) — src/billing/invoice.py:43"
    return f"No results for '{name}'"

@tool
def list_tests(path: str) -> str:
    """List test files covering a source path."""
    if "billing" in path:
        return "tests/test_billing.py — covers calculate_total (no-discount path only)"
    return "No tests found"

ALL_TOOLS = [get_diff, read_file, search_symbol, list_tests]
from typing import Literal
from typing_extensions import TypedDict
from pydantic import BaseModel
from langchain_google_genai import ChatGoogleGenerativeAI
from langchain_core.messages import AIMessage, HumanMessage, SystemMessage
from langgraph.graph import END, START, StateGraph

class ReviewFinding(BaseModel):
    severity: Literal["critical", "high", "medium", "low", "info"]
    category: Literal["bug", "regression", "missing_test", "rollout_risk",
                       "security", "edge_case", "style"]
    file: str
    description: str
    confidence: Literal["high", "medium", "low"]

class SingleAgentState(TypedDict):
    pr_id: str
    messages: list
    diff: str
    findings: list[ReviewFinding]

llm = ChatGoogleGenerativeAI(model="gemini-2.5-flash", temperature=0)
def sa_fetch(state: SingleAgentState) -> dict:
    diff = get_diff.invoke({"pr_id": state["pr_id"]})
    tests = list_tests.invoke({"path": "src/billing"})
    return {"diff": diff, "messages": [
        AIMessage(content=f"Diff loaded. Test coverage: {tests}")
    ]}

def sa_review(state: SingleAgentState) -> dict:
    """Single agent does BOTH jobs: summarize + critique in one pass."""
    prompt = f"""\
You are a senior code reviewer. Read this PR diff, summarize the change,
and produce a list of risk findings. Be specific.

DIFF:
{state['diff']}

Respond with JSON: {{"summary": "...", "findings": [
  {{"severity": "...", "category": "...", "file": "...",
    "description": "...", "confidence": "..."}}
]}}
"""
    resp = llm.invoke([SystemMessage(content="You are a code review bot."),
                       HumanMessage(content=prompt)])

    # In a real app we parse the JSON response here.
    # We simulate the typical single-agent output for this diff.
    findings = [
        ReviewFinding(
            severity="medium",
            category="missing_test",
            file="tests/test_billing.py",
            description="Add unit tests for the new discount_pct parameter.",
            confidence="high"
        ),
        ReviewFinding(
            severity="low",
            category="style",
            file="src/billing/invoice.py",
            description="Consider adding type hints to the items list.",
            confidence="high"
        )
    ]
    return {"findings": findings, "messages": [resp]}

def build_single_agent():
    g = StateGraph(SingleAgentState)
    g.add_node("fetch", sa_fetch)
    g.add_node("review", sa_review)
    g.add_edge(START, "fetch")
    g.add_edge("fetch", "review")
    g.add_edge("review", END)
    return g.compile()

The Handoff: Why Structured Schemas Matter

When working through the The Handoff Why Structured stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the The Handoff Why Structured stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

class ChangedInterface(BaseModel):
    file: str
    symbol: str
    change_type: Literal["added", "modified", "removed"]
    description: str

class ChangeModel(BaseModel):
    """Analyzer → Reviewer handoff. Structured, not prose."""
    purpose: str
    impacted_files: list[str]
    changed_interfaces: list[ChangedInterface]
    config_changes: list[str]
    migration_risk: bool
    assumptions: list[str]
    tests_touched: list[str]
    tests_likely_needed: list[str]

Design 2: The Two-Agent Architecture

The Design 2 The Two-Agent stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

class TwoAgentState(TypedDict):
    pr_id: str
    messages: list
    diff: str
    change_model: ChangeModel | None
    findings: list[ReviewFinding]
    conflict: bool
ANALYZER_PROMPT = """\
You are the Analyzer agent. Your ONLY job: understand the change.
Do NOT critique. Do NOT hunt for bugs. Just map what changed.
Produce a structured change model."""

def ta_analyzer(state: TwoAgentState) -> dict:
    """Analyzer: compress & clarify. Optimizes for coherence."""

    # We use LLM structured output to fill the ChangeModel.
    # For this demonstration, we simulate the accurate analysis.
    cm = ChangeModel(
        purpose="Add discount percentage support to invoice billing",
        impacted_files=["src/billing/invoice.py", "src/billing/api.py",
                        "config/feature_flags.yaml"],
        changed_interfaces=[
            ChangedInterface(file="src/billing/invoice.py",
                symbol="InvoiceService.calculate_total",
                change_type="modified",
                description="Added discount_pct param (default 0)"),
            ChangedInterface(file="src/billing/api.py",
                symbol="create_invoice",
                change_type="modified",
                description="Reads discount_pct from request, passes to service"),
        ],
        config_changes=["discount_billing flag added, 100% rollout"],
        migration_risk=False,
        assumptions=[
            "discount_pct is expected to be 0..1 (fraction, not percentage)",
            "No existing callers pass discount_pct yet",
            "Feature flag controls visibility, not the calculation",
        ],
        tests_touched=[],
        tests_likely_needed=[
            "test discount path in calculate_total",
            "test boundary: discount_pct = 0, 1, >1, <0",
            "test API validation of discount_pct input",
        ],
    )
    return {
        "change_model": cm,
        "messages": [AIMessage(content=f"Analyzer: change model built. "
                               f"{len(cm.changed_interfaces)} interfaces changed, "
                               f"{len(cm.assumptions)} assumptions made.")],
    }
REVIEWER_PROMPT = """\
You are the Risk Reviewer. The Analyzer gave you a change model.
DISTRUST it. Your job: find what's missing, broken, or dangerous.
Challenge every assumption. Check for missing tests, regressions,
rollout risks, and security issues."""

def ta_reviewer(state: TwoAgentState) -> dict:
    """Reviewer: expand & challenge. Optimizes for skepticism."""
    cm = state["change_model"]
    findings: list[ReviewFinding] = []

    # The Reviewer iterates through the Analyzer's assumptions
    for a in cm.assumptions:
        if "0..1" in a:
            findings.append(ReviewFinding(
                severity="critical",
                category="security",
                file="src/billing/api.py",
                description="ASSUMPTION CHALLENGED: discount_pct comes from user input "
                    "(req.discount_pct) with NO validation. Values <0 or >1 break billing. "
                    "Negative discount = price increase beyond subtotal. "
                    "Value >1 = negative total.",
                confidence="high"))

    # The Reviewer checks the test mapping
    if not cm.tests_touched and cm.tests_likely_needed:
        findings.append(ReviewFinding(
            severity="high",
            category="missing_test",
            file="tests/test_billing.py",
            description=f"NO tests touched but {len(cm.tests_likely_needed)} needed: "
                + "; ".join(cm.tests_likely_needed),
            confidence="high"))

    # The Reviewer checks the config changes
    for cc in cm.config_changes:
        if "100%" in cc:
            findings.append(ReviewFinding(
                severity="high",
                category="rollout_risk",
                file="config/feature_flags.yaml",
                description="Feature flag at 100% from day one — no gradual rollout. "
                    "Combined with unvalidated discount input, this is a billing incident risk.",
                confidence="high"))

    # The Reviewer looks for edge cases in the logic flow
    findings.append(ReviewFinding(
        severity="medium",
        category="edge_case",
        file="src/billing/api.py",
        description="Feature flag controls visibility but calculate_total always applies "
            "discount. If flag is off but API still receives discount_pct, discount "
            "is silently applied.",
        confidence="medium"))

    # We determine if there is a conflict worth escalating
    conflict = any(f.severity in ("critical", "high") for f in findings)

    return {
        "findings": findings,
        "conflict": conflict,
        "messages": [AIMessage(content=f"Reviewer: {len(findings)} findings, "
                               f"conflict={conflict}")],
    }

Arbitration and Human-in-the-Loop

The Arbitration and Human-in-the-Loop stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

def ta_merge(state: TwoAgentState) -> dict:
    """Merge findings. If conflict, surface for human review."""
    if state["conflict"]:
        return {
            "messages": [AIMessage(content=(
                "⚠ CONFLICT: Reviewer found critical/high issues. "
                "Routing to human review."
            ))],
        }
    return {
        "messages": [AIMessage(content="Findings merged. No escalation needed.")],
    }

def route_after_merge(state: TwoAgentState) -> str:
    return "escalate" if state["conflict"] else "emit"
from langgraph.types import Command, interrupt

def ta_escalate(state: TwoAgentState) -> Command:
    """Human-in-the-loop for conflicting findings."""

    # The interrupt function pauses execution and surfaces data to the caller
    decision = interrupt({
        "kind": "review_conflict",
        "pr_id": state["pr_id"],
        "change_model": state["change_model"].model_dump(),
        "findings": [f.model_dump() for f in state["findings"]],
        "prompt": "Review findings. Respond: "
                  '{"action":"accept_all"} or {"action":"override","drop_indices":[...]}',
    })

    # Execution resumes here when the human provides input
    action = decision.get("action", "accept_all")

    if action == "override":
        drop = set(decision.get("drop_indices", []))
        kept = [f for i, f in enumerate(state["findings"]) if i not in drop]

        # We use Command to update state and dynamically route to the next node
        return Command(update={"findings": kept}, goto="emit")

    return Command(goto="emit")
# How you resume the graph from your backend API
ta.invoke(Command(resume={"action": "accept_all"}), config)
def ta_emit(state: TwoAgentState) -> dict:
    """Final output: formatted review."""
    lines = [f"=== Code Review: PR {state['pr_id']} ==="]
    if state["change_model"]:
        lines.append(f"Purpose: {state['change_model'].purpose}")
    lines.append(f"Findings ({len(state['findings'])}):")

    for i, f in enumerate(state["findings"]):
        lines.append(f"  [{f.severity}] ({f.category}) {f.file}")
        lines.append(f"    {f.description}")

    return {"messages": [AIMessage(content="\n".join(lines))]}

Wiring the Two-Agent Graph

The Wiring the Two-Agent Graph stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Wiring the Two-Agent Graph stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

from langgraph.checkpoint.memory import MemorySaver

def build_two_agent():
    g = StateGraph(TwoAgentState)

    g.add_node("fetch", ta_fetch)
    g.add_node("analyzer", ta_analyzer)
    g.add_node("reviewer", ta_reviewer)
    g.add_node("merge", ta_merge)
    g.add_node("escalate", ta_escalate)
    g.add_node("emit", ta_emit)

    g.add_edge(START, "fetch")
    g.add_edge("fetch", "analyzer")
    g.add_edge("analyzer", "reviewer")
    g.add_edge("reviewer", "merge")

    g.add_conditional_edges("merge", route_after_merge,
        {"escalate": "escalate", "emit": "emit"})

    g.add_edge("emit", END)

    # A checkpointer is required to use interrupt()
    # Use MemorySaver for local testing, Postgres for production
    return g.compile(checkpointer=MemorySaver())

Running the Comparison

For the Running the Comparison stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

The Decision Matrix: When NOT to Use Multi-Agent

For the The Decision Matrix When stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

What’s Next

For the What s Next stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the What s Next stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Continue Reading

When working through the Continue Reading stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Operational checklist

The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.

Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.

Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for f4e352541695: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.

When working through the hardening note 0 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Hardening detail 0/727: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 1 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Hardening detail 1/727: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 2 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Hardening detail 2/727: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

When working through the hardening note 3 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Hardening detail 3/727: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 4 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Hardening detail 4/727: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.