首页 / 文章 / 实用笔记:多智能体系统——何时两个智能体会战胜一个(以及何时不会)

实用笔记:多智能体系统——何时两个智能体会战胜一个(以及何时不会)

《实用笔记:多智能体系统——何时两个智能体会战胜一个(以及何时不会)》的操作指南,为采用该设计模式的团队提供契约、校验规则及可直接插入的代码片段。

3090 词

本指南将逐步构建从原始材料到可运行系统的完整流程,主题为“多智能体系统:当两个智能体战胜一个时(以及无法战胜时)”。重点在于可操作的步骤、明确的检查点,以及可直接放入代码库的代码,无需猜测其用途。 在概览阶段,应在修改代码之前明确输入参数、各步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行相应步骤,而无需猜测隐藏状态。 配置信息应与应用程序代码分开存放。环境文件、密钥存储和功能开关应集中于一个位置,以便操作人员无需查看整个系统结构即可进行审核。

代码审查的问题

在处理“代码审查问题”阶段时,首先写下相关契约:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 同时记录正常流程和异常恢复流程。重试机制、人工审核环节以及错误消息处理都是产品本身的组成部分,而非后续需要补充的功能。 在成本较高的操作之后设置检查点。当操作员重新尝试某个节点时,恢复流程不应再次调用相同的大型语言模型。

--- a/src/billing/invoice.py
+++ b/src/billing/invoice.py
@@ -42,7 +42,9 @@
 class InvoiceService:
-    def calculate_total(self, items):
-        return sum(i.price * i.qty for i in items)
+    def calculate_total(self, items, discount_pct=0):
+        subtotal = sum(i.price * i.qty for i in items)
+        return subtotal * (1 - discount_pct)

--- a/src/billing/api.py
+++ b/src/billing/api.py
@@ -18,6 +18,8 @@
 @router.post("/invoice")
 def create_invoice(req: InvoiceRequest):
+    discount = req.discount_pct   # NEW: from user input
     svc = InvoiceService()
-    total = svc.calculate_total(req.items)
+    total = svc.calculate_total(req.items, discount)
     return {"total": total}

--- a/config/feature_flags.yaml
+++ b/config/feature_flags.yaml
@@ -5,3 +5,4 @@
 flags:
   new_dashboard: true
+  discount_billing: true   # rollout: 100% immediately

设计1:单智能体审查器

在处理“设计1:单智能体”阶段时,首先需写下相关契约:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 建议使用小型、可测试的单元,而非庞大的脚本。当某个步骤失败时,故障应指向单一责任模块,而非复杂的流程链。 在成本较高的步骤之后设置检查点。当操作员重新尝试后续节点时,恢复流程不应再次调用相同的大型语言模型。

from langchain_core.tools import tool
import textwrap

MOCK_DIFF = """...""" # The diff shown above
MOCK_FILES = {
    "src/billing/invoice.py": "class InvoiceService:\n    def calculate_total(self, items, discount_pct=0):\n        subtotal = sum(i.price * i.qty for i in items)\n        return subtotal * (1 - discount_pct)\n",
    "src/billing/api.py": '@router.post("/invoice")\ndef create_invoice(req: InvoiceRequest):\n    discount = req.discount_pct\n    svc = InvoiceService()\n    total = svc.calculate_total(req.items, discount)\n    return {"total": total}\n',
    "tests/test_billing.py": "def test_calculate_total():\n    # only tests no-discount path\n    assert svc.calculate_total(items) == 300\n",
}

@tool
def get_diff(pr_id: str) -> str:
    """Fetch the PR diff."""
    return MOCK_DIFF

@tool
def read_file(path: str) -> str:
    """Read a file from the repo."""
    return MOCK_FILES.get(path, f"FILE NOT FOUND: {path}")

@tool
def search_symbol(name: str) -> str:
    """Search for a symbol across the codebase."""
    if "discount" in name.lower():
        return "Found: InvoiceService.calculate_total(discount_pct) — src/billing/invoice.py:43"
    return f"No results for '{name}'"

@tool
def list_tests(path: str) -> str:
    """List test files covering a source path."""
    if "billing" in path:
        return "tests/test_billing.py — covers calculate_total (no-discount path only)"
    return "No tests found"

ALL_TOOLS = [get_diff, read_file, search_symbol, list_tests]
from typing import Literal
from typing_extensions import TypedDict
from pydantic import BaseModel
from langchain_google_genai import ChatGoogleGenerativeAI
from langchain_core.messages import AIMessage, HumanMessage, SystemMessage
from langgraph.graph import END, START, StateGraph

class ReviewFinding(BaseModel):
    severity: Literal["critical", "high", "medium", "low", "info"]
    category: Literal["bug", "regression", "missing_test", "rollout_risk",
                       "security", "edge_case", "style"]
    file: str
    description: str
    confidence: Literal["high", "medium", "low"]

class SingleAgentState(TypedDict):
    pr_id: str
    messages: list
    diff: str
    findings: list[ReviewFinding]

llm = ChatGoogleGenerativeAI(model="gemini-2.5-flash", temperature=0)
def sa_fetch(state: SingleAgentState) -> dict:
    diff = get_diff.invoke({"pr_id": state["pr_id"]})
    tests = list_tests.invoke({"path": "src/billing"})
    return {"diff": diff, "messages": [
        AIMessage(content=f"Diff loaded. Test coverage: {tests}")
    ]}

def sa_review(state: SingleAgentState) -> dict:
    """Single agent does BOTH jobs: summarize + critique in one pass."""
    prompt = f"""\
You are a senior code reviewer. Read this PR diff, summarize the change,
and produce a list of risk findings. Be specific.

DIFF:
{state['diff']}

Respond with JSON: {{"summary": "...", "findings": [
  {{"severity": "...", "category": "...", "file": "...",
    "description": "...", "confidence": "..."}}
]}}
"""
    resp = llm.invoke([SystemMessage(content="You are a code review bot."),
                       HumanMessage(content=prompt)])

    # In a real app we parse the JSON response here.
    # We simulate the typical single-agent output for this diff.
    findings = [
        ReviewFinding(
            severity="medium",
            category="missing_test",
            file="tests/test_billing.py",
            description="Add unit tests for the new discount_pct parameter.",
            confidence="high"
        ),
        ReviewFinding(
            severity="low",
            category="style",
            file="src/billing/invoice.py",
            description="Consider adding type hints to the items list.",
            confidence="high"
        )
    ]
    return {"findings": findings, "messages": [resp]}

def build_single_agent():
    g = StateGraph(SingleAgentState)
    g.add_node("fetch", sa_fetch)
    g.add_node("review", sa_review)
    g.add_edge(START, "fetch")
    g.add_edge("fetch", "review")
    g.add_edge("review", END)
    return g.compile()

交接流程:为何结构化方案至关重要

在处理“结构化交接原因”阶段时,首先需写下相关契约:所需输入、成功信号以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持透明可溯。 将此阶段视为输入与已验证输出之间的契约。为相关产物命名,明确成功判定标准,杜绝无声的半完成状态。 在成本较高的步骤之后设置检查点。当操作员重新尝试后续节点时,恢复流程不应再次调用相同的大型语言模型。 在处理“结构化交接原因”阶段时,首先需写下相关契约:所需输入、成功信号以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持透明可溯。 将配置信息置于应用程序代码之外。环境文件、密钥存储及功能开关应集中存放于一处,以便操作员无需查看整个流程即可进行审计。

class ChangedInterface(BaseModel):
    file: str
    symbol: str
    change_type: Literal["added", "modified", "removed"]
    description: str

class ChangeModel(BaseModel):
    """Analyzer → Reviewer handoff. Structured, not prose."""
    purpose: str
    impacted_files: list[str]
    changed_interfaces: list[ChangedInterface]
    config_changes: list[str]
    migration_risk: bool
    assumptions: list[str]
    tests_touched: list[str]
    tests_likely_needed: list[str]

设计2:双智能体架构

在设计2的双智能体阶段中,若能将其视为可度量的结构,则效果最佳。在扩大范围之前,需记录一个成功的用例、一个失败案例以及回滚说明。 同时记录正常流程与恢复流程。重试机制、人工审核环节以及死信处理都是产品本身的组成部分,而非后续需要补充的功能。 保持图结构的扁平化与类型化。嵌套的数据块会掩盖哪个节点编写了哪个字段的信息,且在中断后会导致流程无法继续。

class TwoAgentState(TypedDict):
    pr_id: str
    messages: list
    diff: str
    change_model: ChangeModel | None
    findings: list[ReviewFinding]
    conflict: bool
ANALYZER_PROMPT = """\
You are the Analyzer agent. Your ONLY job: understand the change.
Do NOT critique. Do NOT hunt for bugs. Just map what changed.
Produce a structured change model."""

def ta_analyzer(state: TwoAgentState) -> dict:
    """Analyzer: compress & clarify. Optimizes for coherence."""

    # We use LLM structured output to fill the ChangeModel.
    # For this demonstration, we simulate the accurate analysis.
    cm = ChangeModel(
        purpose="Add discount percentage support to invoice billing",
        impacted_files=["src/billing/invoice.py", "src/billing/api.py",
                        "config/feature_flags.yaml"],
        changed_interfaces=[
            ChangedInterface(file="src/billing/invoice.py",
                symbol="InvoiceService.calculate_total",
                change_type="modified",
                description="Added discount_pct param (default 0)"),
            ChangedInterface(file="src/billing/api.py",
                symbol="create_invoice",
                change_type="modified",
                description="Reads discount_pct from request, passes to service"),
        ],
        config_changes=["discount_billing flag added, 100% rollout"],
        migration_risk=False,
        assumptions=[
            "discount_pct is expected to be 0..1 (fraction, not percentage)",
            "No existing callers pass discount_pct yet",
            "Feature flag controls visibility, not the calculation",
        ],
        tests_touched=[],
        tests_likely_needed=[
            "test discount path in calculate_total",
            "test boundary: discount_pct = 0, 1, >1, <0",
            "test API validation of discount_pct input",
        ],
    )
    return {
        "change_model": cm,
        "messages": [AIMessage(content=f"Analyzer: change model built. "
                               f"{len(cm.changed_interfaces)} interfaces changed, "
                               f"{len(cm.assumptions)} assumptions made.")],
    }
REVIEWER_PROMPT = """\
You are the Risk Reviewer. The Analyzer gave you a change model.
DISTRUST it. Your job: find what's missing, broken, or dangerous.
Challenge every assumption. Check for missing tests, regressions,
rollout risks, and security issues."""

def ta_reviewer(state: TwoAgentState) -> dict:
    """Reviewer: expand & challenge. Optimizes for skepticism."""
    cm = state["change_model"]
    findings: list[ReviewFinding] = []

    # The Reviewer iterates through the Analyzer's assumptions
    for a in cm.assumptions:
        if "0..1" in a:
            findings.append(ReviewFinding(
                severity="critical",
                category="security",
                file="src/billing/api.py",
                description="ASSUMPTION CHALLENGED: discount_pct comes from user input "
                    "(req.discount_pct) with NO validation. Values <0 or >1 break billing. "
                    "Negative discount = price increase beyond subtotal. "
                    "Value >1 = negative total.",
                confidence="high"))

    # The Reviewer checks the test mapping
    if not cm.tests_touched and cm.tests_likely_needed:
        findings.append(ReviewFinding(
            severity="high",
            category="missing_test",
            file="tests/test_billing.py",
            description=f"NO tests touched but {len(cm.tests_likely_needed)} needed: "
                + "; ".join(cm.tests_likely_needed),
            confidence="high"))

    # The Reviewer checks the config changes
    for cc in cm.config_changes:
        if "100%" in cc:
            findings.append(ReviewFinding(
                severity="high",
                category="rollout_risk",
                file="config/feature_flags.yaml",
                description="Feature flag at 100% from day one — no gradual rollout. "
                    "Combined with unvalidated discount input, this is a billing incident risk.",
                confidence="high"))

    # The Reviewer looks for edge cases in the logic flow
    findings.append(ReviewFinding(
        severity="medium",
        category="edge_case",
        file="src/billing/api.py",
        description="Feature flag controls visibility but calculate_total always applies "
            "discount. If flag is off but API still receives discount_pct, discount "
            "is silently applied.",
        confidence="medium"))

    # We determine if there is a conflict worth escalating
    conflict = any(f.severity in ("critical", "high") for f in findings)

    return {
        "findings": findings,
        "conflict": conflict,
        "messages": [AIMessage(content=f"Reviewer: {len(findings)} findings, "
                               f"conflict={conflict}")],
    }

仲裁机制与人工干预

将仲裁与人工干预阶段视为可度量的对象来处理,效果最佳。在扩大范围之前,先记录一份完美的操作日志、一个故障案例以及回滚说明。 优先选择小型且可测试的单元,而非庞大的脚本。当某一步骤出现故障时,故障应指向单一责任主体,而非复杂的流程链。 保持图结构简洁且类型明确。嵌套的数据块会掩盖哪个节点编写了哪个字段的信息,还会在中断后导致流程无法继续。

def ta_merge(state: TwoAgentState) -> dict:
    """Merge findings. If conflict, surface for human review."""
    if state["conflict"]:
        return {
            "messages": [AIMessage(content=(
                "⚠ CONFLICT: Reviewer found critical/high issues. "
                "Routing to human review."
            ))],
        }
    return {
        "messages": [AIMessage(content="Findings merged. No escalation needed.")],
    }

def route_after_merge(state: TwoAgentState) -> str:
    return "escalate" if state["conflict"] else "emit"
from langgraph.types import Command, interrupt

def ta_escalate(state: TwoAgentState) -> Command:
    """Human-in-the-loop for conflicting findings."""

    # The interrupt function pauses execution and surfaces data to the caller
    decision = interrupt({
        "kind": "review_conflict",
        "pr_id": state["pr_id"],
        "change_model": state["change_model"].model_dump(),
        "findings": [f.model_dump() for f in state["findings"]],
        "prompt": "Review findings. Respond: "
                  '{"action":"accept_all"} or {"action":"override","drop_indices":[...]}',
    })

    # Execution resumes here when the human provides input
    action = decision.get("action", "accept_all")

    if action == "override":
        drop = set(decision.get("drop_indices", []))
        kept = [f for i, f in enumerate(state["findings"]) if i not in drop]

        # We use Command to update state and dynamically route to the next node
        return Command(update={"findings": kept}, goto="emit")

    return Command(goto="emit")
# How you resume the graph from your backend API
ta.invoke(Command(resume={"action": "accept_all"}), config)
def ta_emit(state: TwoAgentState) -> dict:
    """Final output: formatted review."""
    lines = [f"=== Code Review: PR {state['pr_id']} ==="]
    if state["change_model"]:
        lines.append(f"Purpose: {state['change_model'].purpose}")
    lines.append(f"Findings ({len(state['findings'])}):")

    for i, f in enumerate(state["findings"]):
        lines.append(f"  [{f.severity}] ({f.category}) {f.file}")
        lines.append(f"    {f.description}")

    return {"messages": [AIMessage(content="\n".join(lines))]}

连接双智能体图结构

将“连接两个智能体图”阶段视为可度量的对象来处理效果最佳。在扩大范围之前,需记录一份理想状态下的完整流程、一个故障案例以及回滚说明。 应将此阶段视为输入与已验证输出之间的契约。为相关成果命名,明确成功标准,绝不允许出现无声的半完成状态。 保持图结构的扁平化与类型化。嵌套的数据块会掩盖哪个节点修改了哪个字段,还会在中断后导致无法继续处理。 将“连接两个智能体图”阶段视为可度量的对象来处理效果最佳。在扩大范围之前,需记录一份理想状态下的完整流程、一个故障案例以及回滚说明。 将配置信息置于应用程序代码之外。环境文件、密钥存储和功能开关应集中存放于一个位置,以便操作人员无需查看整个图结构即可进行审计。

from langgraph.checkpoint.memory import MemorySaver

def build_two_agent():
    g = StateGraph(TwoAgentState)

    g.add_node("fetch", ta_fetch)
    g.add_node("analyzer", ta_analyzer)
    g.add_node("reviewer", ta_reviewer)
    g.add_node("merge", ta_merge)
    g.add_node("escalate", ta_escalate)
    g.add_node("emit", ta_emit)

    g.add_edge(START, "fetch")
    g.add_edge("fetch", "analyzer")
    g.add_edge("analyzer", "reviewer")
    g.add_edge("reviewer", "merge")

    g.add_conditional_edges("merge", route_after_merge,
        {"escalate": "escalate", "emit": "emit"})

    g.add_edge("emit", END)

    # A checkpointer is required to use interrupt()
    # Use MemorySaver for local testing, Postgres for production
    return g.compile(checkpointer=MemorySaver())

运行比较

在“执行比较”阶段,应在修改代码之前明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 需同时记录正常流程和异常恢复流程。重试机制、人工审核环节以及错误处理都是产品本身的组成部分,而非后续需要补充的功能。 对于涉及资金支出或修改生产数据的操作,必须经过人工审批。编译时的配置并不等同于业务功能的完整性。

决策矩阵:何时不应使用多智能体系统

在“决策矩阵”阶段,应在修改代码之前明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。相比冗长的脚本,更应优先使用小型、可测试的单元。当某个步骤失败时,故障原因应能指向单一责任主体,而非复杂的流程链。对于涉及资金支出或修改生产数据的操作,必须经过人工审批。仅靠编译时的配置并不能保证业务的完整性。

下一步是什么

在“下一步做什么”阶段,应在修改代码之前明确输入参数、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 将此阶段视为输入与已验证输出之间的契约。为相关成果命名,定义成功检测标准,并拒绝默许的半完成状态。 对于涉及资金支出或修改生产数据的操作,必须经过人工审批。编译时的连接关系并不等同于业务上的完整性。 在“下一步做什么”阶段,应在修改代码之前明确输入参数、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 将配置信息置于应用程序代码之外。环境文件、密钥存储以及功能开关应集中存放于一个操作人员可以审核的地方,无需阅读整个系统结构。

继续阅读

在进入“继续阅读”阶段时,首先写下相关契约:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 同时记录正常流程和异常恢复流程。重试机制、人工审核环节以及错误处理方式都是产品本身的组成部分,而非后续需要补充的功能。 在成本较高的步骤之后设置检查点。当操作员重新尝试某个节点时,恢复流程不应再次调用相同的大型语言模型。

运营检查清单

将“运营检查清单”阶段视为可衡量的工作面,效果会更好。在扩大范围之前,先记录一份最佳操作示例、一个失败案例以及回滚说明。 在功能结果之外,还需记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外收费。

保持图结构扁平且具有类型约束。嵌套的数据块会隐藏是哪个节点修改了哪个字段,还会在中断后导致无法继续执行。

只要预算允许,就在持续集成过程中使用测试用例而非真实的付费 API 来对关键路径进行压力测试。

将配置信息与应用程序代码分开。环境文件、密钥存储以及功能开关应集中存放于一个位置,以便操作人员无需查看整个图结构即可进行审计。

保持图结构扁平且具有类型约束。嵌套的数据块会隐藏是哪个节点修改了哪个字段,还会在中断后导致无法继续执行。

在升级技术栈之前,先冻结现有版本,为关键路径生成标准化的操作记录,并确认回滚步骤。共享环境需要设置速率限制、租户验证机制,以及明确的密钥轮换负责人。与其追求花哨的一次性演示,不如注重扎实的可靠性。

f4e352541695的批量处理说明:不要将提供者密钥放入代码仓库,为每个会话设置令牌使用上限,并将转录内容存储在评估测试用例的旁边,以便后续更换模型时仍能保持可比性。

在处理强化安全措施的第0阶段时,首先明确相关规范:所需输入、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改始终符合要求。同时要在功能测试结果旁记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外费用。

强化安全措施细节0/727:针对此条说明,需测量实际执行时间、错误类型以及令牌消耗情况,然后根据固定的评估标准而非主观判断来决定是否保留该修改。

在将硬化处理视为可测量的表面时,第一阶段的效果最佳。在扩大范围之前,需记录一份理想的运行日志、一个故障案例以及回滚说明。同时文档化正常流程与恢复流程。重试机制、人工审核环节以及错误处理都属于产品本身的组成部分,而非后续需要补充的内容。

硬化处理细节1/727:针对该步骤需测量耗时、错误类型以及令牌使用情况,然后依据固定的问题集而非个人经验来判断是否保留该变更。

对于硬化处理的第二阶段,在修改代码之前需明确输入参数、该步骤的负责人以及完成标准。操作人员应能够从已知的检查点重新执行该步骤,而无需猜测隐藏状态。应将此阶段视为输入与验证后输出之间的契约,为相关成果命名、定义成功检测标准,并拒绝默许的半完成状态。

强化措施细节2/727:为该记录测量墙钟时间、错误类型以及代币消耗情况,然后依据固定的问题清单而非个人经验来判断是否保留该变更。

在处理强化措施笔记的第三阶段时,首先写下合约的相关内容:所需输入、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改更加规范。 应将配置信息置于应用程序代码之外。环境文件、密钥存储以及功能开关应集中存放于一个位置,以便操作人员无需查看整个系统结构即可进行审计。

强化措施细节3/727:为该记录测量墙钟时间、错误类型以及代币消耗情况,然后依据固定的问题清单而非个人经验来判断是否保留该变更。

在将加固措施视为可测量的表面时,第4阶段的效果最佳。在扩大范围之前,需记录一份理想的测试用例、一个故障案例以及回滚说明。 相比庞大的脚本,应优先选择小型且可测试的单元。当某个步骤出现故障时,故障原因应能指向单一责任主体,而非复杂的流程链。

加固细节4/727:需为该措施测量执行时间、错误类型以及令牌消耗情况,然后依据固定的问题集而非个人经验来判断是否保留该变更。