实用笔记:多智能体系统——何时两个智能体会战胜一个(以及何时不会)
《实用笔记:多智能体系统——何时两个智能体会战胜一个(以及何时不会)》的操作指南,为采用该设计模式的团队提供契约、校验规则及可直接插入的代码片段。
本指南将逐步构建从原始材料到可运行系统的完整流程,主题为“多智能体系统:当两个智能体战胜一个时(以及无法战胜时)”。重点在于可操作的步骤、明确的检查点,以及可直接放入代码库的代码,无需猜测其用途。 在概览阶段,应在修改代码之前明确输入参数、各步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行相应步骤,而无需猜测隐藏状态。 配置信息应与应用程序代码分开存放。环境文件、密钥存储和功能开关应集中于一个位置,以便操作人员无需查看整个系统结构即可进行审核。
代码审查的问题
在处理“代码审查问题”阶段时,首先写下相关契约:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 同时记录正常流程和异常恢复流程。重试机制、人工审核环节以及错误消息处理都是产品本身的组成部分,而非后续需要补充的功能。 在成本较高的操作之后设置检查点。当操作员重新尝试某个节点时,恢复流程不应再次调用相同的大型语言模型。
--- a/src/billing/invoice.py
+++ b/src/billing/invoice.py
@@ -42,7 +42,9 @@
class InvoiceService:
- def calculate_total(self, items):
- return sum(i.price * i.qty for i in items)
+ def calculate_total(self, items, discount_pct=0):
+ subtotal = sum(i.price * i.qty for i in items)
+ return subtotal * (1 - discount_pct)
--- a/src/billing/api.py
+++ b/src/billing/api.py
@@ -18,6 +18,8 @@
@router.post("/invoice")
def create_invoice(req: InvoiceRequest):
+ discount = req.discount_pct # NEW: from user input
svc = InvoiceService()
- total = svc.calculate_total(req.items)
+ total = svc.calculate_total(req.items, discount)
return {"total": total}
--- a/config/feature_flags.yaml
+++ b/config/feature_flags.yaml
@@ -5,3 +5,4 @@
flags:
new_dashboard: true
+ discount_billing: true # rollout: 100% immediately
设计1:单智能体审查器
在处理“设计1:单智能体”阶段时,首先需写下相关契约:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 建议使用小型、可测试的单元,而非庞大的脚本。当某个步骤失败时,故障应指向单一责任模块,而非复杂的流程链。 在成本较高的步骤之后设置检查点。当操作员重新尝试后续节点时,恢复流程不应再次调用相同的大型语言模型。
from langchain_core.tools import tool
import textwrap
MOCK_DIFF = """...""" # The diff shown above
MOCK_FILES = {
"src/billing/invoice.py": "class InvoiceService:\n def calculate_total(self, items, discount_pct=0):\n subtotal = sum(i.price * i.qty for i in items)\n return subtotal * (1 - discount_pct)\n",
"src/billing/api.py": '@router.post("/invoice")\ndef create_invoice(req: InvoiceRequest):\n discount = req.discount_pct\n svc = InvoiceService()\n total = svc.calculate_total(req.items, discount)\n return {"total": total}\n',
"tests/test_billing.py": "def test_calculate_total():\n # only tests no-discount path\n assert svc.calculate_total(items) == 300\n",
}
@tool
def get_diff(pr_id: str) -> str:
"""Fetch the PR diff."""
return MOCK_DIFF
@tool
def read_file(path: str) -> str:
"""Read a file from the repo."""
return MOCK_FILES.get(path, f"FILE NOT FOUND: {path}")
@tool
def search_symbol(name: str) -> str:
"""Search for a symbol across the codebase."""
if "discount" in name.lower():
return "Found: InvoiceService.calculate_total(discount_pct) — src/billing/invoice.py:43"
return f"No results for '{name}'"
@tool
def list_tests(path: str) -> str:
"""List test files covering a source path."""
if "billing" in path:
return "tests/test_billing.py — covers calculate_total (no-discount path only)"
return "No tests found"
ALL_TOOLS = [get_diff, read_file, search_symbol, list_tests]
from typing import Literal
from typing_extensions import TypedDict
from pydantic import BaseModel
from langchain_google_genai import ChatGoogleGenerativeAI
from langchain_core.messages import AIMessage, HumanMessage, SystemMessage
from langgraph.graph import END, START, StateGraph
class ReviewFinding(BaseModel):
severity: Literal["critical", "high", "medium", "low", "info"]
category: Literal["bug", "regression", "missing_test", "rollout_risk",
"security", "edge_case", "style"]
file: str
description: str
confidence: Literal["high", "medium", "low"]
class SingleAgentState(TypedDict):
pr_id: str
messages: list
diff: str
findings: list[ReviewFinding]
llm = ChatGoogleGenerativeAI(model="gemini-2.5-flash", temperature=0)
def sa_fetch(state: SingleAgentState) -> dict:
diff = get_diff.invoke({"pr_id": state["pr_id"]})
tests = list_tests.invoke({"path": "src/billing"})
return {"diff": diff, "messages": [
AIMessage(content=f"Diff loaded. Test coverage: {tests}")
]}
def sa_review(state: SingleAgentState) -> dict:
"""Single agent does BOTH jobs: summarize + critique in one pass."""
prompt = f"""\
You are a senior code reviewer. Read this PR diff, summarize the change,
and produce a list of risk findings. Be specific.
DIFF:
{state['diff']}
Respond with JSON: {{"summary": "...", "findings": [
{{"severity": "...", "category": "...", "file": "...",
"description": "...", "confidence": "..."}}
]}}
"""
resp = llm.invoke([SystemMessage(content="You are a code review bot."),
HumanMessage(content=prompt)])
# In a real app we parse the JSON response here.
# We simulate the typical single-agent output for this diff.
findings = [
ReviewFinding(
severity="medium",
category="missing_test",
file="tests/test_billing.py",
description="Add unit tests for the new discount_pct parameter.",
confidence="high"
),
ReviewFinding(
severity="low",
category="style",
file="src/billing/invoice.py",
description="Consider adding type hints to the items list.",
confidence="high"
)
]
return {"findings": findings, "messages": [resp]}
def build_single_agent():
g = StateGraph(SingleAgentState)
g.add_node("fetch", sa_fetch)
g.add_node("review", sa_review)
g.add_edge(START, "fetch")
g.add_edge("fetch", "review")
g.add_edge("review", END)
return g.compile()
交接流程:为何结构化方案至关重要
在处理“结构化交接原因”阶段时,首先需写下相关契约:所需输入、成功信号以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持透明可溯。 将此阶段视为输入与已验证输出之间的契约。为相关产物命名,明确成功判定标准,杜绝无声的半完成状态。 在成本较高的步骤之后设置检查点。当操作员重新尝试后续节点时,恢复流程不应再次调用相同的大型语言模型。 在处理“结构化交接原因”阶段时,首先需写下相关契约:所需输入、成功信号以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持透明可溯。 将配置信息置于应用程序代码之外。环境文件、密钥存储及功能开关应集中存放于一处,以便操作员无需查看整个流程即可进行审计。
class ChangedInterface(BaseModel):
file: str
symbol: str
change_type: Literal["added", "modified", "removed"]
description: str
class ChangeModel(BaseModel):
"""Analyzer → Reviewer handoff. Structured, not prose."""
purpose: str
impacted_files: list[str]
changed_interfaces: list[ChangedInterface]
config_changes: list[str]
migration_risk: bool
assumptions: list[str]
tests_touched: list[str]
tests_likely_needed: list[str]
设计2:双智能体架构
在设计2的双智能体阶段中,若能将其视为可度量的结构,则效果最佳。在扩大范围之前,需记录一个成功的用例、一个失败案例以及回滚说明。 同时记录正常流程与恢复流程。重试机制、人工审核环节以及死信处理都是产品本身的组成部分,而非后续需要补充的功能。 保持图结构的扁平化与类型化。嵌套的数据块会掩盖哪个节点编写了哪个字段的信息,且在中断后会导致流程无法继续。
class TwoAgentState(TypedDict):
pr_id: str
messages: list
diff: str
change_model: ChangeModel | None
findings: list[ReviewFinding]
conflict: bool
ANALYZER_PROMPT = """\
You are the Analyzer agent. Your ONLY job: understand the change.
Do NOT critique. Do NOT hunt for bugs. Just map what changed.
Produce a structured change model."""
def ta_analyzer(state: TwoAgentState) -> dict:
"""Analyzer: compress & clarify. Optimizes for coherence."""
# We use LLM structured output to fill the ChangeModel.
# For this demonstration, we simulate the accurate analysis.
cm = ChangeModel(
purpose="Add discount percentage support to invoice billing",
impacted_files=["src/billing/invoice.py", "src/billing/api.py",
"config/feature_flags.yaml"],
changed_interfaces=[
ChangedInterface(file="src/billing/invoice.py",
symbol="InvoiceService.calculate_total",
change_type="modified",
description="Added discount_pct param (default 0)"),
ChangedInterface(file="src/billing/api.py",
symbol="create_invoice",
change_type="modified",
description="Reads discount_pct from request, passes to service"),
],
config_changes=["discount_billing flag added, 100% rollout"],
migration_risk=False,
assumptions=[
"discount_pct is expected to be 0..1 (fraction, not percentage)",
"No existing callers pass discount_pct yet",
"Feature flag controls visibility, not the calculation",
],
tests_touched=[],
tests_likely_needed=[
"test discount path in calculate_total",
"test boundary: discount_pct = 0, 1, >1, <0",
"test API validation of discount_pct input",
],
)
return {
"change_model": cm,
"messages": [AIMessage(content=f"Analyzer: change model built. "
f"{len(cm.changed_interfaces)} interfaces changed, "
f"{len(cm.assumptions)} assumptions made.")],
}
REVIEWER_PROMPT = """\
You are the Risk Reviewer. The Analyzer gave you a change model.
DISTRUST it. Your job: find what's missing, broken, or dangerous.
Challenge every assumption. Check for missing tests, regressions,
rollout risks, and security issues."""
def ta_reviewer(state: TwoAgentState) -> dict:
"""Reviewer: expand & challenge. Optimizes for skepticism."""
cm = state["change_model"]
findings: list[ReviewFinding] = []
# The Reviewer iterates through the Analyzer's assumptions
for a in cm.assumptions:
if "0..1" in a:
findings.append(ReviewFinding(
severity="critical",
category="security",
file="src/billing/api.py",
description="ASSUMPTION CHALLENGED: discount_pct comes from user input "
"(req.discount_pct) with NO validation. Values <0 or >1 break billing. "
"Negative discount = price increase beyond subtotal. "
"Value >1 = negative total.",
confidence="high"))
# The Reviewer checks the test mapping
if not cm.tests_touched and cm.tests_likely_needed:
findings.append(ReviewFinding(
severity="high",
category="missing_test",
file="tests/test_billing.py",
description=f"NO tests touched but {len(cm.tests_likely_needed)} needed: "
+ "; ".join(cm.tests_likely_needed),
confidence="high"))
# The Reviewer checks the config changes
for cc in cm.config_changes:
if "100%" in cc:
findings.append(ReviewFinding(
severity="high",
category="rollout_risk",
file="config/feature_flags.yaml",
description="Feature flag at 100% from day one — no gradual rollout. "
"Combined with unvalidated discount input, this is a billing incident risk.",
confidence="high"))
# The Reviewer looks for edge cases in the logic flow
findings.append(ReviewFinding(
severity="medium",
category="edge_case",
file="src/billing/api.py",
description="Feature flag controls visibility but calculate_total always applies "
"discount. If flag is off but API still receives discount_pct, discount "
"is silently applied.",
confidence="medium"))
# We determine if there is a conflict worth escalating
conflict = any(f.severity in ("critical", "high") for f in findings)
return {
"findings": findings,
"conflict": conflict,
"messages": [AIMessage(content=f"Reviewer: {len(findings)} findings, "
f"conflict={conflict}")],
}
仲裁机制与人工干预
将仲裁与人工干预阶段视为可度量的对象来处理,效果最佳。在扩大范围之前,先记录一份完美的操作日志、一个故障案例以及回滚说明。 优先选择小型且可测试的单元,而非庞大的脚本。当某一步骤出现故障时,故障应指向单一责任主体,而非复杂的流程链。 保持图结构简洁且类型明确。嵌套的数据块会掩盖哪个节点编写了哪个字段的信息,还会在中断后导致流程无法继续。
def ta_merge(state: TwoAgentState) -> dict:
"""Merge findings. If conflict, surface for human review."""
if state["conflict"]:
return {
"messages": [AIMessage(content=(
"⚠ CONFLICT: Reviewer found critical/high issues. "
"Routing to human review."
))],
}
return {
"messages": [AIMessage(content="Findings merged. No escalation needed.")],
}
def route_after_merge(state: TwoAgentState) -> str:
return "escalate" if state["conflict"] else "emit"
from langgraph.types import Command, interrupt
def ta_escalate(state: TwoAgentState) -> Command:
"""Human-in-the-loop for conflicting findings."""
# The interrupt function pauses execution and surfaces data to the caller
decision = interrupt({
"kind": "review_conflict",
"pr_id": state["pr_id"],
"change_model": state["change_model"].model_dump(),
"findings": [f.model_dump() for f in state["findings"]],
"prompt": "Review findings. Respond: "
'{"action":"accept_all"} or {"action":"override","drop_indices":[...]}',
})
# Execution resumes here when the human provides input
action = decision.get("action", "accept_all")
if action == "override":
drop = set(decision.get("drop_indices", []))
kept = [f for i, f in enumerate(state["findings"]) if i not in drop]
# We use Command to update state and dynamically route to the next node
return Command(update={"findings": kept}, goto="emit")
return Command(goto="emit")
# How you resume the graph from your backend API
ta.invoke(Command(resume={"action": "accept_all"}), config)
def ta_emit(state: TwoAgentState) -> dict:
"""Final output: formatted review."""
lines = [f"=== Code Review: PR {state['pr_id']} ==="]
if state["change_model"]:
lines.append(f"Purpose: {state['change_model'].purpose}")
lines.append(f"Findings ({len(state['findings'])}):")
for i, f in enumerate(state["findings"]):
lines.append(f" [{f.severity}] ({f.category}) {f.file}")
lines.append(f" {f.description}")
return {"messages": [AIMessage(content="\n".join(lines))]}
连接双智能体图结构
将“连接两个智能体图”阶段视为可度量的对象来处理效果最佳。在扩大范围之前,需记录一份理想状态下的完整流程、一个故障案例以及回滚说明。 应将此阶段视为输入与已验证输出之间的契约。为相关成果命名,明确成功标准,绝不允许出现无声的半完成状态。 保持图结构的扁平化与类型化。嵌套的数据块会掩盖哪个节点修改了哪个字段,还会在中断后导致无法继续处理。 将“连接两个智能体图”阶段视为可度量的对象来处理效果最佳。在扩大范围之前,需记录一份理想状态下的完整流程、一个故障案例以及回滚说明。 将配置信息置于应用程序代码之外。环境文件、密钥存储和功能开关应集中存放于一个位置,以便操作人员无需查看整个图结构即可进行审计。
from langgraph.checkpoint.memory import MemorySaver
def build_two_agent():
g = StateGraph(TwoAgentState)
g.add_node("fetch", ta_fetch)
g.add_node("analyzer", ta_analyzer)
g.add_node("reviewer", ta_reviewer)
g.add_node("merge", ta_merge)
g.add_node("escalate", ta_escalate)
g.add_node("emit", ta_emit)
g.add_edge(START, "fetch")
g.add_edge("fetch", "analyzer")
g.add_edge("analyzer", "reviewer")
g.add_edge("reviewer", "merge")
g.add_conditional_edges("merge", route_after_merge,
{"escalate": "escalate", "emit": "emit"})
g.add_edge("emit", END)
# A checkpointer is required to use interrupt()
# Use MemorySaver for local testing, Postgres for production
return g.compile(checkpointer=MemorySaver())
运行比较
在“执行比较”阶段,应在修改代码之前明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 需同时记录正常流程和异常恢复流程。重试机制、人工审核环节以及错误处理都是产品本身的组成部分,而非后续需要补充的功能。 对于涉及资金支出或修改生产数据的操作,必须经过人工审批。编译时的配置并不等同于业务功能的完整性。
决策矩阵:何时不应使用多智能体系统
在“决策矩阵”阶段,应在修改代码之前明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。相比冗长的脚本,更应优先使用小型、可测试的单元。当某个步骤失败时,故障原因应能指向单一责任主体,而非复杂的流程链。对于涉及资金支出或修改生产数据的操作,必须经过人工审批。仅靠编译时的配置并不能保证业务的完整性。
下一步是什么
在“下一步做什么”阶段,应在修改代码之前明确输入参数、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 将此阶段视为输入与已验证输出之间的契约。为相关成果命名,定义成功检测标准,并拒绝默许的半完成状态。 对于涉及资金支出或修改生产数据的操作,必须经过人工审批。编译时的连接关系并不等同于业务上的完整性。 在“下一步做什么”阶段,应在修改代码之前明确输入参数、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 将配置信息置于应用程序代码之外。环境文件、密钥存储以及功能开关应集中存放于一个操作人员可以审核的地方,无需阅读整个系统结构。
继续阅读
在进入“继续阅读”阶段时,首先写下相关契约:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 同时记录正常流程和异常恢复流程。重试机制、人工审核环节以及错误处理方式都是产品本身的组成部分,而非后续需要补充的功能。 在成本较高的步骤之后设置检查点。当操作员重新尝试某个节点时,恢复流程不应再次调用相同的大型语言模型。
运营检查清单
将“运营检查清单”阶段视为可衡量的工作面,效果会更好。在扩大范围之前,先记录一份最佳操作示例、一个失败案例以及回滚说明。 在功能结果之外,还需记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外收费。
保持图结构扁平且具有类型约束。嵌套的数据块会隐藏是哪个节点修改了哪个字段,还会在中断后导致无法继续执行。
只要预算允许,就在持续集成过程中使用测试用例而非真实的付费 API 来对关键路径进行压力测试。
将配置信息与应用程序代码分开。环境文件、密钥存储以及功能开关应集中存放于一个位置,以便操作人员无需查看整个图结构即可进行审计。
保持图结构扁平且具有类型约束。嵌套的数据块会隐藏是哪个节点修改了哪个字段,还会在中断后导致无法继续执行。
在升级技术栈之前,先冻结现有版本,为关键路径生成标准化的操作记录,并确认回滚步骤。共享环境需要设置速率限制、租户验证机制,以及明确的密钥轮换负责人。与其追求花哨的一次性演示,不如注重扎实的可靠性。
f4e352541695的批量处理说明:不要将提供者密钥放入代码仓库,为每个会话设置令牌使用上限,并将转录内容存储在评估测试用例的旁边,以便后续更换模型时仍能保持可比性。
在处理强化安全措施的第0阶段时,首先明确相关规范:所需输入、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改始终符合要求。同时要在功能测试结果旁记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外费用。
强化安全措施细节0/727:针对此条说明,需测量实际执行时间、错误类型以及令牌消耗情况,然后根据固定的评估标准而非主观判断来决定是否保留该修改。
在将硬化处理视为可测量的表面时,第一阶段的效果最佳。在扩大范围之前,需记录一份理想的运行日志、一个故障案例以及回滚说明。同时文档化正常流程与恢复流程。重试机制、人工审核环节以及错误处理都属于产品本身的组成部分,而非后续需要补充的内容。
硬化处理细节1/727:针对该步骤需测量耗时、错误类型以及令牌使用情况,然后依据固定的问题集而非个人经验来判断是否保留该变更。
对于硬化处理的第二阶段,在修改代码之前需明确输入参数、该步骤的负责人以及完成标准。操作人员应能够从已知的检查点重新执行该步骤,而无需猜测隐藏状态。应将此阶段视为输入与验证后输出之间的契约,为相关成果命名、定义成功检测标准,并拒绝默许的半完成状态。
强化措施细节2/727:为该记录测量墙钟时间、错误类型以及代币消耗情况,然后依据固定的问题清单而非个人经验来判断是否保留该变更。
在处理强化措施笔记的第三阶段时,首先写下合约的相关内容:所需输入、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改更加规范。 应将配置信息置于应用程序代码之外。环境文件、密钥存储以及功能开关应集中存放于一个位置,以便操作人员无需查看整个系统结构即可进行审计。
强化措施细节3/727:为该记录测量墙钟时间、错误类型以及代币消耗情况,然后依据固定的问题清单而非个人经验来判断是否保留该变更。
在将加固措施视为可测量的表面时,第4阶段的效果最佳。在扩大范围之前,需记录一份理想的测试用例、一个故障案例以及回滚说明。 相比庞大的脚本,应优先选择小型且可测试的单元。当某个步骤出现故障时,故障原因应能指向单一责任主体,而非复杂的流程链。
加固细节4/727:需为该措施测量执行时间、错误类型以及令牌消耗情况,然后依据固定的问题集而非个人经验来判断是否保留该变更。