首页 / 文章 / 实用提示:您使用的AI智能体框架很可能不合适,具体原因如下

实用提示:您使用的AI智能体框架很可能不合适,具体原因如下

《实用笔记》操作指南:您的 AI 智能体框架很可能选错了,具体方法如下——为采用该模式的团队提供合同、校验机制以及可直接插入的代码模块。

2234 词

本指南将逐步构建从原材料到可运行系统的完整流程,主题为:您当前使用的AI智能体框架可能并不合适,以下是正确选择的方法。重点在于可操作的步骤、明确的检查点,以及可直接放入代码仓库的代码,无需猜测其用途。 在概览阶段,应在修改代码之前明确输入参数、各步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行相应步骤,而无需推测隐藏状态。 除了功能结果外,还需记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在流程从演示环境过渡到共享环境时出现意外费用。

大家都反着问的问题

在处理“大家都会问的问题”这一阶段时,首先写下相关契约:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 将配置信息置于应用程序代码之外。环境文件、密钥存储以及功能开关应集中存放于一个位置,这样操作人员无需查看整个系统结构即可进行审计。 在耗时较高的步骤之后设置检查点。当操作人员重新执行后续节点时,恢复流程不应再次调用相同的大型语言模型。

轴1:你的分支结构需要多高的确定性?

在处理“轴1:确定性如何”这一阶段时,首先写下相关约定:所需输入、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 同时记录正常流程和恢复流程。重试机制、人工审核环节以及死信处理都是产品本身的组成部分,而非后续的优化内容。 在成本较高的步骤之后设置检查点。当操作员重新尝试某个后续节点时,恢复流程不应再次调用相同的大型语言模型。

# A branch where non-determinism is FINE — picking a tone for a summary email.
# If the agent occasionally phrases things slightly differently, nobody's paged.
def draft_summary_tone(context: dict) -> str:
    return llm_call(
        prompt=f"Summarize this incident in a {context['audience']}-appropriate tone.",
        temperature=0.7,  # variability here is a feature, not a bug
    )
# A branch where non-determinism is NOT fine — deciding whether to page a human
# at 4am versus auto-remediating. This must be code, not a prompt.
def route_alert(alert: dict) -> str:
    if alert["severity"] == "critical" and alert["service"] in PAGE_ALWAYS_SERVICES:
        return "page_oncall"
    if alert["auto_remediation_available"] and alert["confidence"] > 0.9:
        return "auto_remediate"
    if alert["severity"] == "critical":
        return "page_oncall"
    return "log_and_monitor"

轴2:一个工作单元的生命周期有多长?

在处理 Axis 2 的“耗时”阶段时,首先需明确相关约定:所需输入、成功信号以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持透明。 建议采用小型、可测试的单元,而非冗长的脚本。当某个步骤失败时,故障应指向单一责任点,而非复杂的流程链。 在成本较高的步骤之后设置检查点。当操作员重新尝试后续节点时,恢复流程不应再次调用相同的 LLM 接口。

# Short-lived: starts and finishes inside one HTTP request.
# This is the "no framework needed" zone — a framework here is pure overhead.
async def handle_summarize_request(request: SummarizeRequest) -> SummarizeResponse:
    text = await fetch_document(request.doc_id)
    summary = await llm_summarize(text, max_tokens=300)
    return SummarizeResponse(summary=summary)
# Long-lived: this alert might sit in "awaiting human ack" for six hours
# while the on-call engineer is asleep, then resume on a completely
# different process after a deploy rotated the pods underneath it.
class AlertTriageWorkflow:
    async def run(self, alert: dict) -> dict:
        decision = await self.classify_and_route(alert)
        if decision == "page_oncall":
            await self.page(alert)
            await self.wait_for_ack(timeout_hours=1)  # this line is the whole ballgame
        ...

在处理 Axis 2 的“耗时”阶段时,首先需明确相关约定:所需输入、成功信号以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持透明。 除了功能结果外,还需记录执行时间以及 Token 或查询成本。提前了解成本情况,可避免在流程从演示环境转向共享环境时出现意外费用。

第三轴:如果某个步骤被执行两次会怎样?

将“第三轴:会发生什么”这一阶段视为可度量的指标会更为有效。在扩大范围之前,先记录一份理想的执行结果、一个失败案例以及回滚说明。 将配置置于应用程序代码之外。环境文件、密钥存储和功能标志应集中存放于一个位置,这样操作人员无需查看整个流程图即可进行审计。 保持流程图的状态结构化且具有类型定义。嵌套的数据块会掩盖哪个节点修改了哪个字段的信息,还会导致在流程中断后无法继续执行。

# BEFORE — looks fine in a demo, is a live incident waiting to happen
async def auto_remediate(alert: dict):
    await restart_service(alert["service"])  # what if this activity gets retried?
# AFTER — idempotent by construction
async def auto_remediate(alert: dict, idempotency_key: str):
    if await remediation_ledger.already_applied(idempotency_key):
        logger.info("remediation already applied, skipping", key=idempotency_key)
        return await remediation_ledger.get_result(idempotency_key)
    result = await restart_service(alert["service"])
    await remediation_ledger.record(idempotency_key, result)
    return result

第四轴:谁需要在后续读取决策结果,以及以何种形式读取?

将“轴4:需要舞台环境”的机制视为可度量的表面来处理效果最佳。在扩大范围之前,先记录一个成功的案例、一个失败案例以及回滚说明。 同时记录正常流程和恢复流程。重试机制、人工审核环节以及死信处理都是产品本身的一部分,而非后续的优化工作。 保持图表状态简洁且具有类型定义。嵌套的数据块会掩盖哪个节点修改了哪个字段的信息,还会在中断后导致流程无法继续。

# A framework-agnostic audit record — this is what actually matters
# in a postmortem, regardless of what orchestrated the steps.
@dataclass
class DecisionRecord:
    alert_id: str
    timestamp: float
    step: str
    reasoning: str        # what the LLM said, verbatim
    decision: str         # the structured outcome, not prose
    confidence: float | None
    human_override: bool

async def log_decision(record: DecisionRecord):
    await audit_store.insert(record)
    # Also emit as a structured log line — cheap insurance for when
    # the audit store itself is the thing that's down during an incident.
    logger.info("agent_decision", **asdict(record))

轴5:您团队的实际开发速度限制是什么?

Axis 5 What’s 阶段作为可度量的对象来处理时效果最佳。在扩大范围之前,先记录一份理想的执行日志、一个故障案例以及回滚说明。 相较于庞大的脚本,应优先选择小型且可测试的单元。当某一步骤出现故障时,故障应指向单一责任点,而非复杂的流程链。 保持图表状态简洁且具有类型定义。嵌套的数据结构会掩盖哪个节点修改了哪个字段的信息,还会在进程中断后导致无法继续执行。 Axis 5 What’s 阶段作为可度量的对象来处理时效果最佳。在扩大范围之前,先记录一份理想的执行日志、一个故障案例以及回滚说明。 除了功能结果外,还需记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外费用。

# Week-one prototype: prove the concept fast, accept the debt knowingly.
from crewai import Agent, Task, Crew

triage_agent = Agent(role="Alert Triage", goal="Decide how to handle infra alerts")
crew = Crew(agents=[triage_agent], tasks=[Task(description="Triage: {alert}", agent=triage_agent)])
crew.kickoff(inputs={"alert": alert_payload})

Axis 6:每个决策的延迟和成本预算是多少?

对于 Axis 6 的 What’s Stage 模型,在修改代码之前需明确输入参数、该步骤的负责人以及结束条件。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 配置信息应置于应用程序代码之外。环境文件、密钥存储以及功能标志应集中存放于一个位置,这样操作人员无需查看整个流程即可进行审核。 对于涉及资金支出或修改生产数据的节点,必须设置人工审批环节。编译时的连接方式并不等同于业务流程的完整性。

# Expensive pattern: every routing decision is its own LLM call,
# multiplied across a multi-agent conversation with several turns.
# At alert volumes (hundreds/day, sometimes bursts of thousands during
# a real incident), this is a real line item, not a rounding error.
async def route_via_llm(alert: dict) -> str:
    return await llm_call(f"How should we handle this alert? {alert}")

# Cheaper, faster, and more auditable: cheap deterministic pre-filtering
# in code, LLM reserved for genuinely ambiguous cases.
async def route_alert_efficiently(alert: dict) -> str:
    if alert["service"] in KNOWN_NOISY_SERVICES and alert["severity"] == "low":
        return "log_and_monitor"          # zero LLM calls for the common case
    if alert["signature"] in KNOWN_REMEDIATION_PLAYBOOK:
        return "auto_remediate"           # deterministic lookup, zero LLM calls
    return await llm_call(f"Novel alert, needs judgment: {alert}")  # LLM only when genuinely needed

整合应用:决策路径而非决策树

在进入“整合阶段”之前,需先明确输入参数、该步骤的负责人以及结束标准,然后再进行代码修改。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 需同时记录正常流程和异常恢复流程。重试机制、人工审核环节以及错误处理都是产品本身的组成部分,而非后续需要补充的内容。 对于涉及资金支出或修改生产数据的操作,必须经过人工审批。仅靠编译时的配置并不能保证业务的完整性。

Is this unit of work stateless and finishes in seconds?
  └─ YES → skip the framework entirely. Plain functions + retries. Ship it.
  └─ NO, continue.

Does it need to survive process restarts / wait on humans for hours-to-days?
  └─ YES → you need durable execution (Temporal or equivalent) as the backbone,
           regardless of what else you pick for the reasoning layer.
  └─ NO, continue.

Are the important branches safety- or compliance-critical
(money, infra changes, irreversible external actions)?
  └─ YES → LangGraph-style explicit graphs, keep LLM scoped to narrow nodes.
  └─ NO, mostly exploratory/creative → CrewAI or AutoGen are legitimate defaults.

Is this still a prototype whose findings might get thrown away?
  └─ YES → optimize for speed of iteration over long-term correctness,
           but write down when you'll revisit that tradeoff.
@activity.defn
async def classify_alert_activity(alert: dict) -> dict:
    # LangGraph-style graph runs here — bounded reasoning, deterministic routing —
    # inside an activity Temporal will retry and time-box like any other side effect.
    result = alert_triage_graph.invoke({"alert": alert, "audit_log": []})
    return {"decision": result["decision"], "confidence": result["confidence"]}

@workflow.defn
class AlertTriageWorkflow:
    def __init__(self):
        self._acked = False

    @workflow.signal
    async def acknowledge(self):
        self._acked = True

    @workflow.run
    async def run(self, alert: dict) -> dict:
        classification = await workflow.execute_activity(
            classify_alert_activity, alert,
            start_to_close_timeout=timedelta(seconds=20),
            retry_policy=workflow.RetryPolicy(maximum_attempts=3),
        )
        if classification["decision"] == "page_oncall":
            await workflow.execute_activity(page_oncall, alert, start_to_close_timeout=timedelta(seconds=10))
            await workflow.wait_condition(lambda: self._acked, timeout=timedelta(hours=1))
            if not self._acked:
                await workflow.execute_activity(escalate_to_secondary, alert, start_to_close_timeout=timedelta(seconds=10))
        elif classification["decision"] == "auto_remediate":
            await workflow.execute_activity(
                auto_remediate, alert, f"remediate-{alert['id']}",
                start_to_close_timeout=timedelta(minutes=2),
                retry_policy=workflow.RetryPolicy(maximum_attempts=2),
            )
        return {"alert_id": alert["id"], "decision": classification["decision"]}

常见且反复出现的错误

在修改代码之前,应避免常见错误:先明确阶段、输入参数、该步骤的负责人以及退出标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 优先选择小型、可测试的单元,而非庞大的脚本。当某个步骤失败时,故障应指向单一责任点,而非复杂的流程链。 对于涉及资金支出或修改生产数据的操作,必须经过人工审批。编译时的连接方式并不等同于业务功能的完整性。 在修改代码之前,应避免常见错误:先明确阶段、输入参数、该步骤的负责人以及退出标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 除了功能结果外,还需记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在流程从演示环境转向共享环境时出现意外账单。

实际答案

在处理“实际答案”阶段时,首先写下相关契约:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 将配置信息置于应用程序代码之外。环境文件、密钥存储以及功能开关应集中存放于一个位置,以便操作人员无需查看整个系统结构即可进行审计。 在耗时较高的步骤之后设置检查点。当操作人员重新尝试后续节点时,恢复流程不应再次调用相同的大型语言模型。

运营检查清单

将“运营检查清单”阶段视为可度量的指标体系,效果最佳。在扩大范围之前,先记录一份标准示例、一个失败案例以及回滚说明。 应将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,明确成功标准,并杜绝无声的半完成状态。

保持图状态扁平且具有类型约束。嵌套的二进制数据会隐藏是哪个节点修改了哪个字段,还会在中断后导致程序无法继续运行。

只要预算允许,就在持续集成过程中使用测试用例而非真实的付费 API 来执行关键路径的冒烟测试。

在功能结果旁记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外账单。

保持图状态扁平且具有类型约束。嵌套的二进制数据会隐藏是哪个节点修改了哪个字段,还会在中断后导致程序无法继续运行。

在升级技术栈之前,先冻结现有版本,为关键路径生成标准测试记录,并确认回滚步骤。共享环境需要设置速率限制、进行租户身份验证,同时明确负责密钥轮换的人员。与其追求花哨的一次性演示,不如注重扎实的可靠性。

72c003459fd6的批处理说明:不要将提供者密钥放入仓库中,设定每会话的令牌上限,并将转录内容存储在评估测试用例旁边,以便后续模型更换时仍能保持可比性。