首页 / 文章 / 实用提示:避免人工智能代理无限循环——工程指南

实用提示:避免人工智能代理无限循环——工程指南

《实用指南》操作流程详解:避免你的 AI 智能体无限循环——专为采用该模式的团队准备的工程指南,涵盖合同、校验机制以及可直接插入的代码模块。

4049 词

可将此内容作为《切勿让你的 AI 智能体无限循环:终止标准的工程指南》中理念的面向操作员的简化版本:清晰的阶段划分、有序的代码模块以及可在交接时保留的恢复说明。 在将范围扩大之前,最好先将“概览”阶段视为一个可量化的基准。先记录一份理想的操作日志、一个故障案例以及回滚说明。 需同时记录正常流程与恢复流程。重试机制、人工审核环节以及死信处理都是产品本身的组成部分,而非后续需要补充的内容。

周五下午的噩梦

在“周五下午的噩梦”阶段,修改代码之前需明确输入参数、该步骤的负责人以及退出标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 优先选择小型、可测试的单元,而非庞大的脚本。当某个步骤失败时,故障应指向单一责任点,而非复杂的流程链。 对于涉及资金支出或修改生产数据的操作,必须经过人工审批。编译时的逻辑连接并不等同于业务功能的完整性。

智能循环的结构解析

在实现“代理阶段”的架构时,应在修改代码之前明确输入参数、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 应将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,定义成功判定标准,并拒绝默许的半完成状态。 对于涉及资金支出或修改生产数据的操作,必须经过人工审批。编译时的逻辑连接并不等同于业务上的完整性。

                   ┌──────────────────────────────────────┐
                   │        Agent Perception Loop         │
                   │       (Perceive → Plan → Act)        │
                   └──────────────────┬───────────────────┘
                                      │
           ┌──────────────────────────┼──────────────────────────┐
           ▼                          ▼                          ▼
┌────────────────────┐    ┌────────────────────┐    ┌────────────────────┐
│ 1. Success Guard   │    │ 2. Resource Caps   │    │ 3. Progress Guard  │
│ (Programmatic Test)│    │ (Tokens/Turns/Time)│    │ (Loop/Hash Detect) │
└────────────────────┘    └────────────────────┘    └────────────────────┘
                                      │
                                      ▼
                   ┌──────────────────────────────────────┐
                   │  4. Human Handoff / Safe Rollback    │
                   └──────────────────────────────────────┘

1. 成功标准:确定性目标验证

在“1个成功标准:确定性”阶段,应在修改代码之前明确输入参数、该步骤的负责人以及退出条件。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。应在功能结果旁记录执行时间以及令牌或查询成本。提前显示成本信息,可避免在流程从演示环境转向共享环境时出现意外费用。对于会产生费用或更改生产数据的操作,必须经过人工审批。编译时的配置并不等同于业务功能的完整性。

隐患:自我评估

在缺陷自我评估阶段,应在修改代码之前明确输入参数、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 配置信息应置于应用程序代码之外。环境文件、密钥存储以及功能开关应集中存放于一个位置,以便操作人员无需查看整个系统结构即可进行审计。 对于涉及资金支出或修改生产数据的操作,必须经过人工审批。编译时的连接方式并不等同于业务功能的完整性。

解决方案:外部程序化验证工具

在解决方案的外部编程阶段,应在修改代码之前明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 需同时记录正常流程和异常恢复流程。重试机制、人工审核环节以及错误处理都是产品本身的组成部分,而非后续需要补充的内容。 对于涉及资金支出或修改生产数据的操作,必须经过人工审批。编译时的配置并不等同于业务功能的完整性。

import { execSync } from 'node:child_process';

export interface VerificationResult {
  success: boolean;
  message: string;
  stepFailed?: string;
}

/** execSync throws on non-zero exit; diagnostics may land on either stream. */
function runOrCapture(command: string, cwd: string): string | null {
  try {
    execSync(command, { cwd, stdio: 'pipe' });
    return null;
  } catch (err: unknown) {
    const e = err as { stdout?: Buffer; stderr?: Buffer };
    const out = e.stdout?.toString() ?? '';
    const errOut = e.stderr?.toString() ?? '';
    return [out, errOut].filter(Boolean).join('\n') || String(err);
  }
}

export class GoalVerifier {
  public static verify(workspacePath: string): VerificationResult {
    const steps: Array<[string, string]> = [
      ['tsc', 'npx tsc --noEmit'],
      ['npm_test', 'npm test'],
    ];

    for (const [stepId, command] of steps) {
      const failure = runOrCapture(command, workspacePath);
      if (failure !== null) {
        return {
          success: false,
          message: `\`${command}\` failed:\n${failure}`,
          stepFailed: stepId,
        };
      }
    }

    return { success: true, message: 'All typechecks and tests passed cleanly.' };
  }
}

2. 资源与预算上限:严格的引擎限制

在“资源与预算”阶段,应在修改代码之前明确输入参数、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。相比冗长的脚本,更应采用小型且可测试的单元。当某个步骤失败时,故障原因应能指向具体的责任主体,而非复杂的流程链。对于涉及资金支出或修改生产数据的操作,必须经过人工审批。仅靠编译时的配置并不能保证业务的完整性。

隐患:无限制的重试

在“无限制重试”阶段,应在修改代码之前明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 将此阶段视为输入与已验证输出之间的契约。为相关成果命名,定义成功判定标准,并杜绝默许的半完成状态。 对于涉及资金支出或修改生产数据的操作,必须经过人工审批。编译时的配置并不等同于业务上的完整性。

解决方案:多维度限制

在解决方案的多维度限制阶段,修改代码之前需明确输入参数、该步骤的负责人以及退出标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 在功能结果旁记录执行时间以及令牌或查询成本。提前了解成本情况可避免在系统从演示环境切换到共享环境时出现意外账单。 对于会耗费资金或修改生产数据的操作,必须经过人工审批。编译时的配置并不等同于业务功能的完整性。 在解决方案的多维度限制阶段,修改代码之前需明确输入参数、该步骤的负责人以及退出标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 需同时记录正常流程和故障恢复流程。重试机制、人工审核环节以及死信处理都是生产环境不可或缺的部分。

CT,之后不要再做优化。

export interface ResourceLimits {
  maxTurns: number;       // e.g. 10 iterations
  maxTotalTokens: number; // e.g. 100_000 input + output
  timeoutMs: number;      // e.g. 120_000 (2 minutes)
}

export class ResourceGuard {
  private readonly startTime = Date.now();
  private totalTokensUsed = 0;
  private currentTurn = 0;

  constructor(private readonly limits: ResourceLimits) {}

  /** Call once per loop iteration, before the model call. */
  public beginTurn(): void {
    this.currentTurn += 1;
  }

  /** Call for every model call, including retries inside a turn. */
  public recordUsage(tokens: number): void {
    this.totalTokensUsed += tokens;
  }

  public getTurnCount(): number {
    return this.currentTurn;
  }

  public checkShouldTerminate(): { terminate: boolean; reason?: string } {
    if (this.currentTurn >= this.limits.maxTurns) {
      return {
        terminate: true,
        reason: `Exceeded turn cap (${this.limits.maxTurns})`,
      };
    }
    if (this.totalTokensUsed >= this.limits.maxTotalTokens) {
      return {
        terminate: true,
        reason: `Exceeded token budget (${this.totalTokensUsed}/${this.limits.maxTotalTokens})`,
      };
    }
    const elapsed = Date.now() - this.startTime;
    if (elapsed >= this.limits.timeoutMs) {
      return {
        terminate: true,
        reason: `Wall-clock timeout reached (${elapsed}ms/${this.limits.timeoutMs}ms)`,
      };
    }
    return { terminate: false };
  }
}

3. 进度守护机制:卡住与偏移检测

在处理“3个进度守护机制卡住”阶段时,首先列出相关约定:所需输入、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 优先选择小型、可测试的单元,而非庞大的脚本。当某个步骤失败时,故障应指向单一责任模块,而非复杂的流程链。 在成本较高的步骤之后设置检查点。当操作员重新尝试后续节点时,恢复流程不应再次调用相同的LLM接口。

常见故障模式

在处理“常见故障模式”阶段时,首先写下契约:所需的输入、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 将这一阶段视为输入与经过验证的输出之间的契约。为相关成果命名,明确成功判定标准,并杜绝无声的半完成状态。 在成本较高的步骤之后设置检查点。当操作员重新尝试后续节点时,恢复流程不应再次调用相同的大型语言模型。

解决方案:工具签名与工作区状态哈希

在处理解决方案工具签名阶段时,首先需明确合同规范:所需输入、成功信号以及部分失败时的处理方式。这份清单能确保后续的代码修改保持一致性。 在功能结果旁记录执行时间以及令牌或查询成本。提前了解成本情况,可避免从演示环境过渡到共享环境时出现意外费用。 需为每次调用记录工具名称、参数哈希值、延迟时间以及最终结果。没有这些记录,调试过程将会浪费大量时间。 在处理解决方案工具签名阶段时,首先需明确合同规范:所需输入、成功信号以及部分失败时的处理方式。这份清单能确保后续的代码修改保持一致性。 需同时记录正常流程和故障恢复流程。重试机制、人工审核环节以及死信处理都是产品功能的一部分,而非后续需要补充的内容。

import { createHash } from 'node:crypto';

export interface ToolCall {
  name: string;
  args: Record<string, unknown>;
}

export class ProgressGuard {
  private readonly recentActionHashes: string[] = [];

  constructor(
    private readonly windowSize = 5,
    private readonly repeatThreshold = 3,
  ) {}

  /** Stable stringify: key order must not change the hash. */
  private hashToolCall(call: ToolCall): string {
    const args = JSON.stringify(call.args, Object.keys(call.args).sort());
    return createHash('sha256').update(`${call.name}:${args}`).digest('hex');
  }

  /** Returns true when the same call has appeared `repeatThreshold` times in the window. */
  public trackAndCheckStuck(call: ToolCall): boolean {
    const actionHash = this.hashToolCall(call);
    const priorOccurrences = this.recentActionHashes.filter((h) => h === actionHash).length;

    this.recentActionHashes.push(actionHash);
    if (this.recentActionHashes.length > this.windowSize) {
      this.recentActionHashes.shift();
    }

    return priorOccurrences + 1 >= this.repeatThreshold;
  }
}

4. 人工干预机制与安全回滚

将“4. 人工干预机制与安全阶段”视为可度量的框架来使用效果最佳。在扩大范围之前,先记录一份理想的操作流程、一个故障案例以及回滚说明。 优先选择小型且可测试的单元,而非庞大的脚本。当某一步骤出现故障时,故障应能指向具体的责任主体,而非复杂的流程链。 保持图表状态简洁且具有类型定义。嵌套的数据结构会掩盖哪个节点编写了哪个字段的信息,还会在流程中断后导致无法继续执行。

import { execSync } from 'node:child_process';
import { writeFileSync } from 'node:fs';
import { join } from 'node:path';

export class AgentEscalationRequiredError extends Error {
  constructor(message: string, public readonly reportPath?: string) {
    super(message);
    this.name = 'AgentEscalationRequiredError';
  }
}

export class AgentCircuitBreaker {
  constructor(
    private readonly workspaceDir: string,
    private readonly reportDir: string, // keep reports OUTSIDE the workspace
  ) {}

  public handleAbort(reason: string, history: unknown[] = []): never {
    console.error(`[CIRCUIT BREAKER] Terminating agent loop: ${reason}`);

    // 1. Park workspace changes recoverably.
    try {
      execSync('git stash push --include-untracked -m "agent-abort"', {
        cwd: this.workspaceDir,
        stdio: 'pipe',
      });
    } catch (err) {
      console.error('git stash failed during abort; workspace left as-is:', err);
    }

    // 2. Write a diagnostic trace for human review.
    const reportPath = this.writeFailureReport(reason, history);

    // 3. Signal the orchestrator.
    throw new AgentEscalationRequiredError(`Agent failed safely. Reason: ${reason}`, reportPath);
  }

  private writeFailureReport(reason: string, history: unknown[]): string {
    const reportPath = join(this.reportDir, `agent_failure_${Date.now()}.json`);
    writeFileSync(
      reportPath,
      JSON.stringify(
        {
          timestamp: new Date().toISOString(),
          workspace: this.workspaceDir,
          reason,
          historyLength: history.length,
          history: history.slice(-10),
        },
        null,
        2,
      ),
      'utf-8',
    );
    return reportPath;
  }
}

案例研究:开放式应用生成器中的全部四项防护措施

在案例研究中,将四个阶段视为可度量的整体来处理效果最佳。在扩大范围之前,先记录一个成功的案例、一个失败案例以及回滚说明。 将这一阶段视为输入与已验证输出之间的契约。为相关成果命名,明确成功标准,绝不允许出现悄无声息的半完成状态。 保持图表状态的简洁性与类型一致性。嵌套的数据块会掩盖哪个节点修改了哪个字段的信息,且在中断后会导致无法继续处理。

 ┌──────────────────────────────────────────────────────────┐
 │ Vague prompt ("build a modern web app locally")          │
 └────────────────────────────┬─────────────────────────────┘
                              ▼
 ┌──────────────────────────────────────────────────────────┐
 │ Phase 1: Dynamic spec synthesis (`ac-matrix.json`)       │
 └────────────────────────────┬─────────────────────────────┘
                              ▼
 ┌──────────────────────────────────────────────────────────┐
 │ Phase 2: Multi-agent execution loop                      │
 │ (Coder agent + design critic + headless E2E verifier)    │
 └────────────────────────────┬─────────────────────────────┘
                              │
    ┌─────────────────────────┼─────────────────────────┐
    ▼                         ▼                         ▼
┌──────────────────┐    ┌──────────────────┐    ┌──────────────────┐
│ Gate 1: Goal     │    │ Gate 2: Resource │    │ Gate 3: Progress │
│ verification     │    │ caps (turns/     │    │ guard (deadlock/ │
│ (build/E2E/ACs)  │    │ token budget)    │    │ repetition)      │
└─────────┬────────┘    └─────────┬────────┘    └─────────┬────────┘
          └───────────────────────┼───────────────────────┘
                                  ▼
 ┌──────────────────────────────────────────────────────────┐
 │ Gate 4: Safe exit OR circuit-breaker rollback            │
 └──────────────────────────────────────────────────────────┘

动态规范生成

将动态规范合成阶段视为可度量的对象时,其效果最佳。在扩大范围之前,需记录一份理想运行案例、一个故障案例以及回滚说明。 在功能结果旁同时记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在系统从演示环境过渡到共享环境时出现意外费用。 保持图结构扁平且类型明确。嵌套的数据块会掩盖哪个节点修改了哪个字段,还会在中断后导致流程无法继续。 将动态规范合成阶段视为可度量的对象时,其效果最佳。在扩大范围之前,需记录一份理想运行案例、一个故障案例以及回滚说明。 需同时记录正常流程与恢复流程。重试机制、人工审核环节以及死信处理都是产品本身的组成部分,而非后续需要补充的功能。

多智能体分工

在多智能体分工阶段,应在修改代码之前明确输入参数、该步骤的负责人以及终止条件。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。相比冗长的脚本,更应采用小型且可测试的单元。当某个步骤失败时,故障原因应能明确指向单一责任方,而非复杂的流程链。对于涉及资金支出或修改生产数据的操作,必须经过人工审批。编译时的逻辑连接并不等同于业务功能的完整性。

主控制框架

在主控制台阶段,修改代码之前需明确输入参数、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 将此阶段视为输入与已验证输出之间的契约。为相关成果命名,定义成功判定标准,并拒绝默许的半完成状态。 在网关处进行身份验证,在数据层再次授权。仅凭承载令牌并不足以界定租户边界。

import { execSync } from 'node:child_process';
import { readFileSync } from 'node:fs';
import { join } from 'node:path';
import { ResourceGuard, ProgressGuard, AgentCircuitBreaker } from './guards';

interface AcceptanceCriterion {
  id: string;
  description: string;
  status: 'PENDING' | 'IN_PROGRESS' | 'DONE';
}

interface AcMatrix {
  items: AcceptanceCriterion[];
}

export interface RunResult {
  success: true;
  turns: number;
  summary: string;
}

export async function runOpenEndedWebAppGenerator(
  userPrompt: string,
  workspacePath: string,
  reportDir: string,
  devServerUrl = 'http://localhost:5173',
  options = { maxTurns: 15, maxTotalTokens: 200_000, timeoutMs: 300_000 },
): Promise<RunResult> {
  const resources = new ResourceGuard(options);
  const progress = new ProgressGuard();
  const circuitBreaker = new AgentCircuitBreaker(workspacePath, reportDir);
  const acMatrixPath = join(workspacePath, 'ac-matrix.json');

  let currentPrompt = userPrompt;
  let designRetries = 0;
  const maxDesignRetries = 3;

  while (true) {
    // GUARD 1: resource ceilings
    const resourceCheck = resources.checkShouldTerminate();
    if (resourceCheck.terminate) {
      circuitBreaker.handleAbort(resourceCheck.reason!);
    }
    resources.beginTurn();

    const turnResult = await llmAgent.step(currentPrompt);
    resources.recordUsage(turnResult.tokensUsed);

    // GUARD 2: deadlock detection (only meaningful when a tool was called)
    let toolOutput = '(no tool call this turn)';
    if (turnResult.toolCall) {
      if (progress.trackAndCheckStuck(turnResult.toolCall)) {
        circuitBreaker.handleAbort('Repeating tool call detected (stuck agent)');
      }
      toolOutput = await executeTool(turnResult.toolCall);
    }

    // GUARD 3: deterministic convergence check
    const verification = await evaluateConvergence(workspacePath, acMatrixPath, devServerUrl);

    if (verification.gateFailed === 'design') {
      designRetries += 1;
      if (designRetries > maxDesignRetries) {
        circuitBreaker.handleAbort(
          `Design gate never converged after ${maxDesignRetries} refinement passes`,
        );
      }
    }

    if (verification.isConverged) {
      return {
        success: true,
        turns: resources.getTurnCount(),
        summary: 'Web app built, tested, and design-reviewed cleanly.',
      };
    }

    currentPrompt = `Tool output:\n${toolOutput}\n\nConvergence status:\n${verification.statusMessage}`;
  }
}

interface ConvergenceResult {
  isConverged: boolean;
  statusMessage: string;
  gateFailed?: 'build' | 'runtime' | 'acs' | 'design';
}

async function evaluateConvergence(
  workspacePath: string,
  acMatrixPath: string,
  devServerUrl: string,
): Promise<ConvergenceResult> {
  // Gate 1: build and typecheck
  try {
    execSync('npx tsc --noEmit && npm run build', { cwd: workspacePath, stdio: 'pipe' });
  } catch (err: unknown) {
    const e = err as { stdout?: Buffer; stderr?: Buffer };
    const output = [e.stdout?.toString(), e.stderr?.toString()].filter(Boolean).join('\n');
    return {
      isConverged: false,
      gateFailed: 'build',
      statusMessage: `Gate 1 failed (build/typecheck):\n${output || String(err)}`,
    };
  }

  // Gate 2: dev server and runtime health
  const e2eResult = await runHeadlessBrowserCheck(devServerUrl);
  if (!e2eResult.noConsoleErrors) {
    return {
      isConverged: false,
      gateFailed: 'runtime',
      statusMessage: `Gate 2 failed (console errors): ${e2eResult.errors.join(', ')}`,
    };
  }

  // Gate 3: acceptance criteria fully complete
  let acMatrix: AcMatrix;
  try {
    acMatrix = JSON.parse(readFileSync(acMatrixPath, 'utf-8')) as AcMatrix;
  } catch {
    return {
      isConverged: false,
      gateFailed: 'acs',
      statusMessage: 'Gate 3 incomplete: `ac-matrix.json` missing or unparseable.',
    };
  }

  const pending = acMatrix.items.filter((ac) => ac.status !== 'DONE');
  if (pending.length > 0) {
    return {
      isConverged: false,
      gateFailed: 'acs',
      statusMessage: `Gate 3 incomplete: ${pending.length} ACs remaining (${pending
        .map((a) => a.id)
        .join(', ')})`,
    };
  }

  // Gate 4: design audit (soft gate — see retry cap in the caller)
  const criticVerdict = await runDesignCriticAgent(e2eResult.screenshots);
  if (criticVerdict.score < 8.5) {
    return {
      isConverged: false,
      gateFailed: 'design',
      statusMessage: `Gate 4 incomplete (design ${criticVerdict.score}/10): ${criticVerdict.feedback}`,
    };
  }

  return { isConverged: true, statusMessage: 'All four convergence gates passed.' };
}

可直接使用的主提示词

在“即用型主提示词”阶段,修改代码之前需明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。应在功能结果旁记录执行时间以及令牌或查询成本。提前显示成本信息,可避免在从演示环境切换到共享环境时出现意外费用。当下一步操作为编写代码或调用工具时,应优先选择具有结构化格式且经过模式验证的输出,而非自由形式的文本。在“即用型主提示词”阶段,修改代码之前需明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。需同时记录正常流程和异常恢复流程。重试机制、人工审核环节以及错误处理方式都是产品本身的一部分,而非后续需要补充的内容。

You are an autonomous lead software engineer, UX designer, and QA verifier. Your goal
is to build a production-quality web application locally, from scratch.

You must operate in a self-terminating agentic loop, running iteratively until the
application is complete, polished, functional, and verified.

================================================================================
1. TARGET APPLICATION SPECIFICATION
================================================================================

[DESCRIBE YOUR APP IDEA HERE — e.g. "A task management web app with local SQLite
persistence, a kanban board with drag-and-drop, priority tags, search/filter
controls, and a dark mode theme."]

================================================================================
2. EXECUTION PROTOCOL
================================================================================

PHASE 1 — DYNAMIC SPEC SYNTHESIS (TURN 1)
Before writing application code or installing dependencies:
  1. Initialize the local project structure (e.g. Vite + React, Next.js, or Node).
  2. Create `ac-matrix.json` in the workspace root defining explicit acceptance
     criteria:
       - Feature ACs: persistence, full CRUD, interactive components, error
         handling, edge cases.
       - Engineering ACs: strict TypeScript (`npx tsc --noEmit`), zero build
         errors, zero linter warnings, dev server boots cleanly.
       - Design ACs: visual hierarchy, responsive layout, dark/light toggle,
         empty states, micro-interactions.

Format:
  {
    "project": "<app-name>",
    "items": [
      { "id": "FEAT-1", "category": "feature",     "description": "Local database persistence for tasks", "status": "PENDING" },
      { "id": "FEAT-2", "category": "feature",     "description": "Drag-and-drop kanban re-ordering",      "status": "PENDING" },
      { "id": "ENG-1",  "category": "engineering", "description": "Clean TypeScript build, zero errors",   "status": "PENDING" },
      { "id": "ENG-2",  "category": "engineering", "description": "Dev server starts with 0 console errors","status": "PENDING" },
      { "id": "DSGN-1", "category": "design",      "description": "Responsive UI with dark/light mode",    "status": "PENDING" }
    ]
  }

Once written, treat `ac-matrix.json` as frozen scope. Do not delete or weaken an
AC to make a gate pass. If an AC turns out to be genuinely infeasible, mark it
BLOCKED with a reason and surface it in the final summary.

PHASE 2 — AUTONOMOUS DEVELOPMENT LOOP
In each turn:
  - Implement features, components, schemas, and routes incrementally.
  - Run local validation after code changes (`npx tsc --noEmit`, `npm run build`).
  - Update AC statuses (PENDING → IN_PROGRESS → DONE) as work is verified.
  - Self-correction rule: if a command fails, read the exact error, fix the root
    cause, and re-verify. Do not repeat the same failing command or edit more
    than twice — change approach instead.

PHASE 3 — THE 4-GATE CONVERGENCE CHECK (MANDATORY)
Do not end execution or declare the project finished until all four gates pass
in the same turn:

Gate 1 — Compiler and build
    `npx tsc --noEmit` and `npm run build` both exit 0 with zero errors.

Gate 2 — Dev server and runtime health
    `npm run dev` boots cleanly with zero unhandled console or network errors.

Gate 3 — Acceptance criteria complete
    Every item in `ac-matrix.json` has "status": "DONE".

Gate 4 — Design audit score >= 8.5/10
    Audit visual hierarchy, color consistency, typography scale, spacing,
    transitions, responsive behavior, and empty states. Score out of 10. If
    below 8.5, refine and re-audit — but no more than 3 design passes total.
    After 3 passes, stop and report the final score as-is.

================================================================================
3. FINAL COMPLETION OUTPUT
================================================================================

Only when all four gates pass, output:
  - Final status: PROJECT COMPLETE & VERIFIED
  - App summary and architecture overview
  - Build and test commands executed, with exit codes
  - Completed acceptance-criteria summary (including any BLOCKED items)
  - Final design score (X/10) and UX highlights
  - Instructions for running the app locally

Then stop.

总结检查清单

在处理总结检查清单阶段时,首先写下合同条款:所需的输入参数、成功标志以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 优先选择小型、可测试的单元,而非庞大的脚本。当某个步骤失败时,故障应指向单一的责任模块,而非复杂的流程链。 在成本较高的步骤之后设置检查点。当操作员重新尝试后续节点时,恢复流程不应再次调用相同的大型语言模型。

结论

在进入结论阶段时,首先写下相关契约:所需的输入参数、成功标志以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持一致性。

操作检查清单

在处理操作检查清单阶段时,同样要先明确契约内容:所需输入、成功标志以及部分失败时的应对措施。这样的清单有助于保证后续代码修改的准确性。

应将配置信息与应用程序代码分开存放。环境文件、密钥存储以及功能开关应集中于一个位置,这样操作人员无需查看整个系统结构即可进行审计。

在成本较高的操作之后设置检查点。当操作员重新尝试后续节点时,恢复流程不应再次对同一次 LLM 调用收费。

锁定依赖版本,并记录用于运行演示的镜像摘要。可重复性比经验知识更重要。

同时记录正常流程和故障恢复流程。重试机制、人工审核环节以及死信处理都是产品的一部分,而非后续需要补充的功能。

在成本较高的操作之后设置检查点。当操作员重新尝试后续节点时,恢复流程不应再次对同一次 LLM 调用收费。

在升级技术栈之前,先冻结版本,为关键流程保存标准操作记录,并确认回滚步骤。共享环境需要设置速率限制、进行租户检查,同时明确密钥轮换的负责人。与其追求花哨的一次性演示,不如注重扎实的可靠性。

c09d8d68f871的批处理说明:不要将提供者密钥放入仓库中,为每个会话设置令牌上限,并将转录内容存储在评估测试用例的旁边,以便后续更换模型时仍能保持可比性。