This article is published in English.
Practical notes: Don’t Let Your AI Agents Loop Forever: An Engineering Guide to
Operable walkthrough of Practical notes: Don’t Let Your AI Agents Loop Forever: An Engineering Guide to: contracts, checks, and drop-in code slots for teams shipping this pattern.
Use this as an operator-facing rebuild of the ideas in “Don’t Let Your AI Agents Loop Forever: An Engineering Guide to Termination Criteria”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
The Friday Afternoon Nightmare
For the The Friday Afternoon Nightmare stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Anatomy of an Agentic Loop
For the Anatomy of an Agentic stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
┌──────────────────────────────────────┐
│ Agent Perception Loop │
│ (Perceive → Plan → Act) │
└──────────────────┬───────────────────┘
│
┌──────────────────────────┼──────────────────────────┐
▼ ▼ ▼
┌────────────────────┐ ┌────────────────────┐ ┌────────────────────┐
│ 1. Success Guard │ │ 2. Resource Caps │ │ 3. Progress Guard │
│ (Programmatic Test)│ │ (Tokens/Turns/Time)│ │ (Loop/Hash Detect) │
└────────────────────┘ └────────────────────┘ └────────────────────┘
│
▼
┌──────────────────────────────────────┐
│ 4. Human Handoff / Safe Rollback │
└──────────────────────────────────────┘
1. Success Criteria: Deterministic Goal Verification
For the 1 Success Criteria Deterministic stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
The pitfall: self-assessment
For the The pitfall self-assessment stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
The solution: external programmatic verifiers
For the The solution external programmatic stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
import { execSync } from 'node:child_process';
export interface VerificationResult {
success: boolean;
message: string;
stepFailed?: string;
}
/** execSync throws on non-zero exit; diagnostics may land on either stream. */
function runOrCapture(command: string, cwd: string): string | null {
try {
execSync(command, { cwd, stdio: 'pipe' });
return null;
} catch (err: unknown) {
const e = err as { stdout?: Buffer; stderr?: Buffer };
const out = e.stdout?.toString() ?? '';
const errOut = e.stderr?.toString() ?? '';
return [out, errOut].filter(Boolean).join('\n') || String(err);
}
}
export class GoalVerifier {
public static verify(workspacePath: string): VerificationResult {
const steps: Array<[string, string]> = [
['tsc', 'npx tsc --noEmit'],
['npm_test', 'npm test'],
];
for (const [stepId, command] of steps) {
const failure = runOrCapture(command, workspacePath);
if (failure !== null) {
return {
success: false,
message: `\`${command}\` failed:\n${failure}`,
stepFailed: stepId,
};
}
}
return { success: true, message: 'All typechecks and tests passed cleanly.' };
}
}
2. Resource and Budget Caps: Hard Engine Ceilings
For the 2 Resource and Budget stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
The pitfall: unbounded retries
For the The pitfall unbounded retries stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
The solution: multi-dimensional limits
For the The solution multi-dimensional limits stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the The solution multi-dimensional limits stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
export interface ResourceLimits {
maxTurns: number; // e.g. 10 iterations
maxTotalTokens: number; // e.g. 100_000 input + output
timeoutMs: number; // e.g. 120_000 (2 minutes)
}
export class ResourceGuard {
private readonly startTime = Date.now();
private totalTokensUsed = 0;
private currentTurn = 0;
constructor(private readonly limits: ResourceLimits) {}
/** Call once per loop iteration, before the model call. */
public beginTurn(): void {
this.currentTurn += 1;
}
/** Call for every model call, including retries inside a turn. */
public recordUsage(tokens: number): void {
this.totalTokensUsed += tokens;
}
public getTurnCount(): number {
return this.currentTurn;
}
public checkShouldTerminate(): { terminate: boolean; reason?: string } {
if (this.currentTurn >= this.limits.maxTurns) {
return {
terminate: true,
reason: `Exceeded turn cap (${this.limits.maxTurns})`,
};
}
if (this.totalTokensUsed >= this.limits.maxTotalTokens) {
return {
terminate: true,
reason: `Exceeded token budget (${this.totalTokensUsed}/${this.limits.maxTotalTokens})`,
};
}
const elapsed = Date.now() - this.startTime;
if (elapsed >= this.limits.timeoutMs) {
return {
terminate: true,
reason: `Wall-clock timeout reached (${elapsed}ms/${this.limits.timeoutMs}ms)`,
};
}
return { terminate: false };
}
}
3. Progress Guards: Stuck and Drift Detection
When working through the 3 Progress Guards Stuck stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Common failure modes
When working through the Common failure modes stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
The solution: tool signatures and workspace state hashing
When working through the The solution tool signatures stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours. When working through the The solution tool signatures stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
import { createHash } from 'node:crypto';
export interface ToolCall {
name: string;
args: Record<string, unknown>;
}
export class ProgressGuard {
private readonly recentActionHashes: string[] = [];
constructor(
private readonly windowSize = 5,
private readonly repeatThreshold = 3,
) {}
/** Stable stringify: key order must not change the hash. */
private hashToolCall(call: ToolCall): string {
const args = JSON.stringify(call.args, Object.keys(call.args).sort());
return createHash('sha256').update(`${call.name}:${args}`).digest('hex');
}
/** Returns true when the same call has appeared `repeatThreshold` times in the window. */
public trackAndCheckStuck(call: ToolCall): boolean {
const actionHash = this.hashToolCall(call);
const priorOccurrences = this.recentActionHashes.filter((h) => h === actionHash).length;
this.recentActionHashes.push(actionHash);
if (this.recentActionHashes.length > this.windowSize) {
this.recentActionHashes.shift();
}
return priorOccurrences + 1 >= this.repeatThreshold;
}
}
4. Human-in-the-Loop and Safe Rollbacks
The 4 Human-in-the-Loop and Safe stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
import { execSync } from 'node:child_process';
import { writeFileSync } from 'node:fs';
import { join } from 'node:path';
export class AgentEscalationRequiredError extends Error {
constructor(message: string, public readonly reportPath?: string) {
super(message);
this.name = 'AgentEscalationRequiredError';
}
}
export class AgentCircuitBreaker {
constructor(
private readonly workspaceDir: string,
private readonly reportDir: string, // keep reports OUTSIDE the workspace
) {}
public handleAbort(reason: string, history: unknown[] = []): never {
console.error(`[CIRCUIT BREAKER] Terminating agent loop: ${reason}`);
// 1. Park workspace changes recoverably.
try {
execSync('git stash push --include-untracked -m "agent-abort"', {
cwd: this.workspaceDir,
stdio: 'pipe',
});
} catch (err) {
console.error('git stash failed during abort; workspace left as-is:', err);
}
// 2. Write a diagnostic trace for human review.
const reportPath = this.writeFailureReport(reason, history);
// 3. Signal the orchestrator.
throw new AgentEscalationRequiredError(`Agent failed safely. Reason: ${reason}`, reportPath);
}
private writeFailureReport(reason: string, history: unknown[]): string {
const reportPath = join(this.reportDir, `agent_failure_${Date.now()}.json`);
writeFileSync(
reportPath,
JSON.stringify(
{
timestamp: new Date().toISOString(),
workspace: this.workspaceDir,
reason,
historyLength: history.length,
history: history.slice(-10),
},
null,
2,
),
'utf-8',
);
return reportPath;
}
}
Case Study: All Four Guardrails in an Open-Ended App Generator
The Case Study All Four stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
┌──────────────────────────────────────────────────────────┐
│ Vague prompt ("build a modern web app locally") │
└────────────────────────────┬─────────────────────────────┘
▼
┌──────────────────────────────────────────────────────────┐
│ Phase 1: Dynamic spec synthesis (`ac-matrix.json`) │
└────────────────────────────┬─────────────────────────────┘
▼
┌──────────────────────────────────────────────────────────┐
│ Phase 2: Multi-agent execution loop │
│ (Coder agent + design critic + headless E2E verifier) │
└────────────────────────────┬─────────────────────────────┘
│
┌─────────────────────────┼─────────────────────────┐
▼ ▼ ▼
┌──────────────────┐ ┌──────────────────┐ ┌──────────────────┐
│ Gate 1: Goal │ │ Gate 2: Resource │ │ Gate 3: Progress │
│ verification │ │ caps (turns/ │ │ guard (deadlock/ │
│ (build/E2E/ACs) │ │ token budget) │ │ repetition) │
└─────────┬────────┘ └─────────┬────────┘ └─────────┬────────┘
└───────────────────────┼───────────────────────┘
▼
┌──────────────────────────────────────────────────────────┐
│ Gate 4: Safe exit OR circuit-breaker rollback │
└──────────────────────────────────────────────────────────┘
Dynamic spec synthesis
The Dynamic spec synthesis stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Dynamic spec synthesis stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Multi-agent division of labor
For the Multi-agent division of labor stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
The master harness
For the The master harness stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.
import { execSync } from 'node:child_process';
import { readFileSync } from 'node:fs';
import { join } from 'node:path';
import { ResourceGuard, ProgressGuard, AgentCircuitBreaker } from './guards';
interface AcceptanceCriterion {
id: string;
description: string;
status: 'PENDING' | 'IN_PROGRESS' | 'DONE';
}
interface AcMatrix {
items: AcceptanceCriterion[];
}
export interface RunResult {
success: true;
turns: number;
summary: string;
}
export async function runOpenEndedWebAppGenerator(
userPrompt: string,
workspacePath: string,
reportDir: string,
devServerUrl = 'http://localhost:5173',
options = { maxTurns: 15, maxTotalTokens: 200_000, timeoutMs: 300_000 },
): Promise<RunResult> {
const resources = new ResourceGuard(options);
const progress = new ProgressGuard();
const circuitBreaker = new AgentCircuitBreaker(workspacePath, reportDir);
const acMatrixPath = join(workspacePath, 'ac-matrix.json');
let currentPrompt = userPrompt;
let designRetries = 0;
const maxDesignRetries = 3;
while (true) {
// GUARD 1: resource ceilings
const resourceCheck = resources.checkShouldTerminate();
if (resourceCheck.terminate) {
circuitBreaker.handleAbort(resourceCheck.reason!);
}
resources.beginTurn();
const turnResult = await llmAgent.step(currentPrompt);
resources.recordUsage(turnResult.tokensUsed);
// GUARD 2: deadlock detection (only meaningful when a tool was called)
let toolOutput = '(no tool call this turn)';
if (turnResult.toolCall) {
if (progress.trackAndCheckStuck(turnResult.toolCall)) {
circuitBreaker.handleAbort('Repeating tool call detected (stuck agent)');
}
toolOutput = await executeTool(turnResult.toolCall);
}
// GUARD 3: deterministic convergence check
const verification = await evaluateConvergence(workspacePath, acMatrixPath, devServerUrl);
if (verification.gateFailed === 'design') {
designRetries += 1;
if (designRetries > maxDesignRetries) {
circuitBreaker.handleAbort(
`Design gate never converged after ${maxDesignRetries} refinement passes`,
);
}
}
if (verification.isConverged) {
return {
success: true,
turns: resources.getTurnCount(),
summary: 'Web app built, tested, and design-reviewed cleanly.',
};
}
currentPrompt = `Tool output:\n${toolOutput}\n\nConvergence status:\n${verification.statusMessage}`;
}
}
interface ConvergenceResult {
isConverged: boolean;
statusMessage: string;
gateFailed?: 'build' | 'runtime' | 'acs' | 'design';
}
async function evaluateConvergence(
workspacePath: string,
acMatrixPath: string,
devServerUrl: string,
): Promise<ConvergenceResult> {
// Gate 1: build and typecheck
try {
execSync('npx tsc --noEmit && npm run build', { cwd: workspacePath, stdio: 'pipe' });
} catch (err: unknown) {
const e = err as { stdout?: Buffer; stderr?: Buffer };
const output = [e.stdout?.toString(), e.stderr?.toString()].filter(Boolean).join('\n');
return {
isConverged: false,
gateFailed: 'build',
statusMessage: `Gate 1 failed (build/typecheck):\n${output || String(err)}`,
};
}
// Gate 2: dev server and runtime health
const e2eResult = await runHeadlessBrowserCheck(devServerUrl);
if (!e2eResult.noConsoleErrors) {
return {
isConverged: false,
gateFailed: 'runtime',
statusMessage: `Gate 2 failed (console errors): ${e2eResult.errors.join(', ')}`,
};
}
// Gate 3: acceptance criteria fully complete
let acMatrix: AcMatrix;
try {
acMatrix = JSON.parse(readFileSync(acMatrixPath, 'utf-8')) as AcMatrix;
} catch {
return {
isConverged: false,
gateFailed: 'acs',
statusMessage: 'Gate 3 incomplete: `ac-matrix.json` missing or unparseable.',
};
}
const pending = acMatrix.items.filter((ac) => ac.status !== 'DONE');
if (pending.length > 0) {
return {
isConverged: false,
gateFailed: 'acs',
statusMessage: `Gate 3 incomplete: ${pending.length} ACs remaining (${pending
.map((a) => a.id)
.join(', ')})`,
};
}
// Gate 4: design audit (soft gate — see retry cap in the caller)
const criticVerdict = await runDesignCriticAgent(e2eResult.screenshots);
if (criticVerdict.score < 8.5) {
return {
isConverged: false,
gateFailed: 'design',
statusMessage: `Gate 4 incomplete (design ${criticVerdict.score}/10): ${criticVerdict.feedback}`,
};
}
return { isConverged: true, statusMessage: 'All four convergence gates passed.' };
}
Ready-to-Use Master Prompt
For the Ready-to-Use Master Prompt stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call. For the Ready-to-Use Master Prompt stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
You are an autonomous lead software engineer, UX designer, and QA verifier. Your goal
is to build a production-quality web application locally, from scratch.
You must operate in a self-terminating agentic loop, running iteratively until the
application is complete, polished, functional, and verified.
================================================================================
1. TARGET APPLICATION SPECIFICATION
================================================================================
[DESCRIBE YOUR APP IDEA HERE — e.g. "A task management web app with local SQLite
persistence, a kanban board with drag-and-drop, priority tags, search/filter
controls, and a dark mode theme."]
================================================================================
2. EXECUTION PROTOCOL
================================================================================
PHASE 1 — DYNAMIC SPEC SYNTHESIS (TURN 1)
Before writing application code or installing dependencies:
1. Initialize the local project structure (e.g. Vite + React, Next.js, or Node).
2. Create `ac-matrix.json` in the workspace root defining explicit acceptance
criteria:
- Feature ACs: persistence, full CRUD, interactive components, error
handling, edge cases.
- Engineering ACs: strict TypeScript (`npx tsc --noEmit`), zero build
errors, zero linter warnings, dev server boots cleanly.
- Design ACs: visual hierarchy, responsive layout, dark/light toggle,
empty states, micro-interactions.
Format:
{
"project": "<app-name>",
"items": [
{ "id": "FEAT-1", "category": "feature", "description": "Local database persistence for tasks", "status": "PENDING" },
{ "id": "FEAT-2", "category": "feature", "description": "Drag-and-drop kanban re-ordering", "status": "PENDING" },
{ "id": "ENG-1", "category": "engineering", "description": "Clean TypeScript build, zero errors", "status": "PENDING" },
{ "id": "ENG-2", "category": "engineering", "description": "Dev server starts with 0 console errors","status": "PENDING" },
{ "id": "DSGN-1", "category": "design", "description": "Responsive UI with dark/light mode", "status": "PENDING" }
]
}
Once written, treat `ac-matrix.json` as frozen scope. Do not delete or weaken an
AC to make a gate pass. If an AC turns out to be genuinely infeasible, mark it
BLOCKED with a reason and surface it in the final summary.
PHASE 2 — AUTONOMOUS DEVELOPMENT LOOP
In each turn:
- Implement features, components, schemas, and routes incrementally.
- Run local validation after code changes (`npx tsc --noEmit`, `npm run build`).
- Update AC statuses (PENDING → IN_PROGRESS → DONE) as work is verified.
- Self-correction rule: if a command fails, read the exact error, fix the root
cause, and re-verify. Do not repeat the same failing command or edit more
than twice — change approach instead.
PHASE 3 — THE 4-GATE CONVERGENCE CHECK (MANDATORY)
Do not end execution or declare the project finished until all four gates pass
in the same turn:
Gate 1 — Compiler and build
`npx tsc --noEmit` and `npm run build` both exit 0 with zero errors.
Gate 2 — Dev server and runtime health
`npm run dev` boots cleanly with zero unhandled console or network errors.
Gate 3 — Acceptance criteria complete
Every item in `ac-matrix.json` has "status": "DONE".
Gate 4 — Design audit score >= 8.5/10
Audit visual hierarchy, color consistency, typography scale, spacing,
transitions, responsive behavior, and empty states. Score out of 10. If
below 8.5, refine and re-audit — but no more than 3 design passes total.
After 3 passes, stop and report the final score as-is.
================================================================================
3. FINAL COMPLETION OUTPUT
================================================================================
Only when all four gates pass, output:
- Final status: PROJECT COMPLETE & VERIFIED
- App summary and architecture overview
- Build and test commands executed, with exit codes
- Completed acceptance-criteria summary (including any BLOCKED items)
- Final design score (X/10) and UX highlights
- Instructions for running the app locally
Then stop.
Summary Checklist
When working through the Summary Checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Conclusion
When working through the Conclusion stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Operational checklist
When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.
Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for c09d8d68f871: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.