Home / Articles / Practical notes: Loop Engineering: Building Self-Improving AI Agents with Four

This article is published in English.

Practical notes: Loop Engineering: Building Self-Improving AI Agents with Four

Operable walkthrough of Practical notes: Loop Engineering: Building Self-Improving AI Agents with Four: contracts, checks, and drop-in code slots for teams shipping this pattern.

7704 words

Use this as an operator-facing rebuild of the ideas in “Loop Engineering: Building Self-Improving AI Agents with Four Nested Loops”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

The Big Picture Breakdown: Engineering Beyond the Prompt

For the The Big Picture Breakdown stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

Why This Matters

For the Why This Matters stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

The Core Idea: Nested Control Loops

For the The Core Idea Nested stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the The Core Idea Nested stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

┌──────────────────────────────────────────────────────┐
│  Loop 4 — Hill Climbing (self-improvement over time) │
│  ┌────────────────────────────────────────────────┐  │
│  │  Loop 3 — Event Queue Poller (orchestration)   │  │
│  │  ┌──────────────────────────────────────────┐  │  │
│  │  │  Loop 2 — Verification & Retry           │  │  │
│  │  │  ┌────────────────────────────────────┐  │  │  │
│  │  │  │  Loop 1 — ReAct Agent              │  │  │  │
│  │  │  │  (think → act → observe → think)   │  │  │  │
│  │  │  └────────────────────────────────────┘  │  │  │
│  │  └──────────────────────────────────────────┘  │  │
│  └────────────────────────────────────────────────┘  │
└──────────────────────────────────────────────────────┘

The Domain: Insurance Underwriting

When working through the The Domain Insurance Underwriting stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

{
  "id": "APP-001",
  "applicant_age": 34,
  "state": "FL",
  "property_type": "residential",
  "coverage_amount": 450000,
  "claim_history_5yr": 1,
  "credit_score_tier": "good",
  "flood_zone": true,
  "has_flood_rider": false,
  "business_use": false
}
"APP-001": {
  "split": "train",
  "expected_decision": "modify",
  "expected_flags": ["flood_rider_required", "windstorm_exclusion"]
}

Loop 1: The ReAct Agent

When working through the Loop 1 The ReAct stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

The Pattern

When working through the The Pattern stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the The Pattern stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

System Prompt
     │
     ▼
[Think] → What information do I need?
     │
     ▼
[Act]   → Call tool (lookup_application, retrieve_rules, ...)
     │
     ▼
[Observe] → Tool returns result
     │
     ▼
[Think] → What does this mean? What next?
     │
    ... (repeat until decision is made)
     │
     ▼
[Act]   → draft_decision (terminal tool call)

The LangGraph Implementation

The The LangGraph Implementation stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

class UnderwritingState(TypedDict):
    application_id: str
    messages: Annotated[list[AnyMessage], add_messages]  # reducer accumulates
    decision: str
    rationale: str
    flags: list[str]
    recommended_premium_adj: float
    lessons: list[str]
    spend: float
    _steps: int
def build_loop1(system_prompt: str, meter: CostMeter):
    model = ChatAnthropic(model="claude-sonnet-4-6").bind_tools(TOOLS)
    def agent_node(state: UnderwritingState) -> dict:
        prompt = system_prompt
        if state.get("lessons"):
            prompt += "\n\nUnderwriting lessons:\n" + "\n".join(
                f"- {l}" for l in state["lessons"]
            )
        msgs = [SystemMessage(content=prompt)] + list(state["messages"])
        response = model.invoke(msgs)
        meter.record_from_message(response)          # track spend per call
        return {
            "messages": [response],
            "_steps": state.get("_steps", 0) + 1,
            "spend": meter.spent,
        }    def tools_node(state: UnderwritingState) -> dict:
        last = state["messages"][-1]
        tool_results, updates = _dispatch_tools(last.tool_calls)
        return {"messages": tool_results, **updates}  # updates captures draft_decision output    def should_continue(state) -> Literal["tools", "__end__"]:
        if state.get("_steps", 0) >= MAX_STEPS:   # hard cap at 8 steps
            return END
        if getattr(state["messages"][-1], "tool_calls", None):
            return "tools"
        return END    g = StateGraph(UnderwritingState)
    g.add_node("agent", agent_node)
    g.add_node("tools", tools_node)
    g.set_entry_point("agent")
    g.add_conditional_edges("agent", should_continue, {"tools": "tools", END: END})
    g.add_edge("tools", "agent")
    return g.compile()

The Four Tools

The The Four Tools stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

@tool
def lookup_application(app_id: str) -> dict:
    """Return the full application record for the given application ID."""
    # reads from applications.json — deterministic, no LLM call
@tool
def retrieve_rules(risk_factor: str) -> list[str]:
    """Return underwriting rules matching the given risk factor keyword.
    Use terms like: flood, claims, credit, coverage, state, business."""
    # keyword match against rules_kb.json - forces explicit rule lookup
@tool
def fetch_risk_signal(signal_type: str) -> dict:
    # simulated external data pull (credit bureaus, flood maps, etc.)
@tool
def draft_decision(
    decision: Literal["accept", "modify", "decline", "refer"],
    rationale: str,
    flags: list[str],
    recommended_premium_adj: float,
) -> dict:
    """Finalize the underwriting decision with structured output."""
    # terminal action - structured schema forces the agent to commit explicitly

Why the Step Cap Matters

The Why the Step Cap stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Why the Step Cap stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Loop 2: Verification and Grader-Based Retry

For the Loop 2 Verification and stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

def run_loop2(
    app_id: str,
    system_prompt: str,
    meter: CostMeter,
    max_retries: int = 3,
) -> tuple[dict, bool]:
    state = run_loop1(app_id, system_prompt, meter)
for attempt in range(max_retries):
        decision = state.get("decision", "ran_out_of_steps")
        if decision == "refer":
            _write_human_queue(app_id, state)   # human-in-the-loop path
            return state, True
        if is_correct(decision, app_id):         # deterministic check
            return state, True
        if attempt >= max_retries - 1:
            return state, False                  # exhausted retries
        # construct targeted feedback for next attempt
        expected = load_historical_decisions()[app_id]["expected_decision"]
        feedback = HumanMessage(content=(
            f"Your decision was '{decision}' but this application requires '{expected}'. "
            f"Review the risk factors carefully and call draft_decision again."
        ))
        extra_messages = list(state.get("messages", [])) + [feedback]
        state = run_loop1(app_id, system_prompt, meter, extra_messages=extra_messages)
    return state, False

Design Decisions Worth Noting

For the Design Decisions Worth Noting stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Loop 3: The Event Queue Poller

For the Loop 3 The Event stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Loop 3 The Event stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

def run_event_loop(
    system_prompt: str,
    meter: CostMeter,
    pending_path: Path = DEFAULT_PENDING,
    state_path: Path = DEFAULT_STATE,
    once: bool = False,
) -> None:
    run_state = load_run_state(state_path)
    verify_judge_checksum(run_state.get("judge_checksum", ""))   # tamper check
total = len(json.loads(pending_path.read_text(encoding="utf-8-sig"))
                if pending_path.exists() else [])
    processed = 0
    while True:
        pending = (json.loads(pending_path.read_text(encoding="utf-8-sig"))
                   if pending_path.exists() else [])
        if not pending:
            if once:
                print("[loop3] queue empty - done")
                break
            print("[loop3] queue empty - sleeping...")
            time.sleep(POLL_INTERVAL)
            continue
        app_id = pending[0]
        processed += 1
        print(f"[loop3] {processed}/{total}  {app_id} ...", flush=True)
        try:
            result_state, passed = run_loop2(app_id, prompt_with_lessons, meter)
        except BudgetExhaustedError:
            raise                                # budget exhaustion is fatal
        except Exception as exc:
            print(f"[loop3] {app_id} ERROR: {exc!r} - skipping")
            # remove from queue and continue - one bad app cannot block the rest
            remaining = [x for x in pending if x != app_id]
            pending_path.write_text(json.dumps(remaining), encoding="utf-8")
            continue
        # ... update state, distill lesson, trigger hill-climbing

Why File-Based Queue?

When working through the Why File-Based Queue stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Tamper Detection via Checksum

When working through the Tamper Detection via Checksum stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

verify_judge_checksum(run_state.get("judge_checksum", ""))
def compute_judge_checksum() -> str:
    content = _HIST.read_bytes()           # historical_decisions.json
    return hashlib.sha256(content).hexdigest()
def verify_judge_checksum(expected: str) -> None:
    actual = compute_judge_checksum()
    if actual != expected:
        raise RuntimeError(
            f"Judge checksum mismatch - historical_decisions.json was tampered with."
        )

Skipping vs Crashing

When working through the Skipping vs Crashing stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Skipping vs Crashing stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Loop 4: Hill Climbing (Self-Improvement)

The Loop 4 Hill Climbing stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

seed prompt
     │
     ▼
evaluate on training set → score, failures
     │
     ▼
┌────────────────────────────────┐
│  for each round:               │
│                                │
│  reflect(current, failures)    │  ← LLM analyzes failure patterns
│       │                        │
│       ▼                        │
│  candidate_prompt              │
│       │                        │
│       ▼                        │
│  evaluate(candidate, train)    │  ← full eval run
│       │                        │
│       ▼                        │
│  audit_prompt(candidate)       │  ← anti-cheating check
│       │                        │
│       ▼                        │
│  if score > best AND clean:    │
│      best = candidate          │  ← keep
│  else:                         │
│      discard                   │  ← revert
└────────────────────────────────┘
     │
     ▼
write best_prompt.txt

The Reflect Step

The The Reflect Step stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

REFLECT_TEMPLATE = """You are improving the system prompt of an insurance underwriting agent.CURRENT PROMPT:
{prompt}
The agent got these decisions WRONG on training applications:
{failures}
Look for PATTERNS in the failures. Infer GENERAL underwriting rules that would fix them.
Do NOT memorize specific application IDs or applicant details.
Write an improved system prompt. Reply with ONLY the new prompt text."""
def reflect(current_prompt: str, failures: list[dict]) -> str:
    client = anthropic.Anthropic()
    failure_text = "\n".join(
        f"- APP {f['app_id']}: agent said '{f['got']}', expected '{f['expected']}'. "
        f"Application data: {f['application']}"
        for f in failures
    )
    response = client.messages.create(
        model="claude-sonnet-4-6",
        max_tokens=1500,
        messages=[{"role": "user", "content": REFLECT_TEMPLATE.format(
            prompt=current_prompt,
            failures=failure_text,
        )}],
    )
    return response.content[0].text.strip()

The Evaluate Step

The The Evaluate Step stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

def evaluate_prompt(system_prompt: str, split: str = "train", meter=None):
    tasks = load_tasks(split)
    failures = []
    for task in tasks:
        state = run_loop1(task["id"], system_prompt, meter)
        decision = state.get("decision", "ran_out_of_steps")
        if not is_correct(decision, task["id"]):
            failures.append({
                "app_id": task["id"],
                "got": decision,
                "expected": task["expected_decision"],
                "application": task["application"],
            })
    score = (len(tasks) - len(failures)) / len(tasks)
    return score, failures

The The Evaluate Step stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

The Gate: Anti-Overfitting and Anti-Cheating

For the The Gate Anti-Overfitting and stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

def improve(seed_prompt: str, lessons: list[str], meter: CostMeter, rounds: int = 5) -> str:
    best = with_lessons(seed_prompt, lessons)
    best_score, failures = evaluate_prompt(best, split="train", meter=meter)
for r in range(1, rounds + 1):
        candidate = reflect(best, failures)
        score, cand_failures = evaluate_prompt(candidate, split="train", meter=meter)
        verdict = audit_prompt(candidate)        # anti-cheat check
        if score > best_score and verdict == "VERDICT: CLEAN":
            best, best_score, failures = candidate, score, cand_failures
            BEST_PROMPT_PATH.write_text(best, encoding="utf-8")
            print(f"round {r}: {score:.0%} KEPT  ({len(cand_failures)} failures)")
        else:
            reason = "cheating" if verdict != "VERDICT: CLEAN" else f"{score:.0%} <= {best_score:.0%}"
            print(f"round {r}: {score:.0%} REVERTED ({reason})")
    return best

The Prompt Auditor

For the The Prompt Auditor stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

AUDIT_SYSTEM = (
    "You are a security auditor for AI systems. "
    "Check whether the prompt contains any hardcoded application IDs (e.g. APP-001) "
    "or hardcoded decisions for specific applications. "
    "These would constitute cheating — memorizing answers instead of learning rules. "
    "Reply with exactly one of:\n"
    "VERDICT: CLEAN\n"
    "VERDICT: CHEATING\n"
    "Followed by a one-sentence explanation."
)
def audit_prompt(evolved_prompt: str) -> str:
    response = client.messages.create(
        model="claude-sonnet-4-6",
        max_tokens=200,
        system=AUDIT_SYSTEM,
        messages=[{"role": "user", "content": f"PROMPT TO AUDIT:\n{evolved_prompt}"}],
    )
    return response.content[0].text.strip().split("\n")[0]   # first line only

Lesson Memory: Cross-Session Knowledge Distillation

For the Lesson Memory Cross-Session Knowledge stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Lesson Memory Cross-Session Knowledge stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

DISTILL_PROMPT = """An insurance underwriting agent made a wrong decision.Application details: {application}
Agent decided: {got}
Correct decision: {expected}
Write ONE short, general underwriting rule that would prevent this mistake.
Do NOT reference this specific application ID or any applicant names/details.
Reply with only the rule, as a single sentence."""
def distill_lesson(application: dict, got: str, expected: str) -> str:
    response = client.messages.create(
        model="claude-sonnet-4-6",
        max_tokens=150,
        messages=[{"role": "user", "content": DISTILL_PROMPT.format(...)}],
    )
    return response.content[0].text.strip()
def with_lessons(base_prompt: str, lessons: list[str]) -> str:
    if not lessons:
        return base_prompt
    lessons_text = "\n".join(f"- {l}" for l in lessons)
    return base_prompt + f"\n\nUnderwriting lessons learned:\n{lessons_text}"
def measure_gain(base_prompt: str, lessons: list[str], split: str = "test", meter=None) -> float:
    stateless_score, _ = evaluate_prompt(base_prompt, split=split, meter=meter)
    augmented = with_lessons(base_prompt, lessons)
    stateful_score, _ = evaluate_prompt(augmented, split=split, meter=meter)
    return stateful_score - stateless_score   # positive = lessons help

Budget Control: The Cost Meter

When working through the Budget Control The Cost stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

BUDGET_USD = 2.00
INPUT_PRICE_PER_TOKEN  = 3.00  / 1_000_000
OUTPUT_PRICE_PER_TOKEN = 15.00 / 1_000_000
class CostMeter:
    def record(self, input_tokens: int, output_tokens: int) -> None:
        if self.spent >= self._budget:              # check BEFORE adding
            raise BudgetExhaustedError(
                f"Budget ${self._budget:.2f} exhausted at ${self.spent:.4f}"
            )
        self.spent += input_tokens * INPUT_PRICE_PER_TOKEN \
                    + output_tokens * OUTPUT_PRICE_PER_TOKEN

Run State: Idempotent Initialization

When working through the Run State Idempotent Initialization stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

def _init_run_state() -> dict:
    path = _STATE_DIR / "run_state.json"
    state = {}
    if path.exists():
        try:
            content = path.read_text(encoding="utf-8-sig").strip()  # handles Windows BOM
            if content:
                state = json.loads(content)
        except (json.JSONDecodeError, UnicodeDecodeError):
            pass                                 # corrupt file → start fresh
    if not state:
        state = dict(_DEFAULT_RUN_STATE)
    state["judge_checksum"] = compute_judge_checksum()   # always refresh
    path.write_text(json.dumps(state, indent=2), encoding="utf-8")
    return state

The Three Run Modes

When working through the The Three Run Modes stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the The Three Run Modes stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

python -m src.main --mode single --app-id APP-001
python -m src.main --mode event-loop --once
python -m src.main --mode improve --rounds 5

Data Flow: End-to-End Trace

The Data Flow End-to-End Trace stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

pending_applications.json
          │
          │  [Loop 3 reads APP-009]
          ▼
    run_loop2("APP-009", prompt, meter)
          │
          │  [Loop 2 calls Loop 1]
          ▼
    run_loop1("APP-009", prompt, meter)
          │
          │  [LangGraph StateGraph]
          ▼
    agent_node → ChatAnthropic.invoke([SystemMsg, HumanMsg])
          │
          │  response: tool_call(lookup_application, {"app_id": "APP-009"})
          ▼
    tools_node → lookup_application.invoke({"app_id": "APP-009"})
          │
          │  returns: {id: APP-009, state: TX, coverage: 1200000, ...}
          ▼
    agent_node → ChatAnthropic.invoke([..., ToolMessage])
          │
          │  response: tool_call(retrieve_rules, {"risk_factor": "coverage"})
          ▼
    tools_node → retrieve_rules.invoke({"risk_factor": "coverage"})
          │
          │  returns: ["Coverage > $1M requires mandatory refer. Flag: high_coverage"]
          ▼
    agent_node → ChatAnthropic.invoke([..., ToolMessage])
          │
          │  response: tool_call(draft_decision, {decision: "refer", ...})
          ▼
    tools_node → draft_decision.invoke({...})
          │
          │  updates state: {decision: "refer", flags: ["high_coverage"], ...}
          ▼
    should_continue → END (no more tool calls)
          │
          ▼
    Loop 2: decision == "refer" → write human_queue.json → return (state, True)
          │
          ▼
    Loop 3: passed=True → update run_state.json → remove APP-009 from queue
          │
          │  [decisions_since_last_improvement becomes 9 >= 8]
          ▼
    Loop 4: improve(seed_prompt, lessons, meter, rounds=5)
          │
          │  reflect → evaluate → audit → gate → write best_prompt.txt
          ▼
    Loop 3: continue with APP-010 using new best_prompt

Key Engineering Lessons

The Key Engineering Lessons stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

1. Separate the oracle from the agent

The 1 Separate the oracle stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The 1 Separate the oracle stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

2. Determinism at the evaluation boundary

For the 2 Determinism at the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

3. The draft_decision tool as structured extraction

For the 3 The draftdecision tool stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

4. Anti-cheating is non-optional in self-improving systems

For the 4 Anti-cheating is non-optional stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the 4 Anti-cheating is non-optional stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

5. BOM and encoding are production bugs

When working through the 5 BOM and encoding stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

6. Budget before cost, not after

When working through the 6 Budget before cost stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

# WRONG: allows one overspend before raising
self.spent += cost
if self.spent > self._budget:
    raise BudgetExhaustedError(...)
# CORRECT: raises before the overspend registers
if self.spent >= self._budget:
    raise BudgetExhaustedError(...)
self.spent += cost

7. flush=True on progress lines in long-running loops

When working through the 7 flush True on stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

print(f"[loop3] {processed}/{total}  {app_id} ...", flush=True)

When working through the 7 flush True on stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

8. Hill-climbing budget must be separate from event loop budget

The 8 Hill-climbing budget must stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

try:
    system_prompt = run_hill_climbing(system_prompt, lessons, meter)
    run_state["decisions_since_last_improvement"] = 0
except BudgetExhaustedError:
    print(f"[loop3] hill-climbing skipped — budget exhausted at ${meter.spent:.4f}")
    run_state["decisions_since_last_improvement"] = 0

Extending This Pattern

The Extending This Pattern stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Running the System

The Running the System stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

# 1. Install
cd underwriting_loop
pip install -e ".[dev]"
export ANTHROPIC_API_KEY=sk-ant-...
# 2. Run a single application (debug mode - cheapest)
python -m src.main --mode single --app-id APP-001
# 3. Drain the pending queue with event loop
python -m src.main --mode event-loop --once
# 4. Run standalone hill-climbing (5 rounds)
python -m src.main --mode improve --rounds 5
# 5. Run the full offline test suite (no API key needed)
pytest tests/ -v

The Running the System stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

seed score: 44%  (9 failures)
round 1: 56% KEPT  (7 failures)
round 2: 56% REVERTED (score 56% <= best 56%)
round 3: 69% KEPT  (5 failures)
round 4: 75% KEPT  (4 failures)
round 5: 69% REVERTED (score 69% <= best 75%)
Final test evaluation...
spend:              $1.8342
test score:         75%
gain (vs baseline): +25%

Conclusion

For the Conclusion stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Pros and Cons of Loop Engineering

For the Pros and Cons of stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Pros

For the Pros stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Pros stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Cons

When working through the Cons stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

When to Use Loop Engineering — And When Not To

When working through the When to Use Loop stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

The Two-Question Test

When working through the The Two-Question Test stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the The Two-Question Test stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Use Loop Engineering When

The Use Loop Engineering When stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Do Not Use Loop Engineering When

The Do Not Use Loop stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

8 Things That Make a Loop Actually Work

The 8 Things That Make stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The 8 Things That Make stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

1. A Checkable Goal

For the 1 A Checkable Goal stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

2. A Hard Stop

For the 2 A Hard Stop stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

3. Good Tools

For the 3 Good Tools stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the 3 Good Tools stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

4. Memory

When working through the 4 Memory stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

5. A Separate Checker

When working through the 5 A Separate Checker stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

6. Plan Before Acting on Complex Tasks

When working through the 6 Plan Before Acting stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the 6 Plan Before Acting stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

7. Logging

The 7 Logging stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

8. Cost Sense

The 8 Cost Sense stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

# The ordering matters:
if self.spent >= self._budget:          # check BEFORE adding
    raise BudgetExhaustedError(...)
self.spent += cost                       # add AFTER the check clears

The Checklist

The The Checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The The Checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Referrences:

For the Referrences stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Operational checklist

When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.

Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.

Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for c0f4a1437d4f: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.