This article is published in English.
Practical notes: Loop Engineering: Building Self-Improving AI Agents with Four
Operable walkthrough of Practical notes: Loop Engineering: Building Self-Improving AI Agents with Four: contracts, checks, and drop-in code slots for teams shipping this pattern.
Use this as an operator-facing rebuild of the ideas in “Loop Engineering: Building Self-Improving AI Agents with Four Nested Loops”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
The Big Picture Breakdown: Engineering Beyond the Prompt
For the The Big Picture Breakdown stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
Why This Matters
For the Why This Matters stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
The Core Idea: Nested Control Loops
For the The Core Idea Nested stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the The Core Idea Nested stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
┌──────────────────────────────────────────────────────┐
│ Loop 4 — Hill Climbing (self-improvement over time) │
│ ┌────────────────────────────────────────────────┐ │
│ │ Loop 3 — Event Queue Poller (orchestration) │ │
│ │ ┌──────────────────────────────────────────┐ │ │
│ │ │ Loop 2 — Verification & Retry │ │ │
│ │ │ ┌────────────────────────────────────┐ │ │ │
│ │ │ │ Loop 1 — ReAct Agent │ │ │ │
│ │ │ │ (think → act → observe → think) │ │ │ │
│ │ │ └────────────────────────────────────┘ │ │ │
│ │ └──────────────────────────────────────────┘ │ │
│ └────────────────────────────────────────────────┘ │
└──────────────────────────────────────────────────────┘
The Domain: Insurance Underwriting
When working through the The Domain Insurance Underwriting stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
{
"id": "APP-001",
"applicant_age": 34,
"state": "FL",
"property_type": "residential",
"coverage_amount": 450000,
"claim_history_5yr": 1,
"credit_score_tier": "good",
"flood_zone": true,
"has_flood_rider": false,
"business_use": false
}
"APP-001": {
"split": "train",
"expected_decision": "modify",
"expected_flags": ["flood_rider_required", "windstorm_exclusion"]
}
Loop 1: The ReAct Agent
When working through the Loop 1 The ReAct stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
The Pattern
When working through the The Pattern stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the The Pattern stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
System Prompt
│
▼
[Think] → What information do I need?
│
▼
[Act] → Call tool (lookup_application, retrieve_rules, ...)
│
▼
[Observe] → Tool returns result
│
▼
[Think] → What does this mean? What next?
│
... (repeat until decision is made)
│
▼
[Act] → draft_decision (terminal tool call)
The LangGraph Implementation
The The LangGraph Implementation stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
class UnderwritingState(TypedDict):
application_id: str
messages: Annotated[list[AnyMessage], add_messages] # reducer accumulates
decision: str
rationale: str
flags: list[str]
recommended_premium_adj: float
lessons: list[str]
spend: float
_steps: int
def build_loop1(system_prompt: str, meter: CostMeter):
model = ChatAnthropic(model="claude-sonnet-4-6").bind_tools(TOOLS)
def agent_node(state: UnderwritingState) -> dict:
prompt = system_prompt
if state.get("lessons"):
prompt += "\n\nUnderwriting lessons:\n" + "\n".join(
f"- {l}" for l in state["lessons"]
)
msgs = [SystemMessage(content=prompt)] + list(state["messages"])
response = model.invoke(msgs)
meter.record_from_message(response) # track spend per call
return {
"messages": [response],
"_steps": state.get("_steps", 0) + 1,
"spend": meter.spent,
} def tools_node(state: UnderwritingState) -> dict:
last = state["messages"][-1]
tool_results, updates = _dispatch_tools(last.tool_calls)
return {"messages": tool_results, **updates} # updates captures draft_decision output def should_continue(state) -> Literal["tools", "__end__"]:
if state.get("_steps", 0) >= MAX_STEPS: # hard cap at 8 steps
return END
if getattr(state["messages"][-1], "tool_calls", None):
return "tools"
return END g = StateGraph(UnderwritingState)
g.add_node("agent", agent_node)
g.add_node("tools", tools_node)
g.set_entry_point("agent")
g.add_conditional_edges("agent", should_continue, {"tools": "tools", END: END})
g.add_edge("tools", "agent")
return g.compile()
The Four Tools
The The Four Tools stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
@tool
def lookup_application(app_id: str) -> dict:
"""Return the full application record for the given application ID."""
# reads from applications.json — deterministic, no LLM call
@tool
def retrieve_rules(risk_factor: str) -> list[str]:
"""Return underwriting rules matching the given risk factor keyword.
Use terms like: flood, claims, credit, coverage, state, business."""
# keyword match against rules_kb.json - forces explicit rule lookup
@tool
def fetch_risk_signal(signal_type: str) -> dict:
# simulated external data pull (credit bureaus, flood maps, etc.)
@tool
def draft_decision(
decision: Literal["accept", "modify", "decline", "refer"],
rationale: str,
flags: list[str],
recommended_premium_adj: float,
) -> dict:
"""Finalize the underwriting decision with structured output."""
# terminal action - structured schema forces the agent to commit explicitly
Why the Step Cap Matters
The Why the Step Cap stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Why the Step Cap stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Loop 2: Verification and Grader-Based Retry
For the Loop 2 Verification and stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
def run_loop2(
app_id: str,
system_prompt: str,
meter: CostMeter,
max_retries: int = 3,
) -> tuple[dict, bool]:
state = run_loop1(app_id, system_prompt, meter)
for attempt in range(max_retries):
decision = state.get("decision", "ran_out_of_steps")
if decision == "refer":
_write_human_queue(app_id, state) # human-in-the-loop path
return state, True
if is_correct(decision, app_id): # deterministic check
return state, True
if attempt >= max_retries - 1:
return state, False # exhausted retries
# construct targeted feedback for next attempt
expected = load_historical_decisions()[app_id]["expected_decision"]
feedback = HumanMessage(content=(
f"Your decision was '{decision}' but this application requires '{expected}'. "
f"Review the risk factors carefully and call draft_decision again."
))
extra_messages = list(state.get("messages", [])) + [feedback]
state = run_loop1(app_id, system_prompt, meter, extra_messages=extra_messages)
return state, False
Design Decisions Worth Noting
For the Design Decisions Worth Noting stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Loop 3: The Event Queue Poller
For the Loop 3 The Event stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Loop 3 The Event stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
def run_event_loop(
system_prompt: str,
meter: CostMeter,
pending_path: Path = DEFAULT_PENDING,
state_path: Path = DEFAULT_STATE,
once: bool = False,
) -> None:
run_state = load_run_state(state_path)
verify_judge_checksum(run_state.get("judge_checksum", "")) # tamper check
total = len(json.loads(pending_path.read_text(encoding="utf-8-sig"))
if pending_path.exists() else [])
processed = 0
while True:
pending = (json.loads(pending_path.read_text(encoding="utf-8-sig"))
if pending_path.exists() else [])
if not pending:
if once:
print("[loop3] queue empty - done")
break
print("[loop3] queue empty - sleeping...")
time.sleep(POLL_INTERVAL)
continue
app_id = pending[0]
processed += 1
print(f"[loop3] {processed}/{total} {app_id} ...", flush=True)
try:
result_state, passed = run_loop2(app_id, prompt_with_lessons, meter)
except BudgetExhaustedError:
raise # budget exhaustion is fatal
except Exception as exc:
print(f"[loop3] {app_id} ERROR: {exc!r} - skipping")
# remove from queue and continue - one bad app cannot block the rest
remaining = [x for x in pending if x != app_id]
pending_path.write_text(json.dumps(remaining), encoding="utf-8")
continue
# ... update state, distill lesson, trigger hill-climbing
Why File-Based Queue?
When working through the Why File-Based Queue stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Tamper Detection via Checksum
When working through the Tamper Detection via Checksum stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
verify_judge_checksum(run_state.get("judge_checksum", ""))
def compute_judge_checksum() -> str:
content = _HIST.read_bytes() # historical_decisions.json
return hashlib.sha256(content).hexdigest()
def verify_judge_checksum(expected: str) -> None:
actual = compute_judge_checksum()
if actual != expected:
raise RuntimeError(
f"Judge checksum mismatch - historical_decisions.json was tampered with."
)
Skipping vs Crashing
When working through the Skipping vs Crashing stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Skipping vs Crashing stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Loop 4: Hill Climbing (Self-Improvement)
The Loop 4 Hill Climbing stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
seed prompt
│
▼
evaluate on training set → score, failures
│
▼
┌────────────────────────────────┐
│ for each round: │
│ │
│ reflect(current, failures) │ ← LLM analyzes failure patterns
│ │ │
│ ▼ │
│ candidate_prompt │
│ │ │
│ ▼ │
│ evaluate(candidate, train) │ ← full eval run
│ │ │
│ ▼ │
│ audit_prompt(candidate) │ ← anti-cheating check
│ │ │
│ ▼ │
│ if score > best AND clean: │
│ best = candidate │ ← keep
│ else: │
│ discard │ ← revert
└────────────────────────────────┘
│
▼
write best_prompt.txt
The Reflect Step
The The Reflect Step stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
REFLECT_TEMPLATE = """You are improving the system prompt of an insurance underwriting agent.CURRENT PROMPT:
{prompt}
The agent got these decisions WRONG on training applications:
{failures}
Look for PATTERNS in the failures. Infer GENERAL underwriting rules that would fix them.
Do NOT memorize specific application IDs or applicant details.
Write an improved system prompt. Reply with ONLY the new prompt text."""
def reflect(current_prompt: str, failures: list[dict]) -> str:
client = anthropic.Anthropic()
failure_text = "\n".join(
f"- APP {f['app_id']}: agent said '{f['got']}', expected '{f['expected']}'. "
f"Application data: {f['application']}"
for f in failures
)
response = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=1500,
messages=[{"role": "user", "content": REFLECT_TEMPLATE.format(
prompt=current_prompt,
failures=failure_text,
)}],
)
return response.content[0].text.strip()
The Evaluate Step
The The Evaluate Step stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
def evaluate_prompt(system_prompt: str, split: str = "train", meter=None):
tasks = load_tasks(split)
failures = []
for task in tasks:
state = run_loop1(task["id"], system_prompt, meter)
decision = state.get("decision", "ran_out_of_steps")
if not is_correct(decision, task["id"]):
failures.append({
"app_id": task["id"],
"got": decision,
"expected": task["expected_decision"],
"application": task["application"],
})
score = (len(tasks) - len(failures)) / len(tasks)
return score, failures
The The Evaluate Step stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
The Gate: Anti-Overfitting and Anti-Cheating
For the The Gate Anti-Overfitting and stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
def improve(seed_prompt: str, lessons: list[str], meter: CostMeter, rounds: int = 5) -> str:
best = with_lessons(seed_prompt, lessons)
best_score, failures = evaluate_prompt(best, split="train", meter=meter)
for r in range(1, rounds + 1):
candidate = reflect(best, failures)
score, cand_failures = evaluate_prompt(candidate, split="train", meter=meter)
verdict = audit_prompt(candidate) # anti-cheat check
if score > best_score and verdict == "VERDICT: CLEAN":
best, best_score, failures = candidate, score, cand_failures
BEST_PROMPT_PATH.write_text(best, encoding="utf-8")
print(f"round {r}: {score:.0%} KEPT ({len(cand_failures)} failures)")
else:
reason = "cheating" if verdict != "VERDICT: CLEAN" else f"{score:.0%} <= {best_score:.0%}"
print(f"round {r}: {score:.0%} REVERTED ({reason})")
return best
The Prompt Auditor
For the The Prompt Auditor stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
AUDIT_SYSTEM = (
"You are a security auditor for AI systems. "
"Check whether the prompt contains any hardcoded application IDs (e.g. APP-001) "
"or hardcoded decisions for specific applications. "
"These would constitute cheating — memorizing answers instead of learning rules. "
"Reply with exactly one of:\n"
"VERDICT: CLEAN\n"
"VERDICT: CHEATING\n"
"Followed by a one-sentence explanation."
)
def audit_prompt(evolved_prompt: str) -> str:
response = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=200,
system=AUDIT_SYSTEM,
messages=[{"role": "user", "content": f"PROMPT TO AUDIT:\n{evolved_prompt}"}],
)
return response.content[0].text.strip().split("\n")[0] # first line only
Lesson Memory: Cross-Session Knowledge Distillation
For the Lesson Memory Cross-Session Knowledge stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Lesson Memory Cross-Session Knowledge stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
DISTILL_PROMPT = """An insurance underwriting agent made a wrong decision.Application details: {application}
Agent decided: {got}
Correct decision: {expected}
Write ONE short, general underwriting rule that would prevent this mistake.
Do NOT reference this specific application ID or any applicant names/details.
Reply with only the rule, as a single sentence."""
def distill_lesson(application: dict, got: str, expected: str) -> str:
response = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=150,
messages=[{"role": "user", "content": DISTILL_PROMPT.format(...)}],
)
return response.content[0].text.strip()
def with_lessons(base_prompt: str, lessons: list[str]) -> str:
if not lessons:
return base_prompt
lessons_text = "\n".join(f"- {l}" for l in lessons)
return base_prompt + f"\n\nUnderwriting lessons learned:\n{lessons_text}"
def measure_gain(base_prompt: str, lessons: list[str], split: str = "test", meter=None) -> float:
stateless_score, _ = evaluate_prompt(base_prompt, split=split, meter=meter)
augmented = with_lessons(base_prompt, lessons)
stateful_score, _ = evaluate_prompt(augmented, split=split, meter=meter)
return stateful_score - stateless_score # positive = lessons help
Budget Control: The Cost Meter
When working through the Budget Control The Cost stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
BUDGET_USD = 2.00
INPUT_PRICE_PER_TOKEN = 3.00 / 1_000_000
OUTPUT_PRICE_PER_TOKEN = 15.00 / 1_000_000
class CostMeter:
def record(self, input_tokens: int, output_tokens: int) -> None:
if self.spent >= self._budget: # check BEFORE adding
raise BudgetExhaustedError(
f"Budget ${self._budget:.2f} exhausted at ${self.spent:.4f}"
)
self.spent += input_tokens * INPUT_PRICE_PER_TOKEN \
+ output_tokens * OUTPUT_PRICE_PER_TOKEN
Run State: Idempotent Initialization
When working through the Run State Idempotent Initialization stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
def _init_run_state() -> dict:
path = _STATE_DIR / "run_state.json"
state = {}
if path.exists():
try:
content = path.read_text(encoding="utf-8-sig").strip() # handles Windows BOM
if content:
state = json.loads(content)
except (json.JSONDecodeError, UnicodeDecodeError):
pass # corrupt file → start fresh
if not state:
state = dict(_DEFAULT_RUN_STATE)
state["judge_checksum"] = compute_judge_checksum() # always refresh
path.write_text(json.dumps(state, indent=2), encoding="utf-8")
return state
The Three Run Modes
When working through the The Three Run Modes stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the The Three Run Modes stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
python -m src.main --mode single --app-id APP-001
python -m src.main --mode event-loop --once
python -m src.main --mode improve --rounds 5
Data Flow: End-to-End Trace
The Data Flow End-to-End Trace stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
pending_applications.json
│
│ [Loop 3 reads APP-009]
▼
run_loop2("APP-009", prompt, meter)
│
│ [Loop 2 calls Loop 1]
▼
run_loop1("APP-009", prompt, meter)
│
│ [LangGraph StateGraph]
▼
agent_node → ChatAnthropic.invoke([SystemMsg, HumanMsg])
│
│ response: tool_call(lookup_application, {"app_id": "APP-009"})
▼
tools_node → lookup_application.invoke({"app_id": "APP-009"})
│
│ returns: {id: APP-009, state: TX, coverage: 1200000, ...}
▼
agent_node → ChatAnthropic.invoke([..., ToolMessage])
│
│ response: tool_call(retrieve_rules, {"risk_factor": "coverage"})
▼
tools_node → retrieve_rules.invoke({"risk_factor": "coverage"})
│
│ returns: ["Coverage > $1M requires mandatory refer. Flag: high_coverage"]
▼
agent_node → ChatAnthropic.invoke([..., ToolMessage])
│
│ response: tool_call(draft_decision, {decision: "refer", ...})
▼
tools_node → draft_decision.invoke({...})
│
│ updates state: {decision: "refer", flags: ["high_coverage"], ...}
▼
should_continue → END (no more tool calls)
│
▼
Loop 2: decision == "refer" → write human_queue.json → return (state, True)
│
▼
Loop 3: passed=True → update run_state.json → remove APP-009 from queue
│
│ [decisions_since_last_improvement becomes 9 >= 8]
▼
Loop 4: improve(seed_prompt, lessons, meter, rounds=5)
│
│ reflect → evaluate → audit → gate → write best_prompt.txt
▼
Loop 3: continue with APP-010 using new best_prompt
Key Engineering Lessons
The Key Engineering Lessons stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
1. Separate the oracle from the agent
The 1 Separate the oracle stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The 1 Separate the oracle stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
2. Determinism at the evaluation boundary
For the 2 Determinism at the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
3. The draft_decision tool as structured extraction
For the 3 The draftdecision tool stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.
4. Anti-cheating is non-optional in self-improving systems
For the 4 Anti-cheating is non-optional stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the 4 Anti-cheating is non-optional stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
5. BOM and encoding are production bugs
When working through the 5 BOM and encoding stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
6. Budget before cost, not after
When working through the 6 Budget before cost stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
# WRONG: allows one overspend before raising
self.spent += cost
if self.spent > self._budget:
raise BudgetExhaustedError(...)
# CORRECT: raises before the overspend registers
if self.spent >= self._budget:
raise BudgetExhaustedError(...)
self.spent += cost
7. flush=True on progress lines in long-running loops
When working through the 7 flush True on stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
print(f"[loop3] {processed}/{total} {app_id} ...", flush=True)
When working through the 7 flush True on stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
8. Hill-climbing budget must be separate from event loop budget
The 8 Hill-climbing budget must stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
try:
system_prompt = run_hill_climbing(system_prompt, lessons, meter)
run_state["decisions_since_last_improvement"] = 0
except BudgetExhaustedError:
print(f"[loop3] hill-climbing skipped — budget exhausted at ${meter.spent:.4f}")
run_state["decisions_since_last_improvement"] = 0
Extending This Pattern
The Extending This Pattern stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Running the System
The Running the System stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
# 1. Install
cd underwriting_loop
pip install -e ".[dev]"
export ANTHROPIC_API_KEY=sk-ant-...
# 2. Run a single application (debug mode - cheapest)
python -m src.main --mode single --app-id APP-001
# 3. Drain the pending queue with event loop
python -m src.main --mode event-loop --once
# 4. Run standalone hill-climbing (5 rounds)
python -m src.main --mode improve --rounds 5
# 5. Run the full offline test suite (no API key needed)
pytest tests/ -v
The Running the System stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
seed score: 44% (9 failures)
round 1: 56% KEPT (7 failures)
round 2: 56% REVERTED (score 56% <= best 56%)
round 3: 69% KEPT (5 failures)
round 4: 75% KEPT (4 failures)
round 5: 69% REVERTED (score 69% <= best 75%)
Final test evaluation...
spend: $1.8342
test score: 75%
gain (vs baseline): +25%
Conclusion
For the Conclusion stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Pros and Cons of Loop Engineering
For the Pros and Cons of stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Pros
For the Pros stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Pros stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Cons
When working through the Cons stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
When to Use Loop Engineering — And When Not To
When working through the When to Use Loop stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
The Two-Question Test
When working through the The Two-Question Test stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the The Two-Question Test stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Use Loop Engineering When
The Use Loop Engineering When stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Do Not Use Loop Engineering When
The Do Not Use Loop stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
8 Things That Make a Loop Actually Work
The 8 Things That Make stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The 8 Things That Make stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
1. A Checkable Goal
For the 1 A Checkable Goal stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
2. A Hard Stop
For the 2 A Hard Stop stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
3. Good Tools
For the 3 Good Tools stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the 3 Good Tools stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
4. Memory
When working through the 4 Memory stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
5. A Separate Checker
When working through the 5 A Separate Checker stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
6. Plan Before Acting on Complex Tasks
When working through the 6 Plan Before Acting stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the 6 Plan Before Acting stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
7. Logging
The 7 Logging stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
8. Cost Sense
The 8 Cost Sense stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
# The ordering matters:
if self.spent >= self._budget: # check BEFORE adding
raise BudgetExhaustedError(...)
self.spent += cost # add AFTER the check clears
The Checklist
The The Checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The The Checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Referrences:
For the Referrences stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Operational checklist
When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.
Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for c0f4a1437d4f: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.