This article is published in English.
Practical notes: Many Hands, One Pen: Orchestrating Agents Without Losing the
Operable walkthrough of Practical notes: Many Hands, One Pen: Orchestrating Agents Without Losing the: contracts, checks, and drop-in code slots for teams shipping this pattern.
Use this as an operator-facing rebuild of the ideas in “Many Hands, One Pen: Orchestrating Agents Without Losing the Plot”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
071 · Default to a Workflow and Make the Agent Earn Its Autonomy
For the 071 Default to a stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
REFUND_LIMIT = 50.00
def llm_step(prompt: str) -> str:
"""Called only where the written rules run out."""
return model.complete(prompt).strip().lower()
def handle_refund(ticket):
if ticket.days_since_purchase > 30:
return deny(ticket, reason="outside window")
if ticket.amount <= REFUND_LIMIT:
return approve(ticket) # a rule, not a judgement call
intent = llm_step(
f"Classify as fraud, defect, or remorse:\n{ticket.body}"
)
if intent == "fraud":
return escalate(ticket, queue="risk")
if intent == "defect":
return approve(ticket)
return route_to_human(ticket)
072 · Skip the Orchestrator When You Already Know the Decomposition
For the 072 Skip the Orchestrator stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
# Before: up to 15 planning calls to rediscover a fixed list
def run_orchestrated(doc):
state = {"doc": doc}
for _ in range(15):
decision = orchestrator.plan(state) # 1 LLM call per pass
if decision.action == "finalize":
break
state = WORKERS[decision.worker](state)
return state
# After: 0 planning calls, same four workers, same result
PIPELINE = [fetch, extract, summarize, format_report]
def run_static(doc):
state = {"doc": doc}
for step in PIPELINE: # 0 LLM calls here
state = step(state)
return state
073 · Split Agents by Context Boundary, Not by Job Title
For the 073 Split Agents by stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the 073 Split Agents by stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
from dataclasses import dataclass
@dataclass
class Subtask:
name: str
needs: set[str] # facts this subtask must read
produces: set[str] # facts this subtask decides
def should_split(a: Subtask, b: Subtask) -> bool:
"""Split only when neither side needs what the other decides."""
shared = (a.needs & b.produces) | (b.needs & a.produces)
return not shared
implement = Subtask("implement", {"spec"}, {"api_shape", "error_semantics"})
test = Subtask("test", {"spec", "api_shape", "error_semantics"}, {"cases"})
assert not should_split(implement, test) # one agent writes both
074 · An Orchestrator Should Emit a Decision, Never Call a Tool
When working through the 074 An Orchestrator Should stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
from typing import Literal
from pydantic import BaseModel
class OrchestratorDecision(BaseModel):
next_action: Literal["delegate", "replan", "finalize"]
target_worker: str | None = None
task_description: str | None = None
reasoning: str
def orchestrator_step(state):
d = decide(state) # model has zero tools attached
if d.next_action == "finalize":
return finalize(state, d.reasoning)
if d.next_action == "replan":
return state.reset_plan(d.reasoning)
worker = WORKERS[d.target_worker] # workers own every tool
result = worker.run(d.task_description)
return state.record(d.target_worker, result)
075 · Give Every Subagent a Typed Output Contract, Not Free Text
When working through the 075 Give Every Subagent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
from pydantic import BaseModel, Field, ValidationError
class SubagentResult(BaseModel):
findings: list[str] = Field(max_length=5) # hard cap, one sentence each
sources: list[str]
open_questions: list[str] = []
completion_status: str # complete | partial | blocked
def parse_or_retry(worker, task, attempts=2):
for _ in range(attempts):
raw = worker.run(task, response_format=SubagentResult)
try:
return SubagentResult.model_validate_json(raw)
except ValidationError as err:
task = f"{task}\n\nRejected: {err}\nReturn only JSON in the schema."
raise RuntimeError(f"{worker.name} returned no valid result")
076 · Fan Out Reads, Funnel Writes Through One Agent
When working through the 076 Fan Out Reads stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the 076 Fan Out Reads stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
import asyncio
READ_TOOLS = ["search_repo", "read_file", "fetch_docs"]
async def gather_context(subtasks):
workers = [Agent(name=t.name, tools=READ_TOOLS) for t in subtasks]
return await asyncio.gather(
*(w.run(t.prompt) for w, t in zip(workers, subtasks))
)
async def build_feature(spec, subtasks):
findings = await gather_context(subtasks) # wide, parallel, read-only
writer = Agent(name="writer", tools=["write_file", "apply_patch"])
return await writer.run(spec, context=findings) # one writer, one pass
077 · Write Termination Conditions Down: Agents Genuinely Do Not Know When to Stop
The 077 Write Termination Conditions stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
CHECKLIST = [
"every requested section exists",
"each claim cites a retrieved source",
"open questions are listed, or explicitly none",
]
def verify_complete(objective, checklist, output) -> bool:
for item in checklist:
verdict = judge(f"Objective: {objective}\nCheck: {item}\n\n{output}")
if not verdict.passed:
log.info("termination blocked by: %s", item)
return False
return judge(f"Does this satisfy the objective?\n{objective}\n\n{output}").passed
def finish(state):
if verify_complete(state.objective, CHECKLIST, state.draft):
return state.done()
return state.keep_working(reason="checklist not satisfied")
Operational checklist
When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.
Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.
Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 64be605517f1: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.
For the hardening note 0 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 0/768: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 1 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 1/768: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 2 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 2/768: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 3 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 3/768: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 4 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 4/768: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 5 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 5/768: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 6 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 6/768: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 7 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 7/768: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.