Home / Articles / Practical notes: The Missing Layer in Production AI Agent Stacks

This article is published in English.

Practical notes: The Missing Layer in Production AI Agent Stacks

Operable walkthrough of Practical notes: The Missing Layer in Production AI Agent Stacks: contracts, checks, and drop-in code slots for teams shipping this pattern.

5562 words

Use this as an operator-facing rebuild of the ideas in “The Missing Layer in Production AI Agent Stacks”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

The discovery

For the The discovery stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

The three layers

For the The three layers stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

HARNESS — what the agent can touch

For the HARNESS what the agent stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the HARNESS what the agent stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

LOOP — when the agent stops

When working through the LOOP when the agent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

GRAPH — where the agent can go

When working through the GRAPH where the agent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

What happens when you confuse the layers

When working through the What happens when you stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the What happens when you stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Confusing the Harness with the Loop

The Confusing the Harness with stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

# CONFUSED LAYER — Harness owns Loop logic
async def query_suppliers(item: str, budget: float, quantity: int, attempts: int) -> dict:
    """Tool decides when the loop stops. Capability + control flow in one function."""
    results = query_mock_marketplace(item, budget, quantity)

    # The tool decides when to stop — couples execution logic with capability
    if len(results) > 5 or attempts >= 3:
        return {"status": "STOP", "suppliers": results}
    return {"status": "CONTINUE", "suppliers": results}


async def research_node(state: AgentState) -> dict:
    # The node just calls the tool and forwards whatever the tool decided.
    # It has no idea why the loop stopped — that logic is hidden inside the tool.
    result = await query_suppliers(
        state["brief"].item, state["brief"].budget, state["brief"].quantity,
        state.get("attempts", 0) + 1,
    )
    return {"suppliers": result["suppliers"], "status": result["status"].lower()}

# ✅ CLEAN SEPARATION — Harness is pure capability, Loop lives in the Node
MAX_RESEARCH_ATTEMPTS = 3

async def query_suppliers_impl(item: str, budget: float, quantity: int) -> dict:
    """Pure Harness capability: schema validation, circuit breaker, marketplace call.

    The tool has no idea a loop exists. It returns data. It does not decide
    when to stop. Circuit breaker + fault injection live here because they are
    capability concerns (is the downstream reachable?), not control-flow concerns.
    """
    breaker = get_circuit_breaker("query_suppliers")
    if not breaker.allow_call():
        raise CircuitBreakerOpenError("query_suppliers", breaker.state)

    maybe_inject("query_suppliers")  # fault injection for chaos tests
    try:
        results = query_mock_marketplace(item, budget, quantity)
        breaker.record_success()
        return {"suppliers": results}
    except Exception as e:
        breaker.record_failure()
        raise


async def research_node(state: AgentState) -> dict:
    """Graph Node / Loop: mechanical stop conditions and routing decisions.

    The node reads state, calls the tool through the registry, and decides
    whether to stop — based on counters in state, NOT on the LLM's judgment
    and NOT on the tool's opinion.
    """
    brief = state["brief"]
    attempts = state.get("attempts", 0) + 1

    # Harness call — schema validation, timeout, output cap all happen inside registry.call
    query_result = await registry.call("query_suppliers", {
        "item": brief.item,
        "budget": brief.budget,
        "quantity": brief.quantity,
    })
    suppliers = query_result.get("suppliers", [])

    # Mechanical stop condition — enforced OUTSIDE the tool, in code, on state.
    # Deterministic. Survives replay. Testable without the tool being live.
    if not suppliers and attempts >= MAX_RESEARCH_ATTEMPTS:
        logger.info(
            "STOP_CONDITION_FIRED path=no_matches item=%s attempts=%d max=%d",
            brief.item, attempts, MAX_RESEARCH_ATTEMPTS,
        )
        return {"suppliers": [], "attempts": attempts, "status": "no_matches"}

    if not suppliers:
        return {"suppliers": [], "attempts": attempts, "status": "running"}

    # Enrich, then mark completed
    enriched = []
    for s in suppliers:
        quote = await registry.call("get_price_quote", {
            "supplier_name": s["name"], "item": brief.item, "quantity": brief.quantity,
        })
        rating = await registry.call("check_seller_rating", {"supplier_name": s["name"]})
        enriched.append({**s, "quote": quote, "rating_info": rating})

    return {"suppliers": enriched, "attempts": attempts, "status": "completed"}


# The Graph owns WHERE the agent can go — conditional edges, not the node body.
def _should_retry_or_end(state: AgentState) -> str:
    """After Research: loop back if still running, else Score or END."""
    if state["status"] == "running":
        return "research"          # self-loop — the retry
    if state["status"] == "no_matches":
        return END                 # mechanical stop → graph terminates
    return "score"                 # success → next node


# Wiring (in build_sourcing_graph):
#   graph.add_conditional_edges("research", _should_retry_or_end,
#       {"research": "research", "score": "score", END: END})

Confusing the Loop with the Graph

The Confusing the Loop with stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

# ❌ CONFUSED LAYER — Implicit topology inside a monolithic ReAct loop
# The loop IS the graph. The LLM is the edge function. There is no "between."
while not agent.is_done():
    action = llm.decide_next_step(history)
    if action == "query":
        data = query_suppliers()
    elif action == "approve":
        # Retrofitting human approval into an implicit loop requires messy pauses.
        # You have to break the loop, externalize state, park the worker, and
        # resume on a webhook — the loop fights you because it was designed to
        # be self-driving.
        pause_execution_and_wait_for_webhook()
    elif action == "stop":
        break


# ✅ CLEAN SEPARATION — Declarative graph with explicit nodes and human interrupt
from langgraph.graph import StateGraph, START, END
from langgraph.types import interrupt


def approve_node(state: AgentState) -> dict:
    """Human checkpoint via interrupt() — a STRUCTURAL pause, not a prompt.

    The graph genuinely suspends here. interrupt() pauses execution and waits
    for a human to call Command(resume=approval_data). The worker is freed.
    If the worker crashes while waiting, Temporal replays the activity, the
    graph re-runs from the last checkpoint, and interrupt() fires again —
    the human re-approves. No data loss.
    """
    selected = state.get("selected_supplier", {})
    brief = state.get("brief", {})

    approval_request = {
        "item": brief.item,
        "quantity": brief.quantity,
        "supplier": selected.get("name", ""),
        "price": selected.get("price", 0),
        "total_cost": selected.get("price", 0) * brief.quantity,
    }

    # interrupt() pauses the graph HERE. The human must call /approve
    # with a resume value to continue.
    approval = interrupt(approval_request)

    approved = approval.get("approved", False)
    comment = approval.get("comment", "")

    return {
        "approval_status": "approved" if approved else "rejected",
        "approval_comment": comment,
        "rejection_count": state.get("rejection_count", 0) + (0 if approved else 1),
    }


def _after_approve(state: AgentState) -> str:
    """After Approve: go to Confirm if approved, retry Research if rejected."""
    if state.get("approval_status") == "approved":
        return "confirm"
    if state.get("rejection_count", 0) + 1 >= 3:
        return END
    return "research"


def build_sourcing_graph(checkpointer=None):
    """Topology is DECLARED DATA, not statement order in a loop body.

        START → Research → (retry) → Score → Decide → Approve → Confirm → END
                                              ↓ (rejected, < 3)
                                             Research
                                              ↓ (rejected, = 3)
                                             END
    """
    graph = StateGraph(AgentState)

    graph.add_node("research", research_node)
    graph.add_node("score", score_node)
    graph.add_node("decide", decide_node)
    graph.add_node("approve", approve_node)      # explicit checkpoint node
    graph.add_node("confirm", confirm_node)

    graph.add_edge(START, "research")
    graph.add_conditional_edges("research", _should_retry_or_end,
        {"research": "research", "score": "score", END: END})
    graph.add_edge("score", "decide")
    graph.add_edge("decide", "approve")
    graph.add_conditional_edges("approve", _after_approve,
        {"confirm": "confirm", "research": "research", END: END})
    graph.add_edge("confirm", END)

    return graph.compile(checkpointer=checkpointer)

Confusing the Graph with the Harness

The Confusing the Graph with stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve. The Confusing the Graph with stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

# ❌ CONFUSED LAYER — Graph node owns the tool call (Harness concern)
import httpx


def supplier_node(state: AgentState) -> dict:
    """The graph node reaches directly into HTTP, auth, retries, and parsing.

    Why that's wrong: the graph is topology. It should not know about HTTP,
    authentication, retries, circuit breakers, or schema validation. Those are
    harness concerns. When the graph owns the tool call:
      - You can't add a circuit breaker without modifying the graph.
      - You can't swap the tool implementation without modifying the graph.
      - You can't test the graph without the tool being live.
      - You can't enforce an egress allowlist — the URL is hardcoded here.
    """
    resp = httpx.get(
        f"https://supplier-api.example.com/suppliers",
        params={"item": state["brief"].item},
        headers={"Authorization": f"Bearer {SUPPLIER_API_KEY}"},
        timeout=10.0,
    )
    resp.raise_for_status()
    suppliers = resp.json()["suppliers"]  # no schema validation, no output cap
    return {"suppliers": suppliers}


# ✅ CLEAN SEPARATION — Graph node calls through the harness; harness owns the call
async def supplier_node(state: AgentState) -> dict:
    """The graph node says 'call this tool with these args' and gets back a
    validated, typed result. The graph doesn't know there's HTTP involved.
    The harness doesn't know there's a graph involved.
    """
    query_result = await registry.call("query_suppliers", {
        "item": state["brief"].item,
        "budget": state["brief"].budget,
        "quantity": state["brief"].quantity,
    })
    suppliers = query_result.get("suppliers", [])

    # Mechanical stop condition — the node's job, not the tool's
    if state.get("attempts", 0) + 1 >= MAX_RESEARCH_ATTEMPTS and not suppliers:
        return {"suppliers": [], "attempts": state.get("attempts", 0) + 1, "status": "no_matches"}

    return {"suppliers": suppliers, "attempts": state.get("attempts", 0) + 1, "status": "completed"}


# What the harness does — the graph never sees this:
#
#   registry.call("query_suppliers", {...})
#     │
#     ├─ 1. validate_input(QuerySuppliersInput, args)        # Pydantic v2
#     ├─ 2. EgressFilter(spec.allowed_egress)                # URL allowlist
#     ├─ 3. SandboxedHTTPClient(egress_filter)               # blocked domains = structural
#     ├─ 4. run_with_timeout(fn, kwargs, spec.timeout_seconds)
#     ├─ 5. enforce_output_cap(result, spec.max_output_bytes)  # 64KB cap
#     └─ 6. validate_output(QuerySuppliersOutput, result)    # garbage caught here
#
# The graph node gets back a validated dict. It never sees the HTTP call,
# the auth header, the retry, the circuit breaker, or the egress filter.
# Swap httpx for aiohttp, swap the supplier API for a new one, add a circuit
# breaker — the graph doesn't change. The harness doesn't know a graph exists.

Who owns which layer in the stack

For the Who owns which layer stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Temporal owns the LOOP’s crash-safety

For the Temporal owns the LOOP stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

LangGraph owns the GRAPH’s topology

For the LangGraph owns the GRAPH stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the LangGraph owns the GRAPH stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

The MCP tool harness is 100% mine

When working through the The MCP tool harness stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

The verification — the thing that separates “the agent said it worked” from “it worked”

When working through the The verification the thing stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

The separation principle

When working through the The separation principle stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the The separation principle stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

The mental model

The The mental model stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

What’s coming in the next episode

The What s coming in stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

References

The References stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The References stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Operational checklist

When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.

Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.

Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for e9b9da736831: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.

When working through the hardening note 0 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Hardening detail 0/826: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 1 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Hardening detail 1/826: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 2 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Hardening detail 2/826: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

When working through the hardening note 3 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Hardening detail 3/826: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 4 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Hardening detail 4/826: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 5 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Hardening detail 5/826: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

When working through the hardening note 6 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Hardening detail 6/826: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 7 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Hardening detail 7/826: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 8 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Hardening detail 8/826: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

When working through the hardening note 9 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Hardening detail 9/826: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 10 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Hardening detail 10/826: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 11 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Hardening detail 11/826: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

When working through the hardening note 12 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Hardening detail 12/826: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 13 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Hardening detail 13/826: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 14 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Hardening detail 14/826: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

When working through the hardening note 15 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Hardening detail 15/826: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 16 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Hardening detail 16/826: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 17 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Hardening detail 17/826: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

When working through the hardening note 18 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Hardening detail 18/826: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 19 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Hardening detail 19/826: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 20 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Hardening detail 20/826: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 0 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Hardening detail 0/845: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 1 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Hardening detail 1/845: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.