Home / Articles / Practical notes: The DeepAgents Architecture, Part 2: Building the Agent Harness

This article is published in English.

Practical notes: The DeepAgents Architecture, Part 2: Building the Agent Harness

Operable walkthrough of Practical notes: The DeepAgents Architecture, Part 2: Building the Agent Harness: contracts, checks, and drop-in code slots for teams shipping this pattern.

5136 words

Use this as an operator-facing rebuild of the ideas in “The DeepAgents Architecture, Part 2: Building the Agent Harness”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Getting Set Up

For the Getting Set Up stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

git clone https://github.com/shubhodayahampiholi/system-design-planner.git
cd system-design-planner
uv sync
cp .env.example .env   # fill in real API keys

How DeepAgents Handles Files and Access Control

For the How DeepAgents Handles Files stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

Choosing Where an Agent’s Files Are Stored

For the Choosing Where an Agent stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the Choosing Where an Agent stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

# src/system_design_planner/agent.py
from deepagents import create_deep_agent

MODEL = "claude-sonnet-5"

def build_planner_agent(*, system_prompt=PLANNER_SYSTEM_PROMPT, **kwargs):
    return create_deep_agent(model=MODEL, system_prompt=system_prompt, **kwargs)

Deciding What the Agent Is Allowed to Do

When working through the Deciding What the Agent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

# src/system_design_planner/permissions.py
from deepagents import FilesystemPermission

DESIGN_SESSION_PERMISSIONS = [
    FilesystemPermission(operations=["read", "write"], paths=["/.env"], mode="deny"),
    FilesystemPermission(operations=["write"], paths=["/knowledge_base/**"], mode="deny"),
    FilesystemPermission(operations=["write"], paths=["/memory/AGENTS.md"], mode="interrupt"),
]

Confirming the Agent Actually Reads What It’s Given

When working through the Confirming the Agent Actually stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

backend = FilesystemBackend(root_dir="knowledge_base")
agent = build_planner_agent(backend=backend)

result = agent.invoke({
    "messages": [{
        "role": "user",
        "content": "List every file in your working directory, then give a one-line summary of what each one covers.",
    }]
})

Delegating Work to Subagents

When working through the Delegating Work to Subagents stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

What Delegation Actually Means

When working through the What Delegation Actually Means stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

Assigning Different Models to Different Subagents

When working through the Assigning Different Models to stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

from deepagents import SubAgent

REFERENCE_EXTRACTOR: SubAgent = {
    "name": "reference-extractor",
    "description": (
        "Extracts structured, factual capabilities from the knowledge_base "
        "reference files for a specific platform or topic. Delegate here "
        "before proposing any subsystem design, so decisions are grounded "
        "in verified platform facts rather than assumption. Do not use this "
        "subagent for design reasoning or tradeoffs - extraction only."
    ),
    "system_prompt": (
        "You are a fact-extraction specialist. Given a topic, find the "
        "relevant file(s) in your working directory, read them, and return "
        "ONLY a bullet list of the concrete facts relevant to that topic - "
        "no narrative, no design opinions, no recommendations. If a fact is "
        "explicitly flagged as an open question or unverified in the source "
        "file, preserve that flag in your output rather than smoothing it "
        "over into a confident-sounding statement."
    ),
    "model": "openai:gpt-5.6-luna",
}

Why a Subagent’s Memory Is Its Own, by Default

When working through the Why a Subagent s stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

def build_governance_specialist(backend) -> SubAgent:
    from deepagents.middleware.memory import MemoryMiddleware

    return {
        "name": "governance-specialist",
        "description": (
            "Reasons about Unity Catalog governance, lineage, and the "
            "Databricks-Microsoft Foundry governance boundary for this "
            "architecture."
        ),
        "system_prompt": (
            "You are a governance specialist for an Azure + Databricks AI "
            "architecture. Before answering, check /memory/AGENTS.md for any "
            "recorded scoping decisions - they are binding constraints on "
            "your recommendation, not suggestions."
        ),
        "middleware": [MemoryMiddleware(backend=backend, sources=["/memory/AGENTS.md"])],
    }

Human Approval and Persistent Memory

When working through the Human Approval and Persistent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

Why Some Actions Need a Person’s Sign-Off

When working through the Why Some Actions Need stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours. When working through the Why Some Actions Need stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

How Approval and Resumption Actually Work

The How Approval and Resumption stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

from langgraph.checkpoint.memory import MemorySaver
from langgraph.types import Command

checkpointer = MemorySaver()

agent = build_planner_agent(
    backend=backend,
    permissions=DESIGN_SESSION_PERMISSIONS,
    checkpointer=checkpointer,
    memory=["/memory/AGENTS.md"],
)
state = agent.get_state(config)

if state.interrupts:
    action = state.interrupts[0].value["action_requests"][0]
    print(f"tool: {action['name']}")
    print(f"args: {action['args']}")
decision = {"type": "approve"}
# or:
decision = {"type": "reject", "message": "Not yet ready — needs client confirmation first."}

result = agent.invoke(Command(resume={"decisions": [decision]}), config=config)

What Happens When the Model Isn’t Told Why It Was Rejected

The What Happens When the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

Why Approving Several Decisions at Once Deserves Extra Care

The Why Approving Several Decisions stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve. The Why Approving Several Decisions stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Coordinating Multiple Specialists at Once

For the Coordinating Multiple Specialists at stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

Letting the Agent Decide Who Should Handle a Question

For the Letting the Agent Decide stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

Genuine Parallel Delegation

For the Genuine Parallel Delegation stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the Genuine Parallel Delegation stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

What Real Synthesis Looks Like

When working through the What Real Synthesis Looks stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

When Two Approved Decisions Contradict Each Other

When working through the When Two Approved Decisions stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

Giving the Agent New Capabilities

When working through the Giving the Agent New stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

Building a Custom Tool

When working through the Building a Custom Tool stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

from langchain_core.tools import tool
from langchain_tavily import TavilySearch

_tavily_instance = None

def _get_tavily() -> TavilySearch:
    global _tavily_instance
    if _tavily_instance is None:
        _tavily_instance = TavilySearch(max_results=3, topic="general")
    return _tavily_instance

@tool
def check_current_standards(query: str) -> str:
    """Search the live web to verify whether a platform capability, naming,
    or integration detail is still current.

    Use this specifically to check something already pulled from
    knowledge_base against what's true right now - not for open-ended
    research.
    """
    result = _get_tavily().invoke({"query": query})
    return str(result)

Giving the Agent Reusable Procedures with Skills

When working through the Giving the Agent Reusable stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

---
name: native-tool-scoping-check
description: Checks whether a request to use "native tooling" is genuinely unambiguous, given the deep current integration between Databricks-native and Azure-native (Microsoft Foundry) services.
license: MIT
---

# Native-Tool Scoping Check

## The procedure
1. Check /memory/AGENTS.md for an existing scoping decision covering this
   question. If one exists, treat it as binding and stop here.
2. If no decision exists, do not assume either interpretation - surface
   the ambiguity explicitly as a scoping question.
3. Once resolved, the decision should be persisted to memory, subject to
   human approval.

Connecting to a Real External System with MCP

When working through the Connecting to a Real stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

import asyncio
from databricks.sdk import WorkspaceClient
from databricks_langchain import DatabricksMCPServer, DatabricksMultiServerMCPClient

async def _fetch_uc_function_tools():
    workspace_client = WorkspaceClient()
    host = workspace_client.config.host
    mcp_client = DatabricksMultiServerMCPClient([
        DatabricksMCPServer(
            name="uc-functions",
            url=f"{host}/api/2.0/mcp/functions/{CATALOG}/{SCHEMA}",
            workspace_client=workspace_client,
        ),
    ])
    return await mcp_client.get_tools()

def get_databricks_uc_function_tools():
    return asyncio.run(_fetch_uc_function_tools())

Building an Interactive Interface

When working through the Building an Interactive Interface stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

A Different Kind of Application

When working through the A Different Kind of stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours. When working through the A Different Kind of stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Showing What the Agent Is Doing, As It Happens

The Showing What the Agent stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

def classify_tool_call(tool_call, mcp_tool_names):
    name = tool_call["name"]
    args = tool_call.get("args", {})

    if name == "task":
        return "subagent", args.get("subagent_type", "?")
    if name in mcp_tool_names:
        return "mcp", name
    if name == "read_file":
        path = str(args.get("file_path", ""))
        if "/skills/project/" in path and path.endswith("SKILL.md"):
            skill_name = path.split("/")[-2]
            return "skill", skill_name
    if name in ("read_file", "write_file", "edit_file", "ls", "glob", "grep", "delete"):
        return "filesystem", name
    return "tool", name

Approving or Rejecting Actions From the Interface

The Approving or Rejecting Actions stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

if st.session_state.pending_interrupt:
    action = st.session_state.pending_interrupt
    st.warning(f"**Approval needed**\n\n**Tool:** `{action['name']}`\n\n**Args:** `{action['args']}`")
    col1, col2 = st.columns(2)
    with col1:
        if st.button("Approve"):
            run_turn(Command(resume={"decisions": [{"type": "approve"}]}))
    with col2:
        reason = st.text_input("Reason for rejecting")
        if st.button("Reject"):
            decision = {"type": "reject", "message": reason or "Rejected by the human reviewer."}
            run_turn(Command(resume={"decisions": [decision]}))

A Known Limit on What the Interface Can Show

The A Known Limit on stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve. The A Known Limit on stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

What This Build Actually Confirms

For the What This Build Actually stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

The Scaffolding Held

For the The Scaffolding Held stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

What Scaffolding Doesn’t Fix

For the What Scaffolding Doesn t stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the What Scaffolding Doesn t stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Why This Matters

When working through the Why This Matters stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

Operational checklist

When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.

Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.

Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for 35d60e28a332: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.

For the hardening note 0 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Hardening detail 0/762: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

When working through the hardening note 1 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Hardening detail 1/762: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 2 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Hardening detail 2/762: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 3 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Hardening detail 3/762: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

When working through the hardening note 4 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Hardening detail 4/762: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 5 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Hardening detail 5/762: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 6 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Hardening detail 6/762: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

When working through the hardening note 7 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Hardening detail 7/762: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 8 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Hardening detail 8/762: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 9 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Hardening detail 9/762: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

When working through the hardening note 10 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Hardening detail 10/762: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.