This article is published in English.
Practical notes: The DeepAgents Architecture, Part 2: Building the Agent Harness
Operable walkthrough of Practical notes: The DeepAgents Architecture, Part 2: Building the Agent Harness: contracts, checks, and drop-in code slots for teams shipping this pattern.
Use this as an operator-facing rebuild of the ideas in “The DeepAgents Architecture, Part 2: Building the Agent Harness”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Getting Set Up
For the Getting Set Up stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.
git clone https://github.com/shubhodayahampiholi/system-design-planner.git
cd system-design-planner
uv sync
cp .env.example .env # fill in real API keys
How DeepAgents Handles Files and Access Control
For the How DeepAgents Handles Files stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.
Choosing Where an Agent’s Files Are Stored
For the Choosing Where an Agent stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the Choosing Where an Agent stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
# src/system_design_planner/agent.py
from deepagents import create_deep_agent
MODEL = "claude-sonnet-5"
def build_planner_agent(*, system_prompt=PLANNER_SYSTEM_PROMPT, **kwargs):
return create_deep_agent(model=MODEL, system_prompt=system_prompt, **kwargs)
Deciding What the Agent Is Allowed to Do
When working through the Deciding What the Agent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
# src/system_design_planner/permissions.py
from deepagents import FilesystemPermission
DESIGN_SESSION_PERMISSIONS = [
FilesystemPermission(operations=["read", "write"], paths=["/.env"], mode="deny"),
FilesystemPermission(operations=["write"], paths=["/knowledge_base/**"], mode="deny"),
FilesystemPermission(operations=["write"], paths=["/memory/AGENTS.md"], mode="interrupt"),
]
Confirming the Agent Actually Reads What It’s Given
When working through the Confirming the Agent Actually stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
backend = FilesystemBackend(root_dir="knowledge_base")
agent = build_planner_agent(backend=backend)
result = agent.invoke({
"messages": [{
"role": "user",
"content": "List every file in your working directory, then give a one-line summary of what each one covers.",
}]
})
Delegating Work to Subagents
When working through the Delegating Work to Subagents stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
What Delegation Actually Means
When working through the What Delegation Actually Means stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
Assigning Different Models to Different Subagents
When working through the Assigning Different Models to stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
from deepagents import SubAgent
REFERENCE_EXTRACTOR: SubAgent = {
"name": "reference-extractor",
"description": (
"Extracts structured, factual capabilities from the knowledge_base "
"reference files for a specific platform or topic. Delegate here "
"before proposing any subsystem design, so decisions are grounded "
"in verified platform facts rather than assumption. Do not use this "
"subagent for design reasoning or tradeoffs - extraction only."
),
"system_prompt": (
"You are a fact-extraction specialist. Given a topic, find the "
"relevant file(s) in your working directory, read them, and return "
"ONLY a bullet list of the concrete facts relevant to that topic - "
"no narrative, no design opinions, no recommendations. If a fact is "
"explicitly flagged as an open question or unverified in the source "
"file, preserve that flag in your output rather than smoothing it "
"over into a confident-sounding statement."
),
"model": "openai:gpt-5.6-luna",
}
Why a Subagent’s Memory Is Its Own, by Default
When working through the Why a Subagent s stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
def build_governance_specialist(backend) -> SubAgent:
from deepagents.middleware.memory import MemoryMiddleware
return {
"name": "governance-specialist",
"description": (
"Reasons about Unity Catalog governance, lineage, and the "
"Databricks-Microsoft Foundry governance boundary for this "
"architecture."
),
"system_prompt": (
"You are a governance specialist for an Azure + Databricks AI "
"architecture. Before answering, check /memory/AGENTS.md for any "
"recorded scoping decisions - they are binding constraints on "
"your recommendation, not suggestions."
),
"middleware": [MemoryMiddleware(backend=backend, sources=["/memory/AGENTS.md"])],
}
Human Approval and Persistent Memory
When working through the Human Approval and Persistent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
Why Some Actions Need a Person’s Sign-Off
When working through the Why Some Actions Need stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours. When working through the Why Some Actions Need stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
How Approval and Resumption Actually Work
The How Approval and Resumption stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
from langgraph.checkpoint.memory import MemorySaver
from langgraph.types import Command
checkpointer = MemorySaver()
agent = build_planner_agent(
backend=backend,
permissions=DESIGN_SESSION_PERMISSIONS,
checkpointer=checkpointer,
memory=["/memory/AGENTS.md"],
)
state = agent.get_state(config)
if state.interrupts:
action = state.interrupts[0].value["action_requests"][0]
print(f"tool: {action['name']}")
print(f"args: {action['args']}")
decision = {"type": "approve"}
# or:
decision = {"type": "reject", "message": "Not yet ready — needs client confirmation first."}
result = agent.invoke(Command(resume={"decisions": [decision]}), config=config)
What Happens When the Model Isn’t Told Why It Was Rejected
The What Happens When the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
Why Approving Several Decisions at Once Deserves Extra Care
The Why Approving Several Decisions stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve. The Why Approving Several Decisions stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Coordinating Multiple Specialists at Once
For the Coordinating Multiple Specialists at stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.
Letting the Agent Decide Who Should Handle a Question
For the Letting the Agent Decide stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.
Genuine Parallel Delegation
For the Genuine Parallel Delegation stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the Genuine Parallel Delegation stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
What Real Synthesis Looks Like
When working through the What Real Synthesis Looks stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
When Two Approved Decisions Contradict Each Other
When working through the When Two Approved Decisions stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
Giving the Agent New Capabilities
When working through the Giving the Agent New stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
Building a Custom Tool
When working through the Building a Custom Tool stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
from langchain_core.tools import tool
from langchain_tavily import TavilySearch
_tavily_instance = None
def _get_tavily() -> TavilySearch:
global _tavily_instance
if _tavily_instance is None:
_tavily_instance = TavilySearch(max_results=3, topic="general")
return _tavily_instance
@tool
def check_current_standards(query: str) -> str:
"""Search the live web to verify whether a platform capability, naming,
or integration detail is still current.
Use this specifically to check something already pulled from
knowledge_base against what's true right now - not for open-ended
research.
"""
result = _get_tavily().invoke({"query": query})
return str(result)
Giving the Agent Reusable Procedures with Skills
When working through the Giving the Agent Reusable stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
---
name: native-tool-scoping-check
description: Checks whether a request to use "native tooling" is genuinely unambiguous, given the deep current integration between Databricks-native and Azure-native (Microsoft Foundry) services.
license: MIT
---
# Native-Tool Scoping Check
## The procedure
1. Check /memory/AGENTS.md for an existing scoping decision covering this
question. If one exists, treat it as binding and stop here.
2. If no decision exists, do not assume either interpretation - surface
the ambiguity explicitly as a scoping question.
3. Once resolved, the decision should be persisted to memory, subject to
human approval.
Connecting to a Real External System with MCP
When working through the Connecting to a Real stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
import asyncio
from databricks.sdk import WorkspaceClient
from databricks_langchain import DatabricksMCPServer, DatabricksMultiServerMCPClient
async def _fetch_uc_function_tools():
workspace_client = WorkspaceClient()
host = workspace_client.config.host
mcp_client = DatabricksMultiServerMCPClient([
DatabricksMCPServer(
name="uc-functions",
url=f"{host}/api/2.0/mcp/functions/{CATALOG}/{SCHEMA}",
workspace_client=workspace_client,
),
])
return await mcp_client.get_tools()
def get_databricks_uc_function_tools():
return asyncio.run(_fetch_uc_function_tools())
Building an Interactive Interface
When working through the Building an Interactive Interface stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
A Different Kind of Application
When working through the A Different Kind of stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours. When working through the A Different Kind of stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Showing What the Agent Is Doing, As It Happens
The Showing What the Agent stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
def classify_tool_call(tool_call, mcp_tool_names):
name = tool_call["name"]
args = tool_call.get("args", {})
if name == "task":
return "subagent", args.get("subagent_type", "?")
if name in mcp_tool_names:
return "mcp", name
if name == "read_file":
path = str(args.get("file_path", ""))
if "/skills/project/" in path and path.endswith("SKILL.md"):
skill_name = path.split("/")[-2]
return "skill", skill_name
if name in ("read_file", "write_file", "edit_file", "ls", "glob", "grep", "delete"):
return "filesystem", name
return "tool", name
Approving or Rejecting Actions From the Interface
The Approving or Rejecting Actions stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
if st.session_state.pending_interrupt:
action = st.session_state.pending_interrupt
st.warning(f"**Approval needed**\n\n**Tool:** `{action['name']}`\n\n**Args:** `{action['args']}`")
col1, col2 = st.columns(2)
with col1:
if st.button("Approve"):
run_turn(Command(resume={"decisions": [{"type": "approve"}]}))
with col2:
reason = st.text_input("Reason for rejecting")
if st.button("Reject"):
decision = {"type": "reject", "message": reason or "Rejected by the human reviewer."}
run_turn(Command(resume={"decisions": [decision]}))
A Known Limit on What the Interface Can Show
The A Known Limit on stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve. The A Known Limit on stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
What This Build Actually Confirms
For the What This Build Actually stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.
The Scaffolding Held
For the The Scaffolding Held stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.
What Scaffolding Doesn’t Fix
For the What Scaffolding Doesn t stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the What Scaffolding Doesn t stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Why This Matters
When working through the Why This Matters stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
Operational checklist
When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.
Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 35d60e28a332: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.
For the hardening note 0 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 0/762: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 1 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 1/762: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 2 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 2/762: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 3 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 3/762: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 4 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 4/762: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 5 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 5/762: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 6 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 6/762: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 7 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 7/762: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 8 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 8/762: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 9 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 9/762: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 10 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 10/762: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.