This article is published in English.
Practical notes: All You Want For Agentic Memory Design
Operable walkthrough of Practical notes: All You Want For Agentic Memory Design: contracts, checks, and drop-in code slots for teams shipping this pattern.
The following notes reconstruct a practical path around “All You Want For Agentic Memory Design”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Table of Contents
The Table of Contents stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
The Agent kept forgetting
The The Agent kept forgetting stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Why Context windows fail (and the study that proved it)
The Why Context windows fail stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Why Context windows fail stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Short-term vs Long-term Memory: The Foundational split
For the Short-term vs Long-term Memory stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
How the industry converged on four memory types
For the How the industry converged stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Production Database decisions: SQL vs Vector vs Graph
For the Production Database decisions SQL stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Production Database decisions SQL stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Layer 1: Buffer memory with Summarization
When working through the Layer 1 Buffer memory stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
from collections import deque
from typing import List, Dict
class BufferMemory:
def __init__(self, max_turns: int = 10):
# Each turn = one user message + one assistant message
self.history: deque = deque(maxlen=max_turns * 2)
def add_message(self, role: str, content: str):
self.history.append({"role": role, "content": content})
def get_context(self) -> List[Dict]:
return list(self.history)
def token_estimate(self) -> int:
total_chars = sum(len(m["content"]) for m in self.history)
return total_chars // 4 # rough approximation
def clear(self):
self.history.clear()
import os
from collections import deque
from typing import List, Dict
from openai import OpenAI
class SummarizingBufferMemory:
def __init__(self, max_turns: int = 8, summary_batch: int = 4):
self.client = OpenAI(
base_url="https://openrouter.ai/api/v1",
api_key=os.environ["OPENROUTER_API_KEY"],
)
self.model = "deepseek/deepseek-v4-flash-0731"
self.recent: deque = deque(maxlen=max_turns * 2)
self.rolling_summary: str = ""
self.summary_batch = summary_batch
def add_message(self, role: str, content: str):
if len(self.recent) == self.recent.maxlen:
self._compress_oldest()
self.recent.append({"role": role, "content": content})
def _compress_oldest(self):
batch = [self.recent.popleft() for _ in range(min(self.summary_batch * 2, len(self.recent)))]
text = "\n".join(f"{m['role']}: {m['content']}" for m in batch)
response = self.client.chat.completions.create(
model=self.model,
max_tokens=300,
messages=[
{
"role": "system",
"content": "You compress conversation history. Reply with the summary only.",
},
{
"role": "user",
"content": (
"Summarize this conversation segment in 2-4 sentences. "
"Preserve any decisions made, constraints stated, and conclusions reached.\n\n"
f"{text}"
),
},
],
)
new_summary = response.choices[0].message.content.strip()
if self.rolling_summary:
self.rolling_summary = f"{self.rolling_summary} | {new_summary}"
else:
self.rolling_summary = new_summary
def get_context(self) -> List[Dict]:
context = []
if self.rolling_summary:
context.append({
"role": "system",
"content": f"[Prior conversation summary: {self.rolling_summary}]",
})
context.extend(list(self.recent))
return context
Layer 2: Episodic memory with Citation validation
When working through the Layer 2 Episodic memory stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
import json
import sqlite3
from datetime import datetime, timedelta
from dataclasses import dataclass, field
@dataclass
class Episode:
task_type: str
user_query: str
outcome: str
tools_used: list
citations: list # source references for validation
duration_seconds: float
success: bool
session_id: str = ""
created_at: str = field(default_factory=lambda: datetime.utcnow().isoformat())
class EpisodicMemory:
EXPIRY_DAYS = 28 # GitHub Copilot's production default
def __init__(self, db_path: str = "episodes.db"):
self.conn = sqlite3.connect(db_path, check_same_thread=False)
self._init_schema()
def _init_schema(self):
self.conn.executescript("""
CREATE TABLE IF NOT EXISTS episodes (
id INTEGER PRIMARY KEY AUTOINCREMENT,
task_type TEXT NOT NULL,
user_query TEXT,
outcome TEXT,
tools_used TEXT,
citations TEXT,
duration_s REAL,
success INTEGER,
session_id TEXT,
created_at TEXT
);
CREATE INDEX IF NOT EXISTS idx_task_type ON episodes (task_type, success, created_at);
""")
self.conn.commit()
def record(self, ep: Episode):
self.conn.execute(
"""
INSERT INTO episodes
(task_type, user_query, outcome, tools_used, citations,
duration_s, success, session_id, created_at)
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)
""",
(
ep.task_type, ep.user_query, ep.outcome,
json.dumps(ep.tools_used), json.dumps(ep.citations),
ep.duration_seconds, int(ep.success),
ep.session_id, ep.created_at,
),
)
self.conn.commit()
def recall_similar(self, task_type: str, limit: int = 3) -> list[Episode]:
cutoff = (datetime.utcnow() - timedelta(days=self.EXPIRY_DAYS)).isoformat()
cursor = self.conn.execute(
"""
SELECT task_type, user_query, outcome, tools_used, citations,
duration_s, success, session_id, created_at
FROM episodes
WHERE task_type = ? AND success = 1 AND created_at > ?
ORDER BY created_at DESC
LIMIT ?
""",
(task_type, cutoff, limit),
)
return [
Episode(
task_type=r[0], user_query=r[1], outcome=r[2],
tools_used=json.loads(r[3]), citations=json.loads(r[4]),
duration_seconds=r[5], success=bool(r[6]),
session_id=r[7], created_at=r[8],
)
for r in cursor.fetchall()
]
def validate_citations(self, episode: Episode, validator_fn) -> bool:
"""
validator_fn(citation: str) -> bool
Check if cited sources are still valid (file exists, URL responds, etc.)
Return False if any citation fails; the episode should be discarded.
"""
return all(validator_fn(c) for c in episode.citations)
Layer 3: Semantic memory with Weaviate Hybrid search
When working through the Layer 3 Semantic memory stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Layer 3 Semantic memory stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
import uuid as uuid_lib
from datetime import datetime
import weaviate
from weaviate.classes.config import Configure, Property, DataType, VectorDistances
from weaviate.classes.query import MetadataQuery, HybridFusion, Filter
def embed(text: str) -> list[float]:
"""
Swap in any embedding source: OpenAI, Cohere, sentence-transformers, etc.
Returns a normalized float vector.
"""
from sentence_transformers import SentenceTransformer
_model = SentenceTransformer("all-MiniLM-L6-v2")
return _model.encode(text, normalize_embeddings=True).tolist()
class WeaviateSemanticMemory:
COLLECTION = "AgentMemory"
def __init__(self, host: str = "localhost", port: int = 8080):
self.client = weaviate.connect_to_local(host=host, port=port)
self._ensure_collection()
def _ensure_collection(self):
if self.client.collections.exists(self.COLLECTION):
return
self.client.collections.create(
name=self.COLLECTION,
# We provide our own vectors; no built-in vectorizer needed.
# Swap to Configure.Vectorizer.text2vec_openai() if you prefer managed embedding.
vectorizer_config=Configure.Vectorizer.none(),
vector_index_config=Configure.VectorIndex.hnsw(
distance_metric=VectorDistances.COSINE
),
properties=[
Property(name="content", data_type=DataType.TEXT),
Property(name="task_type", data_type=DataType.TEXT),
Property(name="source", data_type=DataType.TEXT),
Property(name="citations", data_type=DataType.TEXT),
Property(name="session_id", data_type=DataType.TEXT),
Property(name="confidence", data_type=DataType.NUMBER),
Property(name="created_at", data_type=DataType.TEXT),
],
)
def store(
self,
content: str,
task_type: str = "",
source: str = "agent",
citations: str = "",
session_id: str = "",
confidence: float = 1.0,
) -> str:
collection = self.client.collections.get(self.COLLECTION)
doc_id = str(uuid_lib.uuid4())
collection.data.insert(
properties={
"content": content,
"task_type": task_type,
"source": source,
"citations": citations,
"session_id": session_id,
"confidence": confidence,
"created_at": datetime.utcnow().isoformat(),
},
vector=embed(content),
uuid=doc_id,
)
return doc_id
def retrieve_hybrid(
self,
query: str,
n_results: int = 5,
min_confidence: float = 0.6,
task_type: str = None,
) -> list[dict]:
collection = self.client.collections.get(self.COLLECTION)
# Filter by confidence floor and optionally by task type
confidence_filter = Filter.by_property("confidence").greater_or_equal(min_confidence)
if task_type:
active_filter = (
Filter.by_property("task_type").equal(task_type) & confidence_filter
)
else:
active_filter = confidence_filter
results = collection.query.hybrid(
query=query,
vector=embed(query),
limit=n_results,
fusion_type=HybridFusion.RELATIVE_SCORE,
filters=active_filter,
return_metadata=MetadataQuery(score=True),
)
return [
{
"content": obj.properties["content"],
"score": obj.metadata.score,
"confidence": obj.properties.get("confidence", 1.0),
"source": obj.properties.get("source", ""),
"citations": obj.properties.get("citations", ""),
"uuid": str(obj.uuid),
}
for obj in results.objects
]
def update_confidence(self, doc_id: str, new_confidence: float):
collection = self.client.collections.get(self.COLLECTION)
collection.data.update(
uuid=doc_id,
properties={"confidence": new_confidence},
)
def close(self):
self.client.close()
Layer 4: Procedural Memory and the Prompt-evolution pattern
The Layer 4 Procedural Memory stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
import json
import os
import re
from datetime import datetime, timedelta, timezone
from openai import OpenAI
MODEL = "deepseek/deepseek-v4-flash-0731"
client = OpenAI(
base_url="https://openrouter.ai/api/v1",
api_key=os.environ["OPENROUTER_API_KEY"],
)
class ProceduralMemory:
def __init__(self, rules_path: str = "procedural_rules.json"):
self.rules_path = rules_path
self.rules: list[dict] = self._load()
def _load(self) -> list[dict]:
try:
with open(self.rules_path) as f:
return json.load(f)
except FileNotFoundError:
return []
def add_rule(self, situation: str, action: str, reason: str, confidence: float = 1.0):
self.rules.append({
"situation": situation,
"action": action,
"reason": reason,
"confidence": float(confidence),
"added_at": datetime.now(timezone.utc).isoformat(),
"trigger_count": 0,
})
self._save()
def get_applicable_rules(self, context: str, min_confidence: float = 0.7) -> list[dict]:
relevant = []
ctx_lower = context.lower()
dirty = False
for rule in self.rules:
if rule.get("confidence", 1.0) < min_confidence:
continue
keywords = rule["situation"].lower().split()
if not keywords:
continue
hits = sum(1 for kw in keywords if kw in ctx_lower)
if hits >= max(1, len(keywords) // 3):
rule["trigger_count"] = rule.get("trigger_count", 0) + 1
relevant.append(rule)
dirty = True
if dirty:
self._save()
relevant.sort(key=lambda r: r.get("confidence", 1.0), reverse=True)
return relevant[:5]
def prune_stale(self, max_age_days: int = 60, min_triggers: int = 2):
cutoff = (datetime.now(timezone.utc) - timedelta(days=max_age_days)).isoformat()
self.rules = [
r for r in self.rules
if r.get("added_at", "") > cutoff or r.get("trigger_count", 0) >= min_triggers
]
self._save()
def _save(self):
tmp = f"{self.rules_path}.tmp"
with open(tmp, "w") as f:
json.dump(self.rules, f, indent=2)
os.replace(tmp, self.rules_path)
def _parse_json(text: str) -> dict:
text = text.strip()
fenced = re.search(r"```(?:json)?\s*(.*?)```", text, re.S)
if fenced:
text = fenced.group(1).strip()
return json.loads(text)
def extract_rule_from_failure(failure_trace: str, memory: ProceduralMemory):
response = client.chat.completions.create(
model=MODEL,
max_tokens=250,
response_format={"type": "json_object"},
messages=[
{
"role": "system",
"content": "You return only a JSON object. No prose, no code fences.",
},
{
"role": "user",
"content": (
"A task failed. Extract one behavioral rule to prevent this failure.\n"
"Return ONLY valid JSON with keys: situation, action, reason, "
"confidence (0.0-1.0)\n"
f"Failure trace:\n{failure_trace}"
),
},
],
)
try:
rule = _parse_json(response.choices[0].message.content)
except (json.JSONDecodeError, AttributeError, TypeError):
return None
if not all(k in rule for k in ("situation", "action", "reason")):
return None
memory.add_rule(
situation=str(rule["situation"]),
action=str(rule["action"]),
reason=str(rule["reason"]),
confidence=float(rule.get("confidence", 1.0)),
)
return rule
Wiring it together: Async memory manager and System Design
The Wiring it together Async stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
import json
from concurrent.futures import ThreadPoolExecutor
from dataclasses import dataclass, field
@dataclass
class TaskResult:
task_type: str
query: str
outcome: str
tools_used: list
citations: list
duration_seconds: float
success: bool
new_facts: list[dict] = field(default_factory=list)
session_id: str = ""
class MemoryManager:
def __init__(
self,
weaviate_host: str = "localhost",
episodic_db: str = "episodes.db",
rules_path: str = "procedural_rules.json",
):
self.buffer = SummarizingBufferMemory(max_turns=8)
self.episodic = EpisodicMemory(db_path=episodic_db)
self.semantic = WeaviateSemanticMemory(host=weaviate_host)
self.procedural = ProceduralMemory(rules_path=rules_path)
self._pool = ThreadPoolExecutor(max_workers=2, thread_name_prefix="memory_write")
def build_context(self, query: str, task_type: str) -> list[dict]:
context: list[dict] = []
# Procedural rules first: they constrain behavior throughout the task
rules = self.procedural.get_applicable_rules(query)
if rules:
rules_text = "\n".join(
f"- When '{r['situation']}': {r['action']} (reason: {r['reason']})"
for r in rules
)
context.append({"role": "user", "content": f"[Behavioral rules:\n{rules_text}]"})
# Semantic facts: domain knowledge and past discoveries
facts = self.semantic.retrieve_hybrid(query, n_results=5, min_confidence=0.6)
if facts:
facts_text = "\n".join(f"- {f['content']}" for f in facts)
context.append({"role": "user", "content": f"[Relevant knowledge:\n{facts_text}]"})
# Past episodes: outcome templates for similar tasks
episodes = self.episodic.recall_similar(task_type, limit=3)
if episodes:
ep_text = "\n".join(
f"- Outcome: {e.outcome} (tools: {', '.join(e.tools_used)})"
for e in episodes
)
context.append({"role": "user", "content": f"[Past similar tasks:\n{ep_text}]"})
# Current conversation last: the model reads this most carefully
context.extend(self.buffer.get_context())
return context
def record_turn(self, role: str, content: str):
self.buffer.add_message(role, content)
def persist(self, result: TaskResult):
# Submit to thread pool and return immediately; never block the caller
self._pool.submit(self._persist_worker, result)
def _persist_worker(self, result: TaskResult):
ep = Episode(
task_type=result.task_type,
user_query=result.query,
outcome=result.outcome,
tools_used=result.tools_used,
citations=result.citations,
duration_seconds=result.duration_seconds,
success=result.success,
session_id=result.session_id,
)
self.episodic.record(ep)
for fact in result.new_facts:
self.semantic.store(**fact)
if not result.success:
extract_rule_from_failure(result.outcome, self.procedural)
def shutdown(self):
self._pool.shutdown(wait=True)
self.semantic.close()
Three decisions that define your architecture
The Three decisions that define stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Three decisions that define stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Memory poisoning: What the attacks look like and How to defend
For the Memory poisoning What the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
the actual recommendation
For the the actual recommendation stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Let’s keep learning together
For the Let s keep learning stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Let s keep learning stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
More Useful Articles
When working through the More Useful Articles stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Operational checklist
For the Operational checklist stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for b038012e06fc: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.
For the hardening note 0 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 0/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 1 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 1/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 2 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 2/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 3 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 3/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 4 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 4/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 5 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 5/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 6 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 6/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 7 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 7/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 8 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 8/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 9 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 9/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 10 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 10/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 11 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 11/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 12 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 12/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.