Home / Articles / Practical notes: All You Want For Agentic Memory Design

This article is published in English.

Practical notes: All You Want For Agentic Memory Design

Operable walkthrough of Practical notes: All You Want For Agentic Memory Design: contracts, checks, and drop-in code slots for teams shipping this pattern.

4588 words

The following notes reconstruct a practical path around “All You Want For Agentic Memory Design”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Table of Contents

The Table of Contents stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

The Agent kept forgetting

The The Agent kept forgetting stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Why Context windows fail (and the study that proved it)

The Why Context windows fail stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Why Context windows fail stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Short-term vs Long-term Memory: The Foundational split

For the Short-term vs Long-term Memory stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

How the industry converged on four memory types

For the How the industry converged stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Production Database decisions: SQL vs Vector vs Graph

For the Production Database decisions SQL stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Production Database decisions SQL stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Layer 1: Buffer memory with Summarization

When working through the Layer 1 Buffer memory stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

from collections import deque
from typing import List, Dict

class BufferMemory:
    def __init__(self, max_turns: int = 10):
        # Each turn = one user message + one assistant message
        self.history: deque = deque(maxlen=max_turns * 2)

    def add_message(self, role: str, content: str):
        self.history.append({"role": role, "content": content})

    def get_context(self) -> List[Dict]:
        return list(self.history)

    def token_estimate(self) -> int:
        total_chars = sum(len(m["content"]) for m in self.history)
        return total_chars // 4  # rough approximation

    def clear(self):
        self.history.clear()
import os
from collections import deque
from typing import List, Dict

from openai import OpenAI


class SummarizingBufferMemory:

    def __init__(self, max_turns: int = 8, summary_batch: int = 4):
        self.client = OpenAI(
            base_url="https://openrouter.ai/api/v1",
            api_key=os.environ["OPENROUTER_API_KEY"],
        )
        self.model = "deepseek/deepseek-v4-flash-0731"
        self.recent: deque = deque(maxlen=max_turns * 2)
        self.rolling_summary: str = ""
        self.summary_batch = summary_batch

    def add_message(self, role: str, content: str):
        if len(self.recent) == self.recent.maxlen:
            self._compress_oldest()
        self.recent.append({"role": role, "content": content})

    def _compress_oldest(self):
        batch = [self.recent.popleft() for _ in range(min(self.summary_batch * 2, len(self.recent)))]
        text = "\n".join(f"{m['role']}: {m['content']}" for m in batch)
        response = self.client.chat.completions.create(
            model=self.model,
            max_tokens=300,
            messages=[
                {
                    "role": "system",
                    "content": "You compress conversation history. Reply with the summary only.",
                },
                {
                    "role": "user",
                    "content": (
                        "Summarize this conversation segment in 2-4 sentences. "
                        "Preserve any decisions made, constraints stated, and conclusions reached.\n\n"
                        f"{text}"
                    ),
                },
            ],
        )
        new_summary = response.choices[0].message.content.strip()
        if self.rolling_summary:
            self.rolling_summary = f"{self.rolling_summary} | {new_summary}"
        else:
            self.rolling_summary = new_summary

    def get_context(self) -> List[Dict]:
        context = []
        if self.rolling_summary:
            context.append({
                "role": "system",
                "content": f"[Prior conversation summary: {self.rolling_summary}]",
            })
        context.extend(list(self.recent))
        return context

Layer 2: Episodic memory with Citation validation

When working through the Layer 2 Episodic memory stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

import json
import sqlite3
from datetime import datetime, timedelta
from dataclasses import dataclass, field

@dataclass
class Episode:
    task_type: str
    user_query: str
    outcome: str
    tools_used: list
    citations: list  # source references for validation
    duration_seconds: float
    success: bool
    session_id: str = ""
    created_at: str = field(default_factory=lambda: datetime.utcnow().isoformat())

class EpisodicMemory:
    EXPIRY_DAYS = 28  # GitHub Copilot's production default

    def __init__(self, db_path: str = "episodes.db"):
        self.conn = sqlite3.connect(db_path, check_same_thread=False)
        self._init_schema()

    def _init_schema(self):
        self.conn.executescript("""
            CREATE TABLE IF NOT EXISTS episodes (
                id          INTEGER PRIMARY KEY AUTOINCREMENT,
                task_type   TEXT NOT NULL,
                user_query  TEXT,
                outcome     TEXT,
                tools_used  TEXT,
                citations   TEXT,
                duration_s  REAL,
                success     INTEGER,
                session_id  TEXT,
                created_at  TEXT
            );
            CREATE INDEX IF NOT EXISTS idx_task_type ON episodes (task_type, success, created_at);
        """)
        self.conn.commit()

    def record(self, ep: Episode):
        self.conn.execute(
            """
            INSERT INTO episodes
                (task_type, user_query, outcome, tools_used, citations,
                 duration_s, success, session_id, created_at)
            VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)
            """,
            (
                ep.task_type, ep.user_query, ep.outcome,
                json.dumps(ep.tools_used), json.dumps(ep.citations),
                ep.duration_seconds, int(ep.success),
                ep.session_id, ep.created_at,
            ),
        )
        self.conn.commit()

    def recall_similar(self, task_type: str, limit: int = 3) -> list[Episode]:
        cutoff = (datetime.utcnow() - timedelta(days=self.EXPIRY_DAYS)).isoformat()
        cursor = self.conn.execute(
            """
            SELECT task_type, user_query, outcome, tools_used, citations,
                   duration_s, success, session_id, created_at
            FROM   episodes
            WHERE  task_type = ? AND success = 1 AND created_at > ?
            ORDER  BY created_at DESC
            LIMIT  ?
            """,
            (task_type, cutoff, limit),
        )
        return [
            Episode(
                task_type=r[0], user_query=r[1], outcome=r[2],
                tools_used=json.loads(r[3]), citations=json.loads(r[4]),
                duration_seconds=r[5], success=bool(r[6]),
                session_id=r[7], created_at=r[8],
            )
            for r in cursor.fetchall()
        ]

    def validate_citations(self, episode: Episode, validator_fn) -> bool:
        """
        validator_fn(citation: str) -> bool
        Check if cited sources are still valid (file exists, URL responds, etc.)
        Return False if any citation fails; the episode should be discarded.
        """
        return all(validator_fn(c) for c in episode.citations)

Layer 3: Semantic memory with Weaviate Hybrid search

When working through the Layer 3 Semantic memory stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Layer 3 Semantic memory stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

import uuid as uuid_lib
from datetime import datetime
import weaviate
from weaviate.classes.config import Configure, Property, DataType, VectorDistances
from weaviate.classes.query import MetadataQuery, HybridFusion, Filter

def embed(text: str) -> list[float]:
    """
    Swap in any embedding source: OpenAI, Cohere, sentence-transformers, etc.
    Returns a normalized float vector.
    """
    from sentence_transformers import SentenceTransformer
    _model = SentenceTransformer("all-MiniLM-L6-v2")
    return _model.encode(text, normalize_embeddings=True).tolist()

class WeaviateSemanticMemory:
    COLLECTION = "AgentMemory"

    def __init__(self, host: str = "localhost", port: int = 8080):
        self.client = weaviate.connect_to_local(host=host, port=port)
        self._ensure_collection()

    def _ensure_collection(self):
        if self.client.collections.exists(self.COLLECTION):
            return
        self.client.collections.create(
            name=self.COLLECTION,
            # We provide our own vectors; no built-in vectorizer needed.
            # Swap to Configure.Vectorizer.text2vec_openai() if you prefer managed embedding.
            vectorizer_config=Configure.Vectorizer.none(),
            vector_index_config=Configure.VectorIndex.hnsw(
                distance_metric=VectorDistances.COSINE
            ),
            properties=[
                Property(name="content",     data_type=DataType.TEXT),
                Property(name="task_type",   data_type=DataType.TEXT),
                Property(name="source",      data_type=DataType.TEXT),
                Property(name="citations",   data_type=DataType.TEXT),
                Property(name="session_id",  data_type=DataType.TEXT),
                Property(name="confidence",  data_type=DataType.NUMBER),
                Property(name="created_at",  data_type=DataType.TEXT),
            ],
        )

    def store(
        self,
        content: str,
        task_type: str = "",
        source: str = "agent",
        citations: str = "",
        session_id: str = "",
        confidence: float = 1.0,
    ) -> str:
        collection = self.client.collections.get(self.COLLECTION)
        doc_id = str(uuid_lib.uuid4())
        collection.data.insert(
            properties={
                "content":    content,
                "task_type":  task_type,
                "source":     source,
                "citations":  citations,
                "session_id": session_id,
                "confidence": confidence,
                "created_at": datetime.utcnow().isoformat(),
            },
            vector=embed(content),
            uuid=doc_id,
        )
        return doc_id

    def retrieve_hybrid(
        self,
        query: str,
        n_results: int = 5,
        min_confidence: float = 0.6,
        task_type: str = None,
    ) -> list[dict]:
        collection = self.client.collections.get(self.COLLECTION)
        # Filter by confidence floor and optionally by task type
        confidence_filter = Filter.by_property("confidence").greater_or_equal(min_confidence)
        if task_type:
            active_filter = (
                Filter.by_property("task_type").equal(task_type) & confidence_filter
            )
        else:
            active_filter = confidence_filter
        results = collection.query.hybrid(
            query=query,
            vector=embed(query),
            limit=n_results,
            fusion_type=HybridFusion.RELATIVE_SCORE,
            filters=active_filter,
            return_metadata=MetadataQuery(score=True),
        )
        return [
            {
                "content":    obj.properties["content"],
                "score":      obj.metadata.score,
                "confidence": obj.properties.get("confidence", 1.0),
                "source":     obj.properties.get("source", ""),
                "citations":  obj.properties.get("citations", ""),
                "uuid":       str(obj.uuid),
            }
            for obj in results.objects
        ]

    def update_confidence(self, doc_id: str, new_confidence: float):
        collection = self.client.collections.get(self.COLLECTION)
        collection.data.update(
            uuid=doc_id,
            properties={"confidence": new_confidence},
        )

    def close(self):
        self.client.close()

Layer 4: Procedural Memory and the Prompt-evolution pattern

The Layer 4 Procedural Memory stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

import json
import os
import re
from datetime import datetime, timedelta, timezone

from openai import OpenAI

MODEL = "deepseek/deepseek-v4-flash-0731"

client = OpenAI(
    base_url="https://openrouter.ai/api/v1",
    api_key=os.environ["OPENROUTER_API_KEY"],
)


class ProceduralMemory:
    def __init__(self, rules_path: str = "procedural_rules.json"):
        self.rules_path = rules_path
        self.rules: list[dict] = self._load()

    def _load(self) -> list[dict]:
        try:
            with open(self.rules_path) as f:
                return json.load(f)
        except FileNotFoundError:
            return []

    def add_rule(self, situation: str, action: str, reason: str, confidence: float = 1.0):
        self.rules.append({
            "situation":     situation,
            "action":        action,
            "reason":        reason,
            "confidence":    float(confidence),
            "added_at":      datetime.now(timezone.utc).isoformat(),
            "trigger_count": 0,
        })
        self._save()

    def get_applicable_rules(self, context: str, min_confidence: float = 0.7) -> list[dict]:
        relevant = []
        ctx_lower = context.lower()
        dirty = False
        for rule in self.rules:
            if rule.get("confidence", 1.0) < min_confidence:
                continue
            keywords = rule["situation"].lower().split()
            if not keywords:
                continue
            hits = sum(1 for kw in keywords if kw in ctx_lower)
            if hits >= max(1, len(keywords) // 3):
                rule["trigger_count"] = rule.get("trigger_count", 0) + 1
                relevant.append(rule)
                dirty = True
        if dirty:
            self._save()
        relevant.sort(key=lambda r: r.get("confidence", 1.0), reverse=True)
        return relevant[:5]

    def prune_stale(self, max_age_days: int = 60, min_triggers: int = 2):
        cutoff = (datetime.now(timezone.utc) - timedelta(days=max_age_days)).isoformat()
        self.rules = [
            r for r in self.rules
            if r.get("added_at", "") > cutoff or r.get("trigger_count", 0) >= min_triggers
        ]
        self._save()

    def _save(self):
        tmp = f"{self.rules_path}.tmp"
        with open(tmp, "w") as f:
            json.dump(self.rules, f, indent=2)
        os.replace(tmp, self.rules_path)


def _parse_json(text: str) -> dict:
    text = text.strip()
    fenced = re.search(r"```(?:json)?\s*(.*?)```", text, re.S)
    if fenced:
        text = fenced.group(1).strip()
    return json.loads(text)


def extract_rule_from_failure(failure_trace: str, memory: ProceduralMemory):
    response = client.chat.completions.create(
        model=MODEL,
        max_tokens=250,
        response_format={"type": "json_object"},
        messages=[
            {
                "role": "system",
                "content": "You return only a JSON object. No prose, no code fences.",
            },
            {
                "role": "user",
                "content": (
                    "A task failed. Extract one behavioral rule to prevent this failure.\n"
                    "Return ONLY valid JSON with keys: situation, action, reason, "
                    "confidence (0.0-1.0)\n"
                    f"Failure trace:\n{failure_trace}"
                ),
            },
        ],
    )

    try:
        rule = _parse_json(response.choices[0].message.content)
    except (json.JSONDecodeError, AttributeError, TypeError):
        return None

    if not all(k in rule for k in ("situation", "action", "reason")):
        return None

    memory.add_rule(
        situation=str(rule["situation"]),
        action=str(rule["action"]),
        reason=str(rule["reason"]),
        confidence=float(rule.get("confidence", 1.0)),
    )
    return rule

Wiring it together: Async memory manager and System Design

The Wiring it together Async stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

import json
from concurrent.futures import ThreadPoolExecutor
from dataclasses import dataclass, field

@dataclass
class TaskResult:
    task_type: str
    query: str
    outcome: str
    tools_used: list
    citations: list
    duration_seconds: float
    success: bool
    new_facts: list[dict] = field(default_factory=list)
    session_id: str = ""

class MemoryManager:
    def __init__(
        self,
        weaviate_host: str = "localhost",
        episodic_db: str = "episodes.db",
        rules_path: str = "procedural_rules.json",
    ):
        self.buffer    = SummarizingBufferMemory(max_turns=8)
        self.episodic  = EpisodicMemory(db_path=episodic_db)
        self.semantic  = WeaviateSemanticMemory(host=weaviate_host)
        self.procedural = ProceduralMemory(rules_path=rules_path)
        self._pool = ThreadPoolExecutor(max_workers=2, thread_name_prefix="memory_write")

    def build_context(self, query: str, task_type: str) -> list[dict]:
        context: list[dict] = []
        # Procedural rules first: they constrain behavior throughout the task
        rules = self.procedural.get_applicable_rules(query)
        if rules:
            rules_text = "\n".join(
                f"- When '{r['situation']}': {r['action']} (reason: {r['reason']})"
                for r in rules
            )
            context.append({"role": "user", "content": f"[Behavioral rules:\n{rules_text}]"})
        # Semantic facts: domain knowledge and past discoveries
        facts = self.semantic.retrieve_hybrid(query, n_results=5, min_confidence=0.6)
        if facts:
            facts_text = "\n".join(f"- {f['content']}" for f in facts)
            context.append({"role": "user", "content": f"[Relevant knowledge:\n{facts_text}]"})
        # Past episodes: outcome templates for similar tasks
        episodes = self.episodic.recall_similar(task_type, limit=3)
        if episodes:
            ep_text = "\n".join(
                f"- Outcome: {e.outcome} (tools: {', '.join(e.tools_used)})"
                for e in episodes
            )
            context.append({"role": "user", "content": f"[Past similar tasks:\n{ep_text}]"})
        # Current conversation last: the model reads this most carefully
        context.extend(self.buffer.get_context())
        return context

    def record_turn(self, role: str, content: str):
        self.buffer.add_message(role, content)

    def persist(self, result: TaskResult):
        # Submit to thread pool and return immediately; never block the caller
        self._pool.submit(self._persist_worker, result)

    def _persist_worker(self, result: TaskResult):
        ep = Episode(
            task_type=result.task_type,
            user_query=result.query,
            outcome=result.outcome,
            tools_used=result.tools_used,
            citations=result.citations,
            duration_seconds=result.duration_seconds,
            success=result.success,
            session_id=result.session_id,
        )
        self.episodic.record(ep)
        for fact in result.new_facts:
            self.semantic.store(**fact)
        if not result.success:
            extract_rule_from_failure(result.outcome, self.procedural)

  def shutdown(self):
        self._pool.shutdown(wait=True)
        self.semantic.close()

Three decisions that define your architecture

The Three decisions that define stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Three decisions that define stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Memory poisoning: What the attacks look like and How to defend

For the Memory poisoning What the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

the actual recommendation

For the the actual recommendation stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Let’s keep learning together

For the Let s keep learning stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Let s keep learning stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

More Useful Articles

When working through the More Useful Articles stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Operational checklist

For the Operational checklist stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.

Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.

Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for b038012e06fc: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.

For the hardening note 0 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Hardening detail 0/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

When working through the hardening note 1 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Hardening detail 1/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 2 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Hardening detail 2/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 3 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Hardening detail 3/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

When working through the hardening note 4 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Hardening detail 4/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 5 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Hardening detail 5/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 6 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Hardening detail 6/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

When working through the hardening note 7 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Hardening detail 7/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 8 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Hardening detail 8/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 9 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Hardening detail 9/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

When working through the hardening note 10 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Hardening detail 10/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 11 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Hardening detail 11/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 12 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Hardening detail 12/804: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.