Home / Articles / Practical notes: Agentic Architectures — Article 12: Are We Ready For An

This article is published in English.

Practical notes: Agentic Architectures — Article 12: Are We Ready For An

Operable walkthrough of Practical notes: Agentic Architectures — Article 12: Are We Ready For An: contracts, checks, and drop-in code slots for teams shipping this pattern.

2821 words

The following notes reconstruct a practical path around “Agentic Architectures — Article 12: Are We Ready For An Agent-Native Memory System?”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

What You’ll Find Here

The What You ll Find stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

The Agent Memory Problem, Stated Precisely

The The Agent Memory Problem stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

+------------------------+-----------------------------+-------------------------+
| Human Memory Type      | Article 7 Implementation    | How It Is Accessed      |
+------------------------+-----------------------------+-------------------------+
| Working memory         | LangGraph state             | Always present in ctx   |
| (active context)       | MemorySaver checkpoints     |                         |
+------------------------+-----------------------------+-------------------------+
| Episodic memory        | DynamoDB + embeddings       | Semantic similarity     |
| (what happened before) | TTL 90 days                 | query at run start      |
+------------------------+-----------------------------+-------------------------+
| Semantic memory        | Bedrock Knowledge Base      | Vector search query     |
| (domain knowledge)     | S3-backed JSONL             | at run start            |
+------------------------+-----------------------------+-------------------------+
| Procedural memory      | DynamoDB validated table    | Task type lookup        |
| (how to do things)     | Success rate tracking       | at run start            |
+------------------------+-----------------------------+-------------------------+

Gap 1: Temporal Awareness

The Gap 1 Temporal Awareness stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Gap 1 Temporal Awareness stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

+----------------------------------+------------------------------------------+
| Temporal Property                | What It Enables                          |
+----------------------------------+------------------------------------------+
| Creation timestamp               | Basic recency weighting in retrieval     |
+----------------------------------+------------------------------------------+
| Last confirmed timestamp         | Distinguish stale from fresh knowledge   |
+----------------------------------+------------------------------------------+
| Confidence decay function        | Facts become less certain over time      |
|                                  | without confirmation                     |
+----------------------------------+------------------------------------------+
| Version history                  | Track how understanding of a topic       |
|                                  | has evolved across runs                  |
+----------------------------------+------------------------------------------+
| Temporal context at retrieval    | "What did I know about X on this date?"  |
|                                  | not just "What do I know about X now?"   |
+----------------------------------+------------------------------------------+
# harness/memory/temporal.py
import time
from typing import List
def apply_temporal_weighting(
    retrieved_facts: List[dict],
    recency_half_life_days: float = 30.0,
) -> List[dict]:
    """
    Weights retrieved facts by recency using exponential decay.
    Facts confirmed recently score higher than stale ones with
    the same semantic similarity.
    """
    now = time.time()
    half_life_seconds = recency_half_life_days * 86400
    for fact in retrieved_facts:
        base_score = fact.get("similarity_score", 0.8)
        last_confirmed = fact.get("last_confirmed_at", fact.get("created_at", now))
        age_seconds = now - last_confirmed
        # Exponential decay: score halves every half_life_days
        import math
        decay_factor = math.exp(-0.693 * age_seconds / half_life_seconds)
        fact["temporal_weighted_score"] = base_score * decay_factor
    return sorted(retrieved_facts, key=lambda f: f["temporal_weighted_score"], reverse=True)

Gap 2: Associative Retrieval

For the Gap 2 Associative Retrieval stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

# harness/memory/associative.py
import boto3
from typing import List, Optional
class AssociativeMemoryIndex:
    """
    Maintains an association graph between memory entities.
    Complements vector search with relationship-based retrieval.
    This is a simplified implementation. Production would use
    Amazon Neptune or a graph database for complex traversals.
    """
    def __init__(
        self,
        table_name: str = "agent-memory-associations",
        region: str = "us-east-1",
    ):
        dynamodb = boto3.resource("dynamodb", region_name=region)
        self.table = dynamodb.Table(table_name)
    def record_association(
        self,
        entity_a: str,
        entity_b: str,
        relationship: str,
        strength: float = 1.0,
        run_id: str = None,
    ):
        """
        Records that two memory entities are related.
        entity_a, entity_b: fact_ids, episode_ids, or concept labels
        relationship: "co-occurred", "caused", "contradicts", "supports"
        strength: 0.0 to 1.0, increases with repeated co-occurrence
        """
        import time
        self.table.update_item(
            Key={"entity_a": entity_a, "entity_b": entity_b},
            UpdateExpression=(
                "SET relationship = :r, "
                "strength = if_not_exists(strength, :z) + :s, "
                "occurrence_count = if_not_exists(occurrence_count, :z) + :one, "
                "last_seen = :now"
            ),
            ExpressionAttributeValues={
                ":r": relationship,
                ":z": 0,
                ":s": strength,
                ":one": 1,
                ":now": int(time.time()),
            }
        )
    def get_associated_entities(
        self,
        entity: str,
        min_strength: float = 0.5,
        max_results: int = 10,
    ) -> List[dict]:
        """Retrieves entities associated with the given entity."""
        response = self.table.query(
            KeyConditionExpression="entity_a = :e",
            FilterExpression="strength >= :s",
            ExpressionAttributeValues={":e": entity, ":s": min_strength},
        )
        items = sorted(
            response.get("Items", []),
            key=lambda x: x.get("strength", 0),
            reverse=True
        )
        return items[:max_results]

Gap 3: Write-Path Intelligence

For the Gap 3 Write-Path Intelligence stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

# harness/memory/annotation.py
from langchain_core.messages import SystemMessage, HumanMessage
from langchain_aws import ChatBedrock
import boto3
import json

MEMORY_ANNOTATION_PROMPT = """
You have just completed a reasoning step. Before continuing, consider:
1. Did you discover something that would be useful in future runs on similar tasks?
2. Did you encounter a pattern you had not seen before?
3. Did something fail that you want to remember to avoid next time?
If yes to any of these, describe what you want to remember in one or two sentences.
If no, respond with null.
Respond with JSON:
{"worth_remembering": true | false, "annotation": "description or null"}
"""

class InlineMemoryAnnotator:
    """
    Runs between agent reasoning steps and asks the agent to flag
    anything worth remembering before the run ends.
    This is experimental. The risk is that the agent annotates
    incorrect conclusions. Pair with the validation gate from Article 7.
    """
    def __init__(self, region: str = "us-east-1"):
        bedrock = boto3.client("bedrock-runtime", region_name=region)
        self.model = ChatBedrock(
            client=bedrock,
            model_id="anthropic.claude-haiku-4-5",
            model_kwargs={"temperature": 0, "max_tokens": 256},
        )
        self._annotations: list = []
    def maybe_annotate(self, last_reasoning_step: str) -> bool:
        """
        Called after each significant reasoning step.
        Returns True if an annotation was recorded.
        """
        response = self.model.invoke([
            SystemMessage(content=MEMORY_ANNOTATION_PROMPT),
            HumanMessage(content=f"Recent reasoning:\n{last_reasoning_step[:1000]}")
        ])
        try:
            result = json.loads(response.content)
            if result.get("worth_remembering") and result.get("annotation"):
                self._annotations.append(result["annotation"])
                return True
        except json.JSONDecodeError:
            pass
        return False
    def get_annotations(self) -> list:
        return list(self._annotations)

Gap 4: Cross-Agent Memory with Isolation

For the Gap 4 Cross-Agent Memory stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Gap 4 Cross-Agent Memory stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

# harness/memory/shared_memory.py
import boto3
from typing import List, Optional

class SharedMemoryPolicy:
    """
    Controls which agent types can read and write which memory namespaces.
    Sharing is opt-in. Default is isolated per agent_id.
    """
    def __init__(self, policies: dict):
        """
        policies example:
        {
          "security_findings": {
            "readers": ["supervisor", "security_reviewer", "code_analyst"],
            "writers": ["security_reviewer"],
          },
          "code_patterns": {
            "readers": ["supervisor", "code_analyst"],
            "writers": ["code_analyst"],
          }
        }
        """
        self.policies = policies
    def can_read(self, agent_id: str, namespace: str) -> bool:
        policy = self.policies.get(namespace, {})
        return agent_id in policy.get("readers", [])
    def can_write(self, agent_id: str, namespace: str) -> bool:
        policy = self.policies.get(namespace, {})
        return agent_id in policy.get("writers", [])
    def filter_retrievable(
        self,
        agent_id: str,
        records: List[dict],
    ) -> List[dict]:
        """
        Filters a list of memory records to only those the agent can read.
        """
        return [
            r for r in records
            if self.can_read(agent_id, r.get("namespace", "private"))
            or r.get("agent_id") == agent_id  # always read own memories
        ]

Gap 5: Forgetting as a First-Class Operation

When working through the Gap 5 Forgetting as stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

# harness/memory/intelligent_forgetting.py
import boto3
import time
from typing import List

class IntelligentForgettingManager:
    """
    Manages memory removal based on content-aware signals,
    not just age. Complements the TTL-based decay from Article 7.
    """
    def __init__(
        self,
        episodic_table: str = "agent-episodic-memory",
        semantic_kb_id: str = None,
        region: str = "us-east-1",
    ):
        dynamodb = boto3.resource("dynamodb", region_name=region)
        self.episodic_table = dynamodb.Table(episodic_table)
        self.kb_id = semantic_kb_id
        self.region = region
    def mark_contradicted(
        self,
        fact_id: str,
        contradicting_run_id: str,
        contradiction_description: str,
    ):
        """
        Marks a semantic fact as contradicted by newer evidence.
        Does not delete immediately: flags for review first.
        """
        self.episodic_table.update_item(
            Key={"fact_id": fact_id},
            UpdateExpression=(
                "SET contradicted = :t, "
                "contradicted_by = :run, "
                "contradiction_note = :note, "
                "needs_review = :t"
            ),
            ExpressionAttributeValues={
                ":t": True,
                ":run": contradicting_run_id,
                ":note": contradiction_description,
            }
        )
    def detect_contradictions(
        self,
        new_fact_content: str,
        existing_facts: List[dict],
        detector_model,
    ) -> List[str]:
        """
        Checks whether a new fact contradicts existing ones.
        Returns list of fact_ids that are contradicted.
        """
        if not existing_facts:
            return []
        from langchain_core.messages import SystemMessage, HumanMessage
        import json
        facts_text = "\n".join([
            f"[{f.get('fact_id', 'unknown')}]: {f.get('content', '')}"
            for f in existing_facts
        ])
        response = detector_model.invoke([
            SystemMessage(content="""
You are checking for contradictions between a new fact and existing facts.
A contradiction means the new fact and an existing fact cannot both be true.
Return JSON:
{"contradicted_ids": ["fact_id_1", ...]}
Return empty list if no contradictions found.
"""),
            HumanMessage(content=f"""
New fact: {new_fact_content}
Existing facts:
{facts_text}
""")
        ])
        try:
            result = json.loads(response.content)
            return result.get("contradicted_ids", [])
        except json.JSONDecodeError:
            return []

What the Ecosystem Is Building

When working through the What the Ecosystem Is stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

What This Means For How You Build Today

When working through the What This Means For stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the What This Means For stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Production Reality Check

The Production Reality Check stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Reference Architecture

The Reference Architecture stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

  Current State (Article 7)          Future Agent-Native System
  -----------------------            --------------------------
  Run starts                         Run starts
      |                                  |
  Retrieve episodes   <-- embedding      Query temporal memory
  Retrieve KB facts   <-- vector         Traverse association graph
  Retrieve procedures <-- exact key      Agent-selected retrieval
      |                                  |
  Agent runs                         Agent runs
      |                                  |
  End-of-run extractor               Inline annotation (agent)
  stores artifacts                   + end-of-run consolidation
      |                                  |
  DynamoDB episodes                  Temporal-aware store
  KB facts (S3/vector)               Association index
  DynamoDB procedures                Policy-gated sharing
                                     Contradiction detection
                                     Intelligent forgetting

Operational checklist

For the Operational checklist stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.

Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.

Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for 8969a47a89be: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.

For the hardening note 0 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Hardening detail 0/837: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

When working through the hardening note 1 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Hardening detail 1/837: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.