首页 / 文章 / 实用笔记:代理架构——第12篇:我们为……做好准备了吗?

实用笔记:代理架构——第12篇:我们为……做好准备了吗?

《实用笔记:代理架构》操作指南——第12篇:我们是否已为采用该模式的团队准备好合同、校验机制以及可直接插入的代码模块?

2821 词

以下笔记围绕“代理架构——第12篇:我们准备好迎接原生代理内存系统了吗?”梳理出一条实用路径。重点在于契约、校验以及可直接插入的代码占位符,而非动机性阐述。 在完成概览阶段时,首先写下契约内容:所需输入、成功信号以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持一致性。 同时记录正常流程与故障恢复路径。重试机制、人工审核环节以及死信处理都是产品本身的组成部分,而非后续的优化内容。

此处包含的内容

“你将发现什么”阶段若被视为可度量的界面,则效果最佳。在扩大范围之前,先记录一份优秀的案例、一个失败实例以及回滚说明。 相较于庞大的脚本,应优先选择小型且可测试的单元。当某一步骤失败时,故障应能指向单一责任主体,而非复杂的流程链。 保持图表状态简洁且具有类型定义。嵌套的数据块会掩盖哪个节点编写了哪个字段的信息,还会在中断后导致无法继续处理。

精确阐述的智能体内存问题

将“智能体记忆问题”阶段视为可测量的对象来处理效果最佳。在扩大范围之前,先记录一份理想的处理结果、一个失败案例以及回滚说明。 把这一阶段视为输入与经过验证的输出之间的契约。为相关成果命名,明确成功标准,绝不允许出现无声的、不完整的处理结果。 保持图结构的状态简洁且具有明确类型。嵌套的数据块会掩盖是哪个节点修改了哪个字段,还会在处理中断后导致无法继续。

+------------------------+-----------------------------+-------------------------+
| Human Memory Type      | Article 7 Implementation    | How It Is Accessed      |
+------------------------+-----------------------------+-------------------------+
| Working memory         | LangGraph state             | Always present in ctx   |
| (active context)       | MemorySaver checkpoints     |                         |
+------------------------+-----------------------------+-------------------------+
| Episodic memory        | DynamoDB + embeddings       | Semantic similarity     |
| (what happened before) | TTL 90 days                 | query at run start      |
+------------------------+-----------------------------+-------------------------+
| Semantic memory        | Bedrock Knowledge Base      | Vector search query     |
| (domain knowledge)     | S3-backed JSONL             | at run start            |
+------------------------+-----------------------------+-------------------------+
| Procedural memory      | DynamoDB validated table    | Task type lookup        |
| (how to do things)     | Success rate tracking       | at run start            |
+------------------------+-----------------------------+-------------------------+

差距1:时间感知能力

将 Gap 1 时间感知阶段视为可测量的对象来处理时效果最佳。在扩大范围之前,先记录一个理想运行案例、一个故障案例以及回滚说明。 在功能结果旁同时记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在系统从演示环境过渡到共享环境时出现意外费用。 保持图结构扁平且类型明确。嵌套的数据块会掩盖哪个节点修改了哪个字段的信息,还会在中断后导致流程无法继续。 将 Gap 1 时间感知阶段视为可测量的对象来处理时效果最佳。在扩大范围之前,先记录一个理想运行案例、一个故障案例以及回滚说明。 需同时记录正常流程与恢复流程。重试机制、人工审核环节以及死信处理都是产品本身的组成部分,而非后续需要补充的功能。

+----------------------------------+------------------------------------------+
| Temporal Property                | What It Enables                          |
+----------------------------------+------------------------------------------+
| Creation timestamp               | Basic recency weighting in retrieval     |
+----------------------------------+------------------------------------------+
| Last confirmed timestamp         | Distinguish stale from fresh knowledge   |
+----------------------------------+------------------------------------------+
| Confidence decay function        | Facts become less certain over time      |
|                                  | without confirmation                     |
+----------------------------------+------------------------------------------+
| Version history                  | Track how understanding of a topic       |
|                                  | has evolved across runs                  |
+----------------------------------+------------------------------------------+
| Temporal context at retrieval    | "What did I know about X on this date?"  |
|                                  | not just "What do I know about X now?"   |
+----------------------------------+------------------------------------------+
# harness/memory/temporal.py
import time
from typing import List
def apply_temporal_weighting(
    retrieved_facts: List[dict],
    recency_half_life_days: float = 30.0,
) -> List[dict]:
    """
    Weights retrieved facts by recency using exponential decay.
    Facts confirmed recently score higher than stale ones with
    the same semantic similarity.
    """
    now = time.time()
    half_life_seconds = recency_half_life_days * 86400
    for fact in retrieved_facts:
        base_score = fact.get("similarity_score", 0.8)
        last_confirmed = fact.get("last_confirmed_at", fact.get("created_at", now))
        age_seconds = now - last_confirmed
        # Exponential decay: score halves every half_life_days
        import math
        decay_factor = math.exp(-0.693 * age_seconds / half_life_seconds)
        fact["temporal_weighted_score"] = base_score * decay_factor
    return sorted(retrieved_facts, key=lambda f: f["temporal_weighted_score"], reverse=True)

Gap 2:关联检索

在 Gap 2 关联检索阶段,修改代码之前需明确输入内容、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。相比冗长的脚本,更应采用小型且可测试的单元。当某个步骤失败时,故障原因应指向单一责任模块,而非复杂的流程链。必须引用实际作为答案依据的段落;没有引用的话,操作人员就无法区分是幻觉内容还是索引缺失所致。

# harness/memory/associative.py
import boto3
from typing import List, Optional
class AssociativeMemoryIndex:
    """
    Maintains an association graph between memory entities.
    Complements vector search with relationship-based retrieval.
    This is a simplified implementation. Production would use
    Amazon Neptune or a graph database for complex traversals.
    """
    def __init__(
        self,
        table_name: str = "agent-memory-associations",
        region: str = "us-east-1",
    ):
        dynamodb = boto3.resource("dynamodb", region_name=region)
        self.table = dynamodb.Table(table_name)
    def record_association(
        self,
        entity_a: str,
        entity_b: str,
        relationship: str,
        strength: float = 1.0,
        run_id: str = None,
    ):
        """
        Records that two memory entities are related.
        entity_a, entity_b: fact_ids, episode_ids, or concept labels
        relationship: "co-occurred", "caused", "contradicts", "supports"
        strength: 0.0 to 1.0, increases with repeated co-occurrence
        """
        import time
        self.table.update_item(
            Key={"entity_a": entity_a, "entity_b": entity_b},
            UpdateExpression=(
                "SET relationship = :r, "
                "strength = if_not_exists(strength, :z) + :s, "
                "occurrence_count = if_not_exists(occurrence_count, :z) + :one, "
                "last_seen = :now"
            ),
            ExpressionAttributeValues={
                ":r": relationship,
                ":z": 0,
                ":s": strength,
                ":one": 1,
                ":now": int(time.time()),
            }
        )
    def get_associated_entities(
        self,
        entity: str,
        min_strength: float = 0.5,
        max_results: int = 10,
    ) -> List[dict]:
        """Retrieves entities associated with the given entity."""
        response = self.table.query(
            KeyConditionExpression="entity_a = :e",
            FilterExpression="strength >= :s",
            ExpressionAttributeValues={":e": entity, ":s": min_strength},
        )
        items = sorted(
            response.get("Items", []),
            key=lambda x: x.get("strength", 0),
            reverse=True
        )
        return items[:max_results]

Gap 3:写入路径智能

在 Gap 3 的 Write-Path Intelligence 阶段,应在修改代码之前明确输入参数、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 将此阶段视为输入与已验证输出之间的契约。为相关成果命名,定义成功判定标准,并拒绝默许的半完成状态。 对于涉及资金支出或修改生产数据的操作,必须经过人工审批。编译时的连接方式并不等同于业务上的完整性。

# harness/memory/annotation.py
from langchain_core.messages import SystemMessage, HumanMessage
from langchain_aws import ChatBedrock
import boto3
import json

MEMORY_ANNOTATION_PROMPT = """
You have just completed a reasoning step. Before continuing, consider:
1. Did you discover something that would be useful in future runs on similar tasks?
2. Did you encounter a pattern you had not seen before?
3. Did something fail that you want to remember to avoid next time?
If yes to any of these, describe what you want to remember in one or two sentences.
If no, respond with null.
Respond with JSON:
{"worth_remembering": true | false, "annotation": "description or null"}
"""

class InlineMemoryAnnotator:
    """
    Runs between agent reasoning steps and asks the agent to flag
    anything worth remembering before the run ends.
    This is experimental. The risk is that the agent annotates
    incorrect conclusions. Pair with the validation gate from Article 7.
    """
    def __init__(self, region: str = "us-east-1"):
        bedrock = boto3.client("bedrock-runtime", region_name=region)
        self.model = ChatBedrock(
            client=bedrock,
            model_id="anthropic.claude-haiku-4-5",
            model_kwargs={"temperature": 0, "max_tokens": 256},
        )
        self._annotations: list = []
    def maybe_annotate(self, last_reasoning_step: str) -> bool:
        """
        Called after each significant reasoning step.
        Returns True if an annotation was recorded.
        """
        response = self.model.invoke([
            SystemMessage(content=MEMORY_ANNOTATION_PROMPT),
            HumanMessage(content=f"Recent reasoning:\n{last_reasoning_step[:1000]}")
        ])
        try:
            result = json.loads(response.content)
            if result.get("worth_remembering") and result.get("annotation"):
                self._annotations.append(result["annotation"])
                return True
        except json.JSONDecodeError:
            pass
        return False
    def get_annotations(self) -> list:
        return list(self._annotations)

Gap 4:具有隔离功能的跨智能体内存

在 Gap 4 跨智能体记忆阶段,修改代码之前需明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 在功能结果旁记录执行时间以及令牌或查询成本。提前显示成本可避免在流程从演示环境转向共享环境时出现意外费用。 对于会消耗资金或修改生产数据的操作,必须经过人工审批。编译时的连接方式并不等同于业务功能的完整性。 在 Gap 4 跨智能体记忆阶段,修改代码之前需明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 需同时记录正常流程和异常恢复流程。重试机制、人工审核环节以及死信处理都是产品功能的一部分,而非后续需要补充的内容。

# harness/memory/shared_memory.py
import boto3
from typing import List, Optional

class SharedMemoryPolicy:
    """
    Controls which agent types can read and write which memory namespaces.
    Sharing is opt-in. Default is isolated per agent_id.
    """
    def __init__(self, policies: dict):
        """
        policies example:
        {
          "security_findings": {
            "readers": ["supervisor", "security_reviewer", "code_analyst"],
            "writers": ["security_reviewer"],
          },
          "code_patterns": {
            "readers": ["supervisor", "code_analyst"],
            "writers": ["code_analyst"],
          }
        }
        """
        self.policies = policies
    def can_read(self, agent_id: str, namespace: str) -> bool:
        policy = self.policies.get(namespace, {})
        return agent_id in policy.get("readers", [])
    def can_write(self, agent_id: str, namespace: str) -> bool:
        policy = self.policies.get(namespace, {})
        return agent_id in policy.get("writers", [])
    def filter_retrievable(
        self,
        agent_id: str,
        records: List[dict],
    ) -> List[dict]:
        """
        Filters a list of memory records to only those the agent can read.
        """
        return [
            r for r in records
            if self.can_read(agent_id, r.get("namespace", "private"))
            or r.get("agent_id") == agent_id  # always read own memories
        ]

第五个差距:将“遗忘”视为一项核心操作

在处理“遗忘”这一阶段时,首先需明确相关约定:所需的输入参数、成功标志,以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 建议采用小型、可测试的单元,而非庞大的脚本。当某个步骤失败时,故障应指向单一的责任模块,而非复杂的流程链。 在成本较高的步骤之后设置检查点。当操作员重新尝试后续节点时,恢复流程不应再次调用相同的大型语言模型。

# harness/memory/intelligent_forgetting.py
import boto3
import time
from typing import List

class IntelligentForgettingManager:
    """
    Manages memory removal based on content-aware signals,
    not just age. Complements the TTL-based decay from Article 7.
    """
    def __init__(
        self,
        episodic_table: str = "agent-episodic-memory",
        semantic_kb_id: str = None,
        region: str = "us-east-1",
    ):
        dynamodb = boto3.resource("dynamodb", region_name=region)
        self.episodic_table = dynamodb.Table(episodic_table)
        self.kb_id = semantic_kb_id
        self.region = region
    def mark_contradicted(
        self,
        fact_id: str,
        contradicting_run_id: str,
        contradiction_description: str,
    ):
        """
        Marks a semantic fact as contradicted by newer evidence.
        Does not delete immediately: flags for review first.
        """
        self.episodic_table.update_item(
            Key={"fact_id": fact_id},
            UpdateExpression=(
                "SET contradicted = :t, "
                "contradicted_by = :run, "
                "contradiction_note = :note, "
                "needs_review = :t"
            ),
            ExpressionAttributeValues={
                ":t": True,
                ":run": contradicting_run_id,
                ":note": contradiction_description,
            }
        )
    def detect_contradictions(
        self,
        new_fact_content: str,
        existing_facts: List[dict],
        detector_model,
    ) -> List[str]:
        """
        Checks whether a new fact contradicts existing ones.
        Returns list of fact_ids that are contradicted.
        """
        if not existing_facts:
            return []
        from langchain_core.messages import SystemMessage, HumanMessage
        import json
        facts_text = "\n".join([
            f"[{f.get('fact_id', 'unknown')}]: {f.get('content', '')}"
            for f in existing_facts
        ])
        response = detector_model.invoke([
            SystemMessage(content="""
You are checking for contradictions between a new fact and existing facts.
A contradiction means the new fact and an existing fact cannot both be true.
Return JSON:
{"contradicted_ids": ["fact_id_1", ...]}
Return empty list if no contradictions found.
"""),
            HumanMessage(content=f"""
New fact: {new_fact_content}
Existing facts:
{facts_text}
""")
        ])
        try:
            result = json.loads(response.content)
            return result.get("contradicted_ids", [])
        except json.JSONDecodeError:
            return []

生态系统正在构建什么

在完成“生态系统是什么”这一阶段时,首先写下契约内容:所需的输入参数、成功标志,以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持一致性。 将这一阶段视为输入与经过验证的输出之间的契约。为相关成果命名,明确成功判定标准,杜绝无声的半完成状态。 在成本较高的步骤之后设置检查点。当操作员重新尝试后续节点时,恢复流程不应再次调用相同的大型语言模型。

这对你今天的开发工作意味着什么

在“这意味着什么”阶段工作时,首先需明确相关约定:所需输入、成功标志以及部分失败时的处理方式。这份清单能确保后续的代码修改保持一致性。 在功能结果旁记录执行时间以及令牌或查询成本。提前了解成本情况,可避免从演示环境过渡到共享环境时出现意外费用。 在耗时较高的步骤之后设置检查点。当操作员重新尝试后续节点时,恢复流程不应再次收取相同的LLM调用费用。 在“这意味着什么”阶段工作时,首先需明确相关约定:所需输入、成功标志以及部分失败时的处理方式。这份清单能确保后续的代码修改保持一致性。 需同时记录正常流程和故障恢复流程。重试机制、人工审核环节以及死信处理都是产品功能的一部分,而非后续需要补充的内容。

生产环境现实检验

将“生产环境现实检验”阶段视为可度量的对象来处理,效果最佳。在扩大范围之前,先记录一份优秀的流程文档、一个失败案例以及回滚说明。 优先选择小型且可测试的单元,而非庞大的脚本。当某个步骤出错时,故障应能指向具体的责任主体,而非复杂的流程链。 保持图表状态简洁且具有类型定义。嵌套的数据块会掩盖哪个节点修改了哪个字段的信息,还会在流程中断后导致无法继续执行。

参考架构

将参考架构阶段视为可度量的对象来处理效果最佳。在扩大范围之前,先记录一份最优案例、一个故障场景以及回滚说明。 把这一阶段视为输入与已验证输出之间的契约。为相关成果命名,明确成功标准,杜绝默许的半完成状态。 保持图结构简洁且类型明确。嵌套的数据块会掩盖哪个节点修改了哪个字段,还会在中断后导致无法继续处理。

  Current State (Article 7)          Future Agent-Native System
  -----------------------            --------------------------
  Run starts                         Run starts
      |                                  |
  Retrieve episodes   <-- embedding      Query temporal memory
  Retrieve KB facts   <-- vector         Traverse association graph
  Retrieve procedures <-- exact key      Agent-selected retrieval
      |                                  |
  Agent runs                         Agent runs
      |                                  |
  End-of-run extractor               Inline annotation (agent)
  stores artifacts                   + end-of-run consolidation
      |                                  |
  DynamoDB episodes                  Temporal-aware store
  KB facts (S3/vector)               Association index
  DynamoDB procedures                Policy-gated sharing
                                     Contradiction detection
                                     Intelligent forgetting

操作检查清单

在操作检查清单阶段,修改代码之前需明确输入内容、该步骤的负责人以及退出标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏的状态。

将配置信息置于应用程序代码之外。环境文件、密钥存储和功能标志应集中存放于一个位置,以便操作人员无需查看整个图结构即可进行审计。

对于会花费资金或修改生产数据的操作,必须经过人工审批。编译时的配置并不等同于业务功能的完整性。

编写一份简短的操作手册:说明如何轮换密钥、如何清空队列、以及如何回滚上一次的导入操作。

同时记录正常流程和故障恢复流程。重试机制、人工审核环节以及死信处理都是产品不可或缺的部分,而非后续才需要补充的功能。

对于会花费资金或修改生产数据的操作,必须经过人工审批。编译时的配置并不等同于业务功能的完整性。

在升级技术栈之前,应先冻结版本,为关键流程保存完整的操作记录,并确认好回滚步骤。共享环境需要设置速率限制、进行租户身份验证,同时要明确负责密钥轮换的人员。与其追求花哨的一次性演示,不如注重扎实的可靠性。

关于8969a47a89be的批量处理说明:不要将提供者密钥放入代码仓库,设定每会话的令牌上限,并将转录内容存储在评估测试用例旁边,以便后续更换模型时仍能保持可比性。

针对强化安全措施的第0阶段,在修改代码之前需明确输入内容、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。同时应在功能结果旁记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外费用。

强化安全措施细节0/837:需测量该步骤的耗时、错误类型以及令牌使用量,然后根据固定的问题集而非主观判断来决定是否保留该更改。

在处理强化措施的第一阶段时,首先写下相关契约:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。

同时记录正常流程和故障恢复流程。重试机制、人工审核环节以及错误处理方式都是产品本身的一部分,而非后续的优化工作。

强化措施细节 1/837:需测量该措施的执行时间、错误类型以及令牌消耗情况,然后依据固定的评估标准而非个人经验来决定是否保留该修改。