首页 / 文章 / 《实用笔记》:用于生产环境RAG的向量数据库——索引构建、混合搜索

《实用笔记》:用于生产环境RAG的向量数据库——索引构建、混合搜索

《实用笔记》操作指南:用于生产环境RAG的向量数据库——索引构建、混合搜索,以及为采用该模式的团队提供的模板与代码示例。

7876 词

以下笔记为“用于生产环境 RAG 的向量数据库:索引构建、混合搜索与检索扩展”提供了实用的实施路径。重点在于契约定义、校验机制以及可直接插入的代码占位符,而非激励性描述。 在完成概览阶段时,首先明确契约内容:所需输入、成功标志以及部分失败时的处理方式。这样的检查清单能确保后续代码修改的规范性。 将配置信息置于应用程序代码之外。环境文件、密钥存储及功能开关应集中存放于一处,以便操作人员无需查看整个系统结构即可进行审计。

向量数据库即搜索算法

将向量数据库视为可测量的表面时,其阶段化开发效果最佳。在扩大范围之前,先记录一个成功的用例、一个失败案例以及回滚说明。同时记录正常流程与恢复流程的文档。重试机制、人工审核环节以及死信处理都是产品本身的一部分,而非后续需要补充的功能。应将分块策略与检索策略分开,当质量指标发生变化时,修改其中一项不应强制要求重新编写另一项。

精确最近邻搜索:无法在大规模应用中使用的基准方法

当将“精确最近邻搜索”阶段视为可度量的表面时,其效果最佳。在扩大范围之前,先记录一份理想的处理结果、一个失败案例以及回滚说明。 相较于庞大的脚本,应优先选择小型且可测试的单元。当某一步骤失败时,故障应指向单一责任模块,而非复杂的流程链。 应将分块策略与检索策略分开。当质量指标发生变化时,修改其中一项不应迫使重新编写另一项。

import faiss
import numpy as np
from typing import Tuple


def build_exact_index(
    embeddings: np.ndarray,
    use_cosine: bool = True
) -> faiss.IndexFlatIP:
    """
    Build a FAISS flat index for exact nearest neighbour search.

    embeddings: (N, D) float32 array.
    use_cosine: If True, normalises a copy of the embeddings and uses inner
                product (equivalent to cosine similarity). The caller's array
                is not mutated.

    Returns a FAISS flat index. Benchmark latency against your corpus and
    latency SLO before deciding whether ANN indexing is necessary.
    """
    dimension = embeddings.shape[1]

    if use_cosine:
        # Copy before normalising to avoid mutating the caller's array.
        embeddings_copy = embeddings.astype(np.float32).copy()
        faiss.normalize_L2(embeddings_copy)
        index = faiss.IndexFlatIP(dimension)
        index.add(embeddings_copy)
        return index
    else:
        index = faiss.IndexFlatL2(dimension)
        index.add(embeddings.astype(np.float32).copy())
        return index


def search_exact(
    index: faiss.IndexFlatIP,
    query_vector: np.ndarray,
    top_k: int = 10
) -> Tuple[np.ndarray, np.ndarray]:
    """
    Search the flat index. Returns (distances, indices).
    query_vector must already be normalised if the index was built with
    normalised embeddings.
    """
    query = query_vector.reshape(1, -1).astype(np.float32)
    faiss.normalize_L2(query)
    distances, indices = index.search(query, top_k)
    return distances[0], indices[0]

HNSW:为何它在生产环境中的向量搜索中如此常见

HNSW的“为何如此”阶段在被视为可度量的对象时效果最佳。在扩大范围之前,需记录一份理想案例、一个失败案例以及回滚说明。 将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,明确成功标准,杜绝默许的半完成状态。 将分块策略与检索策略分开。当质量指标发生变化时,修改其中一项不应强制要求重新编写另一项。 HNSW的“为何如此”阶段在被视为可度量的对象时效果最佳。在扩大范围之前,需记录一份理想案例、一个失败案例以及回滚说明。 将配置信息置于应用程序代码之外。环境文件、密钥存储和功能开关应集中存放于一处,以便操作人员无需查看整个结构即可进行审计。

图结构的构建方式

在“图表生成方式”阶段,应在修改代码之前明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 需同时记录正常流程和异常恢复流程。重试机制、人工审核环节以及错误处理都是产品本身的组成部分,而非后续需要补充的功能。 必须引用那些为答案提供依据的段落。如果没有引用,操作人员就无法区分是虚假信息还是索引缺失导致的错误。

决定召回率与延迟权衡的参数

在修改代码之前,需先确定定义阶段的参数、输入内容、该步骤的负责人以及退出标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。相比冗长的脚本,更应采用小型且可测试的单元。当某个步骤失败时,故障原因应能明确指向某个具体的责任模块,而非整个复杂的流程。必须引用实际作为答案依据的段落;没有引用的话,操作人员就无法区分是虚假信息还是索引缺失所致。

import faiss
import numpy as np
from typing import Tuple


def build_hnsw_index(
    embeddings: np.ndarray,
    m: int = 32,
    ef_construction: int = 200,
    ef_search: int = 100,
    use_cosine: bool = True
) -> faiss.IndexHNSWFlat:
    """
    Build a FAISS HNSW index for approximate nearest neighbour search.

    m: Graph connectivity parameter. Higher = better recall potential, more memory.
       Starting range for banking policy corpora: 16 to 32. Benchmark your corpus.
    ef_construction: Candidates explored during index build. Higher = better graph quality.
       One-time cost at index build; does not affect query latency.
    ef_search: Candidates explored at query time. Controls recall-latency trade-off.
       Can be changed without rebuilding. Starting range: 50 to 200.
    use_cosine: Normalise embeddings and use inner product (cosine similarity).

    Note: FAISS HNSW does not support GPU acceleration. For GPU-accelerated ANN,
    use IndexIVFPQ variants.
    """
    dimension = embeddings.shape[1]

    # Copy before normalising to avoid mutating the caller's array.
    embeddings_to_index = embeddings.astype(np.float32).copy()

    if use_cosine:
        faiss.normalize_L2(embeddings_to_index)
        index = faiss.IndexHNSWFlat(dimension, m, faiss.METRIC_INNER_PRODUCT)
    else:
        index = faiss.IndexHNSWFlat(dimension, m, faiss.METRIC_L2)

    index.hnsw.efConstruction = ef_construction
    index.hnsw.efSearch = ef_search
    index.add(embeddings_to_index)
    return index


def search_hnsw(
    index: faiss.IndexHNSWFlat,
    query_vector: np.ndarray,
    top_k: int = 10
) -> Tuple[np.ndarray, np.ndarray]:
    """
    Search the HNSW index. Returns (scores, indices).
    query_vector must be normalised if the index was built with normalised embeddings.
    """
    query = query_vector.reshape(1, -1).astype(np.float32)
    faiss.normalize_L2(query)
    scores, indices = index.search(query, top_k)
    return scores[0], indices[0]

HNSW内存需求

在HNSW内存需求阶段,修改代码之前需明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 应将此阶段视为输入与验证后输出之间的契约。为相关成果命名,定义成功判定标准,并拒绝默许的半完成状态。 需引用实际作为答案依据的段落。没有引用的话,操作人员就无法区分幻觉内容与索引缺失问题。 在HNSW内存需求阶段,修改代码之前需明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 将配置信息置于应用程序代码之外。环境文件、密钥存储及功能标志应集中存放于操作人员可审计的位置,无需阅读整个系统结构。

IVF:适用于内存受限环境的倒排文件索引技术

在处理IVF倒排文件索引阶段时,首先明确需求规范:所需的输入参数、成功标志以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改不会偏离原有设计。 需同时记录正常流程与异常恢复路径。重试机制、人工审核环节以及错误消息处理都是产品功能的一部分,而非后续需要补充的内容。 在调整提示词之前,应先使用固定的问题集来测试检索效果。仅仅更换提示词很难改善较差的检索性能。

import faiss
import numpy as np
from typing import Tuple


def build_ivf_index(
    embeddings: np.ndarray,
    nlist: int = 1024,
    nprobe: int = 64,
    use_cosine: bool = True
) -> faiss.IndexIVFFlat:
    """
    Build a FAISS IVF flat index.

    nlist: Number of Voronoi cells. A common starting heuristic is sqrt(N),
           where N is corpus size. For 100K vectors: 300-1000. For 1M: 1024-4096.
           Validate empirically.
    nprobe: Number of cells searched at query time. Higher = better recall, slower.
            Set based on your recall benchmark results.
    use_cosine: Use inner product on normalised vectors.

    Requires training on a representative sample before adding vectors.
    """
    dimension = embeddings.shape[1]
    embeddings_to_index = embeddings.astype(np.float32).copy()

    if use_cosine:
        faiss.normalize_L2(embeddings_to_index)
        quantiser = faiss.IndexFlatIP(dimension)
        index = faiss.IndexIVFFlat(quantiser, dimension, nlist, faiss.METRIC_INNER_PRODUCT)
    else:
        quantiser = faiss.IndexFlatL2(dimension)
        index = faiss.IndexIVFFlat(quantiser, dimension, nlist)

    # Use a random representative sample for training. Using the first N records
    # risks training on a non-representative slice if the corpus is ordered by
    # date, jurisdiction, or document type.
    n_available = len(embeddings_to_index)
    desired_training_size = min(n_available, 40 * nlist)
    rng = np.random.default_rng(seed=42)
    training_indices = rng.choice(n_available, size=desired_training_size, replace=False)
    training_sample = embeddings_to_index[training_indices]
    index.train(training_sample)

    index.nprobe = nprobe
    index.add(embeddings_to_index)
    return index

结合产品量化技术的IVF

在处理产品量化阶段的试管婴儿式开发时,首先需明确合同条款:所需输入、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 相比庞大的脚本,应优先选择小型且可测试的单元。当某个步骤失败时,故障应指向单一责任模块,而非复杂的流程链。 在调整提示词之前,需先用固定的问题集来衡量召回率。仅仅更换提示词很难解决检索效果不佳的问题。

import faiss
import numpy as np
from typing import Tuple


def build_ivfpq_index(
    embeddings: np.ndarray,
    nlist: int = 1024,
    m_subvectors: int = 8,
    bits_per_code: int = 8,
    nprobe: int = 64
) -> Tuple[faiss.IndexIVFPQ, faiss.IndexFlatIP]:
    """
    Build a FAISS IVF-PQ index paired with a flat index for exact re-scoring.

    m_subvectors: Number of sub-vectors. Must divide dimension evenly.
                  For 1536 dimensions: m=8 (192 dims each), m=16 (96 dims each).
                  Select based on the storage-recall trade-off for your corpus.
    bits_per_code: Bits per sub-vector code. 8 bits = 256 centroids per sub-vector.
                   Lower bits = smaller code, larger recall degradation.

    Returns (pq_index, flat_index).
    Use pq_index to retrieve top-N candidates cheaply; use flat_index to re-score
    those candidates with full float32 precision.
    """
    dimension = embeddings.shape[1]
    assert dimension % m_subvectors == 0, (
        f"Dimension {dimension} must be divisible by m_subvectors {m_subvectors}"
    )

    norm_embeddings = embeddings.astype(np.float32).copy()
    faiss.normalize_L2(norm_embeddings)

    # Compressed IVF-PQ index for broad retrieval
    quantiser = faiss.IndexFlatIP(dimension)
    pq_index = faiss.IndexIVFPQ(
        quantiser, dimension, nlist, m_subvectors, bits_per_code,
        faiss.METRIC_INNER_PRODUCT
    )
    training_size = min(len(norm_embeddings), 50 * nlist)
    rng = np.random.default_rng(seed=42)
    training_indices = rng.choice(len(norm_embeddings), size=training_size, replace=False)
    pq_index.train(norm_embeddings[training_indices])
    pq_index.nprobe = nprobe
    pq_index.add(norm_embeddings)

    # Flat index for exact re-scoring of PQ candidates
    flat_index = faiss.IndexFlatIP(dimension)
    flat_index.add(norm_embeddings)

    return pq_index, flat_index


def two_stage_search(
    pq_index: faiss.IndexIVFPQ,
    flat_index: faiss.IndexFlatIP,
    query_vector: np.ndarray,
    top_k: int = 10,
    candidate_multiplier: int = 10
) -> Tuple[np.ndarray, np.ndarray]:
    """
    Two-stage retrieval: broad PQ candidate recall followed by exact flat re-scoring.

    Stage 1: IVF-PQ retrieves top_k * candidate_multiplier candidates cheaply.
    Stage 2: The flat index re-scores those candidates with full float32 precision.

    The flat index must have been built with the same normalised embeddings added
    in the same corpus order so that IVF-PQ indices align to flat index positions.

    candidate_multiplier: Higher values improve recall at higher latency cost.
    """
    query = query_vector.reshape(1, -1).astype(np.float32)
    faiss.normalize_L2(query)

    n_candidates = top_k * candidate_multiplier
    _, candidate_indices = pq_index.search(query, n_candidates)

    valid_mask = candidate_indices[0] >= 0
    valid_candidates = candidate_indices[0][valid_mask]

    if len(valid_candidates) == 0:
        return np.array([]), np.array([])

    # Reconstruct candidate vectors from the flat index and score them exactly.
    candidate_vectors = np.zeros(
        (len(valid_candidates), flat_index.d), dtype=np.float32
    )
    for i, idx in enumerate(valid_candidates):
        flat_index.reconstruct(int(idx), candidate_vectors[i])

    exact_scores = (candidate_vectors @ query.T).flatten()
    reranked_order = np.argsort(exact_scores)[::-1][:top_k]

    final_indices = valid_candidates[reranked_order]
    final_scores = exact_scores[reranked_order]
    return final_scores, final_indices

2026年向量数据库对比

在处理向量数据库对比阶段时,首先需明确相关约定:所需输入、成功标志以及部分失败时的处理方式。这样的清单能确保后续的代码修改始终符合约定。 将此阶段视为输入与验证后输出之间的契约。为相关成果命名,定义成功检测标准,杜绝无声的半完成状态。 在调整提示词之前,先使用固定的问题集来衡量召回率。仅仅更换提示词很难改善较差的检索效果。 在处理向量数据库对比阶段时,首先需明确相关约定:所需输入、成功标志以及部分失败时的处理方式。这样的清单能确保后续的代码修改始终符合约定。 将配置信息置于应用程序代码之外。环境文件、密钥存储以及功能开关应集中存放,以便操作人员无需查看整个系统结构即可进行审计。

FAISS

将 FAISS 阶段视为可度量的表面时,其效果最佳。在扩大范围之前,先记录一个成功的案例、一个失败案例以及回滚说明。同时记录正常流程和恢复流程。重试机制、人工审核环节以及死信处理都是产品本身的一部分,而非后续需要补充的内容。应将分块策略与检索策略分开,当质量指标发生变化时,修改其中一项不应迫使重新编写另一项。

pgvector

将 pgvector 阶段视为可测量的表面来处理时效果最佳。在扩大范围之前,先记录一个成功的案例、一个失败案例以及回滚说明。 优先选择小型且可测试的单元,而非庞大的脚本。当某个步骤失败时,故障应指向单一的责任主体,而非复杂的流程链。 将分块策略与检索策略分开。当质量指标发生变化时,修改其中一项不应迫使重新编写另一项。

-- Enable the pgvector extension
CREATE EXTENSION IF NOT EXISTS vector;

-- Policy chunk table with vector and structured metadata
CREATE TABLE policy_chunks (
    chunk_id          UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    document_id       TEXT NOT NULL,
    document_version  TEXT NOT NULL,
    policy_id         TEXT,
    jurisdiction      TEXT,
    effective_date    DATE,
    section           TEXT,
    content_type      TEXT NOT NULL,
    chunk_text        TEXT NOT NULL,
    classification    TEXT NOT NULL DEFAULT 'INTERNAL',
    permitted_roles   TEXT[] NOT NULL DEFAULT '{}',
    embedding_model   TEXT NOT NULL,
    embedding         vector(1536),
    indexed_at        TIMESTAMPTZ DEFAULT NOW()
);

-- HNSW index for cosine similarity retrieval
CREATE INDEX ON policy_chunks
    USING hnsw (embedding vector_cosine_ops)
    WITH (m = 16, ef_construction = 200);

-- Partial index for jurisdiction-scoped retrieval (common query pattern)
CREATE INDEX ON policy_chunks
    USING hnsw (embedding vector_cosine_ops)
    WHERE jurisdiction = 'EU';

-- Standard indexes for metadata filter columns
CREATE INDEX ON policy_chunks (policy_id);
CREATE INDEX ON policy_chunks (jurisdiction);
CREATE INDEX ON policy_chunks (classification);
CREATE INDEX ON policy_chunks (effective_date);
import psycopg2
import numpy as np
from typing import List, Dict, Optional


def search_policy_chunks(
    query_embedding: List[float],
    jurisdiction: Optional[str] = None,
    classification_ceiling: str = "INTERNAL",
    permitted_role: Optional[str] = None,
    top_k: int = 10,
    ef_search: int = 100,
    connection_string: str = "postgresql://user:password@localhost:5432/rag_db"
) -> List[Dict]:
    """
    Retrieve policy chunks from pgvector with jurisdiction and access filtering.

    ef_search: Controls the HNSW recall-latency trade-off for this session.
               Set per-session; does not require index rebuild.
    """
    conn = psycopg2.connect(connection_string)
    cur = conn.cursor()

    cur.execute(f"SET hnsw.ef_search = {ef_search};")

    classification_levels = {"PUBLIC": 0, "INTERNAL": 1, "CONFIDENTIAL": 2}
    max_level = classification_levels.get(classification_ceiling, 1)
    permitted_classifications = [
        k for k, v in classification_levels.items() if v <= max_level
    ]

    filters = ["classification = ANY(%s)"]
    params: List = [permitted_classifications]

    if jurisdiction:
        filters.append("jurisdiction = %s")
        params.append(jurisdiction)

    if permitted_role:
        filters.append("%s = ANY(permitted_roles) OR cardinality(permitted_roles) = 0")
        params.append(permitted_role)

    where_clause = " AND ".join(filters)
    embedding_str = "[" + ",".join(str(x) for x in query_embedding) + "]"

    query = f"""
        SELECT
            chunk_id,
            document_id,
            document_version,
            policy_id,
            jurisdiction,
            effective_date,
            section,
            content_type,
            chunk_text,
            1 - (embedding <=> %s::vector) AS cosine_similarity
        FROM policy_chunks
        WHERE {where_clause}
        ORDER BY embedding <=> %s::vector
        LIMIT %s;
    """

    params_with_embedding = [embedding_str] + params + [embedding_str, top_k]
    cur.execute(query, params_with_embedding)
    rows = cur.fetchall()

    columns = [
        "chunk_id", "document_id", "document_version", "policy_id",
        "jurisdiction", "effective_date", "section", "content_type",
        "chunk_text", "cosine_similarity"
    ]
    results = [dict(zip(columns, row)) for row in rows]

    cur.close()
    conn.close()
    return results

Qdrant

将 Qdrant 阶段视为可度量的对象时,其效果最佳。在扩大范围之前,需记录一份理想的处理结果、一个失败案例以及回滚说明。 应将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,明确成功标准,杜绝默许的半完成状态。 应将分块策略与检索策略分开。当质量指标发生变化时,修改其中一项不应强制要求重新编写另一项。 将 Qdrant 阶段视为可度量的对象时,其效果最佳。在扩大范围之前,需记录一份理想的处理结果、一个失败案例以及回滚说明。 应将配置置于应用程序代码之外。环境文件、密钥存储和功能开关应集中存放于一处,以便操作人员无需查看整个系统结构即可进行审计。

from qdrant_client import QdrantClient
from qdrant_client.models import (
    VectorParams, Distance, HnswConfigDiff,
    PointStruct, Filter, FieldCondition, MatchValue, MatchAny,
    SparseVectorParams, SparseIndexParams, SparseVector
)
from typing import List, Dict, Optional

client = QdrantClient(host="localhost", port=6333)

COLLECTION_NAME = "banking_policy"
DENSE_VECTOR_NAME = "dense"
SPARSE_VECTOR_NAME = "sparse"


def create_policy_collection(
    dimension: int = 1536,
    m: int = 16,
    ef_construction: int = 200
) -> None:
    """
    Create a Qdrant collection configured for both dense and sparse vectors.
    """
    client.recreate_collection(
        collection_name=COLLECTION_NAME,
        vectors_config={
            DENSE_VECTOR_NAME: VectorParams(
                size=dimension,
                distance=Distance.COSINE,
                hnsw_config=HnswConfigDiff(
                    m=m,
                    ef_construct=ef_construction,
                    full_scan_threshold=10000
                )
            )
        },
        sparse_vectors_config={
            SPARSE_VECTOR_NAME: SparseVectorParams(
                index=SparseIndexParams(on_disk=False)
            )
        }
    )


def upsert_policy_chunks(chunks: List[Dict]) -> None:
    """
    Index policy chunks with dense vectors, sparse vectors, and metadata payloads.

    Each chunk dict must contain:
        chunk_id, dense_vector, sparse_indices, sparse_values,
        chunk_text, document_id, document_version, policy_id,
        jurisdiction, effective_date, content_type, classification,
        permitted_roles
    """
    points = [
        PointStruct(
            id=chunk["chunk_id"],
            vector={
                DENSE_VECTOR_NAME: chunk["dense_vector"],
                SPARSE_VECTOR_NAME: SparseVector(
                    indices=chunk["sparse_indices"],
                    values=chunk["sparse_values"]
                )
            },
            payload={
                "chunk_text": chunk["chunk_text"],
                "document_id": chunk["document_id"],
                "document_version": chunk["document_version"],
                "policy_id": chunk.get("policy_id"),
                "jurisdiction": chunk.get("jurisdiction"),
                "effective_date": chunk.get("effective_date"),
                "content_type": chunk["content_type"],
                "classification": chunk["classification"],
                "permitted_roles": chunk.get("permitted_roles", []),
                "status": "active"
            }
        )
        for chunk in chunks
    ]
    client.upsert(collection_name=COLLECTION_NAME, points=points)


def search_dense_filtered(
    dense_query: List[float],
    jurisdiction: Optional[str] = None,
    permitted_classifications: List[str] = None,
    top_k: int = 10,
    score_threshold: float = 0.3
) -> List[Dict]:
    """
    Dense vector search with integrated payload filtering.
    Filtering is applied inside the HNSW graph traversal, not as a post-filter.
    """
    if permitted_classifications is None:
        permitted_classifications = ["PUBLIC", "INTERNAL"]

    must_conditions = [
        FieldCondition(
            key="classification",
            match=MatchAny(any=permitted_classifications)
        ),
        FieldCondition(key="status", match=MatchValue(value="active"))
    ]

    if jurisdiction:
        must_conditions.append(
            FieldCondition(key="jurisdiction", match=MatchValue(value=jurisdiction))
        )

    search_filter = Filter(must=must_conditions)

    results = client.search(
        collection_name=COLLECTION_NAME,
        query_vector=(DENSE_VECTOR_NAME, dense_query),
        query_filter=search_filter,
        limit=top_k,
        score_threshold=score_threshold,
        with_payload=True
    )

    return [
        {"chunk_id": hit.id, "score": hit.score, **hit.payload}
        for hit in results
    ]


def search_hybrid_qdrant(
    dense_query: List[float],
    sparse_query_indices: List[int],
    sparse_query_values: List[float],
    jurisdiction: Optional[str] = None,
    permitted_classifications: List[str] = None,
    top_k: int = 10
) -> List[Dict]:
    """
    Hybrid search using both dense and sparse vectors with access-control filtering.

    This uses Qdrant's native prefetch-and-fuse API. Both the dense and sparse
    signals contribute to retrieval. The Fusion.RRF strategy applies Reciprocal
    Rank Fusion internally.
    """
    from qdrant_client.models import Prefetch, FusionQuery, Fusion

    if permitted_classifications is None:
        permitted_classifications = ["PUBLIC", "INTERNAL"]

    must_conditions = [
        FieldCondition(
            key="classification",
            match=MatchAny(any=permitted_classifications)
        ),
        FieldCondition(key="status", match=MatchValue(value="active"))
    ]
    if jurisdiction:
        must_conditions.append(
            FieldCondition(key="jurisdiction", match=MatchValue(value=jurisdiction))
        )
    search_filter = Filter(must=must_conditions)

    results = client.query_points(
        collection_name=COLLECTION_NAME,
        prefetch=[
            Prefetch(
                query=dense_query,
                using=DENSE_VECTOR_NAME,
                limit=top_k * 5,
                filter=search_filter
            ),
            Prefetch(
                query=SparseVector(
                    indices=sparse_query_indices,
                    values=sparse_query_values
                ),
                using=SPARSE_VECTOR_NAME,
                limit=top_k * 5,
                filter=search_filter
            ),
        ],
        query=FusionQuery(fusion=Fusion.RRF),
        limit=top_k,
        with_payload=True
    )

    return [
        {"chunk_id": hit.id, "score": hit.score, **hit.payload}
        for hit in results.points
    ]

Weaviate

在 Weaviate 阶段,修改代码之前需先明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 需同时记录正常流程与异常恢复流程。重试机制、人工审核环节以及错误消息处理都是产品本身的组成部分,而非后续需要补充的功能。 必须引用那些真正作为答案依据的段落。如果没有引用,操作人员就无法区分是幻觉内容还是索引缺失导致的错误。

Milvus

在 Milvus 阶段,修改代码之前需明确输入参数、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。相比庞大的脚本,更应采用小型且可测试的单元。当某个步骤失败时,故障原因应能指向单一责任模块,而非复杂的流程链。必须引用实际作为答案依据的段落;没有引用的话,操作人员就无法区分是幻觉内容还是索引缺失所致。

ChromaDB

在处理ChromaDB阶段时,应在修改代码之前明确输入参数、该步骤的负责人以及终止标准。操作员应能够从已知的检查点重新运行该步骤,而无需猜测其中的隐藏状态。 应将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,设定成功检测标准,并杜绝无声的半完成状态。 必须引用实际作为答案依据的段落。没有引用的话,操作员就无法区分是幻觉内容还是索引缺失导致的错误。 在处理ChromaDB阶段时,应在修改代码之前明确输入参数、该步骤的负责人以及终止标准。操作员应能够从已知的检查点重新运行该步骤,而无需猜测其中的隐藏状态。 应将配置信息置于应用程序代码之外。环境文件、密钥存储以及功能开关应集中存放于一个位置,以便操作员无需查看整个系统结构即可进行审计。

Pinecone

在处理 Pinecone 阶段时,首先需写下契约:所需的输入参数、成功信号以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持一致性。 同时记录正常流程与异常恢复路径。重试机制、人工审核环节以及死信处理都是产品本身的组成部分,而非后续需要优化的内容。 在调整提示词之前,先使用固定的问题集来衡量检索覆盖率。仅仅更换提示词很难解决检索效果不佳的问题。

混合搜索:密集检索与稀疏检索的结合

在处理混合搜索的密集组合阶段时,首先需明确相关约定:所需的输入参数、成功标志,以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 相较于庞大的脚本,应优先选择小型且可测试的单元。当某个步骤出现故障时,故障点应指向单一责任模块,而非复杂的流程链。 在调整提示词之前,需先使用固定的问题集来衡量召回率。仅仅更换提示词往往无法改善较差的检索效果。

加权互反排名融合

在处理加权互反排名融合阶段时,首先需明确相关约定:所需输入、成功标志以及部分失败时的处理方式。这样的清单能确保后续的代码修改始终符合约定。 将这一阶段视为输入与验证后输出之间的契约。为相关产物命名,定义成功检测标准,杜绝无声的半完成状态。 在调整提示词之前,先使用固定的问题集来衡量召回率。仅仅更换提示词很难改善较差的检索效果。 在处理加权互反排名融合阶段时,首先需明确相关约定:所需输入、成功标志以及部分失败时的处理方式。这样的清单能确保后续的代码修改始终符合约定。 将配置信息置于应用程序代码之外。环境文件、密钥存储以及功能开关应集中存放于一个位置,以便操作人员无需查看整个系统结构即可进行审计。

from typing import List, Dict, Tuple
from collections import defaultdict


def weighted_reciprocal_rank_fusion(
    dense_results: List[Tuple[str, float]],
    sparse_results: List[Tuple[str, float]],
    k: int = 60,
    dense_weight: float = 0.6,
    sparse_weight: float = 0.4
) -> List[Tuple[str, float]]:
    """
    Fuse dense vector search results with sparse BM25 results using weighted RRF.

    dense_results: List of (chunk_id, dense_score) sorted by dense score descending.
    sparse_results: List of (chunk_id, sparse_score) sorted by sparse score descending.
    k: RRF constant. Higher k reduces the influence of top-ranked documents.
       Conventional default: 60.
    dense_weight / sparse_weight: Relative weights. Tune against your evaluation set.
       For corpora with high-precision identifier queries, increase sparse_weight.

    Returns fused list of (chunk_id, rrf_score) sorted by rrf_score descending.
    """
    rrf_scores: Dict[str, float] = defaultdict(float)

    for rank, (chunk_id, _) in enumerate(dense_results, start=1):
        rrf_scores[chunk_id] += dense_weight * (1.0 / (k + rank))

    for rank, (chunk_id, _) in enumerate(sparse_results, start=1):
        rrf_scores[chunk_id] += sparse_weight * (1.0 / (k + rank))

    return sorted(rrf_scores.items(), key=lambda x: x[1], reverse=True)

调整密集-稀疏权重

将“调整密集-稀疏权重”这一阶段视为可度量的指标来处理效果最佳。在扩大范围之前,先记录一个成功的案例、一个失败案例以及回滚说明。 同时记录正常流程和恢复流程。重试机制、人工审核环节以及死信处理都是产品本身的一部分,而非后续需要补充的内容。 应将分块策略与检索策略分开。当质量指标发生变化时,修改其中一项不应强制要求重新编写另一项。

元数据过滤:检索范围与授权边界

将“元数据过滤检索范围”阶段视为可度量的对象来处理,效果最佳。在扩大检索范围之前,先记录一份理想的处理结果、一个失败案例以及回滚说明。 优先选择小型且可测试的单元,而非庞大的脚本。当某个步骤出现故障时,故障应指向单一责任模块,而非复杂的流程链。 将分块策略与检索策略分开。当质量指标发生变化时,修改其中一项不应迫使重新编写另一项。

from typing import List, Dict, Optional
from enum import Enum


class ClassificationLevel(Enum):
    PUBLIC = 0
    INTERNAL = 1
    CONFIDENTIAL = 2


def build_access_filter(
    user_classification_ceiling: str,
    user_jurisdiction: Optional[str] = None,
    user_roles: Optional[List[str]] = None
) -> Dict:
    """
    Build a Qdrant-compatible filter dict enforcing access control rules.

    user_classification_ceiling: Highest classification the user can see.
    user_jurisdiction: If set, restrict to chunks applicable to that jurisdiction.
    user_roles: If set, restrict to chunks permitted for those roles.

    Integrate with your identity provider at request time, not at index time.
    This filter represents one layer of the authorisation model; it does not
    replace identity verification, audit logging, tenant isolation, or
    downstream response controls.
    """
    ceiling = ClassificationLevel[user_classification_ceiling].value
    permitted = [
        level.name
        for level in ClassificationLevel
        if level.value <= ceiling
    ]

    must_conditions = [
        {"key": "classification", "match": {"any": permitted}},
        {"key": "status", "match": {"value": "active"}}
    ]

    if user_jurisdiction:
        must_conditions.append(
            {"key": "jurisdiction", "match": {"value": user_jurisdiction}}
        )

    if user_roles:
        # Chunks with empty permitted_roles are accessible to all roles.
        must_conditions.append({
            "should": [
                {"key": "permitted_roles", "match": {"any": user_roles}},
                {"is_empty": {"key": "permitted_roles"}}
            ]
        })

    return {"must": must_conditions}

增量索引:无需重建即可添加新文档

将“增量索引:添加新内容”阶段视为可度量的工作面时,其效果最佳。在扩大范围之前,需记录一份理想样本、一个失败案例以及回滚说明。 应将此阶段视为输入与已验证输出之间的契约。为相关成果命名,明确成功标准,杜绝默许的半完成状态。 需将分块策略与检索策略分开。当质量指标发生变化时,修改其中一项不应强制重新编写另一项。 将“增量索引:添加新内容”阶段视为可度量的工作面时,其效果最佳。在扩大范围之前,需记录一份理想样本、一个失败案例以及回滚说明。 应将配置置于应用程序代码之外。环境文件、密钥存储及功能开关应集中存放于一处,以便操作人员无需查看整个系统结构即可进行审计。

import logging
from typing import List, Dict
from datetime import datetime

logger = logging.getLogger(__name__)


class IncrementalIndexManager:
    """
    Manages incremental updates to a Qdrant collection using soft deletion.

    Production pattern:
    1. New chunks are inserted immediately with status='active'.
    2. Superseded chunks are marked status='deleted' (soft delete).
    3. Retrieval filters exclude deleted chunks without graph rebuild.
    4. Full rebuild is triggered on schedule or when deleted fraction exceeds threshold.
    """

    def __init__(self, qdrant_client, collection_name: str):
        self.client = qdrant_client
        self.collection = collection_name
        self.deleted_threshold = 0.15  # Rebuild when 15% of index is soft-deleted

    def upsert_policy_version(
        self,
        new_chunks: List[Dict],
        superseded_chunk_ids: List[str],
        policy_id: str,
        new_version: str
    ) -> Dict:
        """
        Insert new policy version chunks and soft-delete superseded ones.
        """
        if superseded_chunk_ids:
            self.client.set_payload(
                collection_name=self.collection,
                payload={
                    "status": "deleted",
                    "deleted_at": datetime.now().isoformat(),
                    "superseded_by_version": new_version
                },
                points=superseded_chunk_ids
            )
            logger.info(
                f"Soft-deleted {len(superseded_chunk_ids)} chunks "
                f"from policy {policy_id}, superseded by version {new_version}"
            )

        from qdrant_client.models import PointStruct
        points = [
            PointStruct(
                id=chunk["chunk_id"],
                vector={"dense": chunk["dense_vector"]},
                payload={
                    **{k: v for k, v in chunk.items()
                       if k not in ("chunk_id", "dense_vector")},
                    "status": "active",
                    "indexed_at": datetime.now().isoformat()
                }
            )
            for chunk in new_chunks
        ]

        self.client.upsert(collection_name=self.collection, points=points)
        logger.info(
            f"Inserted {len(new_chunks)} chunks for policy {policy_id} version {new_version}"
        )

        return {
            "inserted": len(new_chunks),
            "soft_deleted": len(superseded_chunk_ids),
            "policy_id": policy_id,
            "new_version": new_version
        }

    def should_rebuild(self) -> bool:
        """Check whether the fraction of soft-deleted vectors justifies a full rebuild."""
        from qdrant_client.models import Filter, FieldCondition, MatchValue

        total = self.client.get_collection(self.collection).vectors_count
        deleted_filter = Filter(must=[
            FieldCondition(key="status", match=MatchValue(value="deleted"))
        ])
        deleted_count = self.client.count(
            collection_name=self.collection,
            count_filter=deleted_filter
        ).count

        fraction = deleted_count / total if total > 0 else 0
        logger.info(
            f"Index health: {deleted_count}/{total} soft-deleted ({fraction:.1%})"
        )
        return fraction >= self.deleted_threshold

嵌入模型版本变更

在嵌入模型版本变更阶段,应在修改代码之前明确输入内容、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 需同时记录正常流程和故障恢复流程。重试机制、人工审核环节以及错误处理都是产品本身的组成部分,而非后续需要补充的功能。 当下一步操作为代码编写或工具调用时,应优先使用具有结构化格式且经过模式验证的输出,而非自由形式的文字描述。

分片与复制:实现单节点之外的扩展

在分片与复制扩展阶段,应在修改代码之前明确输入参数、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。相比冗长的脚本,更应采用小型且可测试的单元。当某个步骤失败时,故障原因应能指向单一责任主体,而非复杂的流程链。必须引用实际作为答案依据的段落;没有引用的话,操作人员就无法区分是虚假信息还是索引缺失所致。

分片策略

在分片策略阶段,应在修改代码之前明确输入参数、该步骤的负责人以及终止标准。操作员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 将此阶段视为输入与已验证输出之间的契约。为相关成果命名,定义成功判定标准,并拒绝默许的半完成状态。 需引用实际作为答案依据的段落。没有引用的话,操作员就无法区分幻觉内容与索引缺失问题。 在分片策略阶段,应在修改代码之前明确输入参数、该步骤的负责人以及终止标准。操作员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 将配置信息置于应用程序代码之外。环境文件、密钥存储以及功能开关应集中存放于一个操作员可审计的位置,无需查看整个系统结构。

容量规划

在开展容量规划阶段时,首先列出相关要求:所需输入、成功标志以及部分故障时的处理方式。这样的清单能确保后续的代码修改保持一致性。 同时记录正常流程和恢复流程。重试机制、人工审核环节以及错误处理方式都是产品本身的一部分,而非后续需要补充的内容。 在调整提示词之前,先使用固定的问题集来衡量检索效果。仅仅更换提示词很难解决检索能力不足的问题。

from dataclasses import dataclass


@dataclass
class VectorIndexCapacityPlan:
    """
    Illustrative capacity model for HNSW vector indexes.
    All figures are approximations for planning purposes.
    Benchmark against your actual implementation and workload.
    """
    n_vectors: int
    dimension: int
    hnsw_m: int = 16
    replication_factor: int = 2
    avg_payload_bytes: int = 2048
    memory_headroom_factor: float = 1.5

    def vector_storage_gb(self) -> float:
        return (self.n_vectors * self.dimension * 4) / (1024 ** 3)

    def hnsw_graph_gb(self) -> float:
        # Approximate; actual graph overhead varies by implementation and configuration.
        return (self.n_vectors * self.hnsw_m * 2 * 8) / (1024 ** 3)

    def payload_storage_gb(self) -> float:
        return (self.n_vectors * self.avg_payload_bytes) / (1024 ** 3)

    def total_index_gb(self) -> float:
        return self.vector_storage_gb() + self.hnsw_graph_gb() + self.payload_storage_gb()

    def memory_per_replica_gb(self) -> float:
        return self.total_index_gb() * self.memory_headroom_factor

    def total_cluster_memory_gb(self) -> float:
        # Each replica holds a full copy of the index.
        return self.memory_per_replica_gb() * self.replication_factor

    def report(self) -> str:
        return (
            f"Illustrative capacity model — {self.n_vectors:,} vectors at {self.dimension}d:\n"
            f"  Vector storage (approx):         {self.vector_storage_gb():.2f} GB\n"
            f"  HNSW graph estimate:             {self.hnsw_graph_gb():.2f} GB\n"
            f"  Payload storage (approx):        {self.payload_storage_gb():.2f} GB\n"
            f"  Index footprint before overhead: {self.total_index_gb():.2f} GB\n"
            f"  Per-replica memory + headroom:   {self.memory_per_replica_gb():.2f} GB\n"
            f"  Total cluster memory (approx):   {self.total_cluster_memory_gb():.2f} GB\n"
            f"  ({self.replication_factor} replicas, each holding a full copy)\n"
            f"  Treat these as planning estimates, not deployment guarantees.\n"
            f"  Benchmark against your implementation before provisioning."
        )


# Illustrative example: banking policy corpus
plan = VectorIndexCapacityPlan(
    n_vectors=500_000,
    dimension=1536,
    hnsw_m=16,
    replication_factor=2,
    avg_payload_bytes=2048,
    memory_headroom_factor=1.5
)
print(plan.report())

完整的检索流程

在处理“完整检索流程”阶段时,首先需明确相关约定:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 相比冗长的脚本,应优先选择小型且可测试的单元。当某个步骤出现故障时,故障原因应能指向单一责任模块,而非整个复杂的流程。 在调整提示词之前,需先使用固定的问题集来衡量检索的召回率。仅仅更换提示词往往无法解决检索效果不佳的问题。

User Query
    ↓
Query Embedding (dense + sparse)
    ↓
Metadata / Authorisation Constraints
    ↓
Dense Vector Retrieval (filtered HNSW)
        +
Sparse Retrieval (BM25)
    ↓
Weighted RRF Fusion
    ↓
Candidate Documents with Provenance Metadata
    ↓
[Part 7: Reranking and Context Assembly]
    ↓
LLM Generation
import logging
import re
from typing import List, Dict, Optional
from dataclasses import dataclass
from collections import defaultdict

logger = logging.getLogger(__name__)


@dataclass
class RetrievalConfig:
    dense_candidate_pool: int = 50
    sparse_candidate_pool: int = 50
    rrf_k: int = 60
    dense_weight: float = 0.6
    sparse_weight: float = 0.4
    final_top_k: int = 10
    score_threshold: float = 0.2


class BankingPolicyRetriever:
    """
    Production retrieval pipeline for a regulated banking policy corpus.
    Combines dense vector search, BM25 sparse retrieval, access filtering,
    and weighted RRF score fusion.

    Retrieval ends at the fused candidate list. Reranking and context assembly
    are handled in Part 7.
    """

    def __init__(
        self,
        qdrant_client,
        collection_name: str,
        embedding_pipeline,
        bm25_index,
        chunk_store: Dict[str, Dict],
        config: Optional[RetrievalConfig] = None
    ):
        self.client = qdrant_client
        self.collection = collection_name
        self.embedder = embedding_pipeline
        self.bm25 = bm25_index
        self.chunk_store = chunk_store
        self.config = config or RetrievalConfig()

    def retrieve(
        self,
        query: str,
        user_classification_ceiling: str = "INTERNAL",
        user_jurisdiction: Optional[str] = None,
        user_roles: Optional[List[str]] = None
    ) -> List[Dict]:
        """
        Full hybrid retrieval with access control.

        Returns top-k chunks with provenance metadata, access-filtered
        for the requesting user's classification ceiling and jurisdiction.
        """
        from qdrant_client.models import Filter, FieldCondition, MatchValue, MatchAny

        query_embedding = self.embedder.embed_query(query)

        classification_levels = {"PUBLIC": 0, "INTERNAL": 1, "CONFIDENTIAL": 2}
        ceiling = classification_levels.get(user_classification_ceiling, 1)
        permitted_classifications = [
            k for k, v in classification_levels.items() if v <= ceiling
        ]

        must_conditions = [
            FieldCondition(
                key="classification",
                match=MatchAny(any=permitted_classifications)
            ),
            FieldCondition(key="status", match=MatchValue(value="active"))
        ]
        if user_jurisdiction:
            must_conditions.append(
                FieldCondition(
                    key="jurisdiction",
                    match=MatchValue(value=user_jurisdiction)
                )
            )

        access_filter = Filter(must=must_conditions)

        # Dense vector search with integrated access filtering
        dense_hits = self.client.search(
            collection_name=self.collection,
            query_vector=("dense", query_embedding),
            query_filter=access_filter,
            limit=self.config.dense_candidate_pool,
            score_threshold=self.config.score_threshold,
            with_payload=True
        )
        dense_results = [(hit.id, hit.score) for hit in dense_hits]

        # Sparse BM25 retrieval with post-retrieval access filtering
        tokens = re.findall(r'\b\w+\b', query.lower())
        bm25_scores = self.bm25.get_scores(tokens)
        sparse_ranked = sorted(enumerate(bm25_scores), key=lambda x: x[1], reverse=True)

        chunk_ids = list(self.chunk_store.keys())
        sparse_results = []
        for corpus_idx, score in sparse_ranked:
            if score <= 0 or len(sparse_results) >= self.config.sparse_candidate_pool:
                break
            chunk_id = chunk_ids[corpus_idx]
            chunk_meta = self.chunk_store.get(chunk_id, {})
            if chunk_meta.get("classification") not in permitted_classifications:
                continue
            if user_jurisdiction and chunk_meta.get("jurisdiction") != user_jurisdiction:
                continue
            if chunk_meta.get("status") != "active":
                continue
            sparse_results.append((chunk_id, score))

        # Weighted RRF fusion
        rrf_scores: Dict[str, float] = defaultdict(float)
        k = self.config.rrf_k

        for rank, (chunk_id, _) in enumerate(dense_results, start=1):
            rrf_scores[chunk_id] += self.config.dense_weight * (1.0 / (k + rank))

        for rank, (chunk_id, _) in enumerate(sparse_results, start=1):
            rrf_scores[chunk_id] += self.config.sparse_weight * (1.0 / (k + rank))

        fused = sorted(rrf_scores.items(), key=lambda x: x[1], reverse=True)
        top_chunk_ids = [cid for cid, _ in fused[:self.config.final_top_k]]

        # Assemble results with provenance metadata
        results = []
        for chunk_id in top_chunk_ids:
            chunk_data = self.chunk_store.get(chunk_id, {})
            results.append({
                "chunk_id": chunk_id,
                "rrf_score": rrf_scores[chunk_id],
                **chunk_data
            })

        logger.info(
            f"Retrieval complete: query={query[:60]!r}, "
            f"dense_candidates={len(dense_results)}, "
            f"sparse_candidates={len(sparse_results)}, "
            f"final_results={len(results)}"
        )
        return results

可观测性:生产环境中需监测的内容

在处理“可观测性:需衡量什么”这一阶段时,首先要写明相关契约:所需的输入、成功标志以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持一致性。 将这一阶段视为输入与经过验证的输出之间的契约。为相关产物命名,明确成功判定标准,杜绝无声的半完成状态。 在调整提示词之前,先使用固定的问题集来测试召回率。仅仅更换提示词很难解决检索效果不佳的问题。 在处理“可观测性:需衡量什么”这一阶段时,首先要写明相关契约:所需的输入、成功标志以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持一致性。 将配置信息置于应用程序代码之外。环境文件、密钥存储以及功能开关应集中存放于一个位置,这样操作人员无需查看整个系统结构即可进行审计。

import time
import logging
from typing import Callable, TypeVar, Any
from functools import wraps

logger = logging.getLogger(__name__)
F = TypeVar("F", bound=Callable[..., Any])


def retrieval_instrumented(func: F) -> F:
    """
    Decorator that adds structured latency logging and empty-result alerting
    to retrieval functions. Wrap your primary retrieve() method in production.
    """
    @wraps(func)
    def wrapper(*args, **kwargs):
        start = time.perf_counter()
        result = None
        error = None

        try:
            result = func(*args, **kwargs)
            return result
        except Exception as e:
            error = str(e)
            raise
        finally:
            elapsed_ms = (time.perf_counter() - start) * 1000
            n_results = len(result) if result is not None else 0

            log_payload = {
                "function": func.__name__,
                "latency_ms": round(elapsed_ms, 2),
                "n_results": n_results,
                "error": error
            }

            query = kwargs.get("query", args[1] if len(args) > 1 else None)
            if query:
                log_payload["query_prefix"] = str(query)[:80]

            if error:
                logger.error("retrieval_error", extra=log_payload)
            elif n_results == 0:
                logger.warning("retrieval_empty_result", extra=log_payload)
            elif elapsed_ms > 500:
                logger.warning("retrieval_high_latency", extra=log_payload)
            else:
                logger.info("retrieval_success", extra=log_payload)

    return wrapper  # type: ignore

返回到1200万欧元的信贷申请方案

将“返回到EUR阶段”的流程视为可度量的操作界面最为有效。在扩大范围之前,需记录一个成功的案例、一个失败案例以及回滚说明。同时记录正常流程与恢复流程的细节。重试机制、人工审核环节以及错误处理都属于产品本身的功能,而非后续需要补充的内容。应将分块策略与检索策略分开处理;当质量指标发生变化时,修改其中一项不应强制要求重新编写另一项。

在进入重排序之前:生产环境检查清单

在将“迁移前准备”阶段视为可度量的工作面时,其效果最佳。在扩大范围之前,先记录一个成功的案例、一个失败案例以及回滚说明。 相较于庞大的脚本,应优先选择小型且可测试的单元。当某个步骤失败时,故障应指向单一责任点,而非复杂的流程链。 应将分块策略与检索策略分开。当质量指标发生变化时,修改其中一项不应迫使重新编写另一项。

下一步是什么

“下一步该做什么”阶段若被视为可度量的工作面,效果最佳。在扩大范围之前,需记录一份理想输出样本、一个失败案例以及回滚说明。 将此阶段视为输入与已验证输出之间的契约。为相关成果命名,明确成功标准,杜绝默许的半完成状态。 将分块策略与检索策略分开。当质量指标发生变化时,修改其中一项不应强制重新编写另一项。 “下一步该做什么”阶段若被视为可度量的工作面,效果最佳。在扩大范围之前,需记录一份理想输出样本、一个失败案例以及回滚说明。 将配置信息置于应用程序代码之外。环境文件、密钥存储及功能开关应集中存放于一处,以便操作人员无需查看整个系统结构即可进行审计。

运营检查清单

在操作检查清单阶段,应在修改代码之前明确输入参数、该步骤的负责人以及完成标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。

在功能结果旁记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外费用。

需注明实际作为答案依据的段落。没有引用的话,操作人员就无法区分是幻觉内容还是索引缺失导致的错误。

除了质量指标外,还需跟踪成本和延迟。虽然答案稍差一些,但成本只有原来的十分之一,那可能是适合生产环境的最佳选择。

锁定依赖版本,并记录用于运行演示的图像摘要。可重复性比传统经验更为可靠。

应优先选择小型、可测试的单元,而非庞大的脚本。当某个步骤出错时,故障应指向单一责任模块,而非复杂的流程链。

在推广该技术栈之前,需冻结版本、为关键路径记录标准输出日志,并确认回滚步骤。共享环境需要设置速率限制、租户验证机制,以及明确的密钥轮换负责人。与其追求花哨的一次性演示,不如注重扎实的可靠性。

针对 fa68a70d815a 的批量说明:请将提供方密钥移出代码仓库,设定单会话令牌上限,并将日志存储在评估用示例文件旁,以便后续模型更换时保持数据可比性。

在强化措施的第0阶段,应在修改代码之前明确输入参数、该步骤的负责人以及完成标准。操作人员应能够从已知的检查点重新执行该步骤,而无需猜测隐藏状态。应将此阶段视为输入与验证后输出之间的契约:为相关成果命名,定义成功判定标准,并拒绝默许部分完成的情况。

强化措施细节0/864:需记录该措施的耗时、错误类型以及令牌消耗情况,然后依据固定的评估标准而非主观经验来决定是否保留该变更。

在处理强化措施的第一阶段时,首先写下相关约定:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 应将配置信息与应用程序代码分开存放。环境文件、密钥存储以及功能开关应集中于一个位置,这样操作人员无需查看整个系统结构即可进行审计。

强化措施细节 1/864:针对该措施记录运行时间、错误类型以及令牌消耗情况,然后依据固定的评估标准而非个人经验来决定是否保留该变更。

将强化措施的第二阶段视为可度量的对象来处理效果最佳。在扩大范围之前,先记录一份理想的操作流程、一个失败案例以及回滚说明。 相比复杂的脚本,更应采用小型且可测试的单元。当某个步骤出现故障时,故障原因应能明确指向某个具体的责任模块,而非整个复杂的流程链。

强化措施细节2/864:为该记录测量运行时间、错误类型以及代币消耗情况,然后依据固定的问题集而非个人经验来判断是否保留该变更。

在强化措施的第3阶段,应在修改代码之前明确输入参数、该步骤的负责人以及完成标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。应在功能结果旁记录时间以及代币或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外费用。

强化措施细节3/864:为该记录测量运行时间、错误类型以及代币消耗情况,然后依据固定的问题集而非个人经验来判断是否保留该变更。

在处理强化措施的第4阶段时,首先写下相关契约:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改始终符合要求。

同时记录正常流程和故障恢复流程。重试机制、人工审核环节以及错误处理方式都是产品本身的组成部分,而非后续才添加的完善措施。

强化措施的第4/864项细节要求:测量该环节的耗时、错误类型以及令牌使用情况,然后依据固定的评估标准而非个人经验来决定是否保留该变更。

将强化措施的第5阶段视为一个可量化的目标面最为有效。在扩大范围之前,先记录一份理想的操作日志、一个故障案例以及回滚说明。

要把这一阶段视为输入参数与验证后输出结果之间的契约。为相关文档命名,明确成功判定标准,杜绝无声的半完成状态。

强化措施细节5/864:为该记录测量运行时间、错误类型以及令牌消耗情况,然后依据固定的问题集而非个人经验来判断是否保留该变更。

在强化措施的第6阶段,应在修改代码之前明确输入参数、该步骤的负责人以及完成标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。配置应置于应用程序代码之外,环境文件、密钥存储以及功能标志应集中存放于一个操作人员可以审核的位置,无需查看整个系统结构。

强化措施细节6/864:为该记录测量运行时间、错误类型以及令牌消耗情况,然后依据固定的问题集而非个人经验来判断是否保留该变更。

在处理强化措施的第7阶段时,首先写下相关约定:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改始终符合预期。 相比冗长的脚本,更应采用小型且可测试的单元。当某个步骤失败时,故障应能指向单一责任点,而非复杂的流程链。

强化措施细节7/864:需测量该步骤的运行时间、错误类型以及代币消耗情况,然后依据固定的评估标准而非主观感受来决定是否保留该修改。

将强化措施的第8阶段视为可量化的目标面处理效果最佳。在扩大范围之前,先记录一份理想的运行示例、一个失败案例以及回滚说明。 在功能结果旁同时记录时间消耗及代币或查询成本。提前明确成本情况,可避免在从演示环境过渡到共享环境时出现意外支出。

强化措施细节 8/864:为该记录测量运行时间、错误类型以及令牌消耗情况,然后依据固定的问题集而非个别案例来决定是否保留该变更。

在强化措施的第9阶段,应在修改代码之前明确输入参数、该步骤的负责人以及完成标准。操作人员应能够从已知的检查点重新执行该步骤,而无需猜测隐藏状态。需同时记录正常流程和故障恢复流程。重试机制、人工审核环节以及错误处理方式都是产品本身的组成部分,而非后续的优化内容。

强化措施细节 9/864:为该记录测量运行时间、错误类型以及令牌消耗情况,然后依据固定的问题集而非个别案例来决定是否保留该变更。

在处理强化措施的第10阶段时,首先写下相关契约:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改始终符合约定。 将这一阶段视为输入与验证后输出之间的契约。为相关产物命名,明确成功判定标准,杜绝默许部分完成的情况。

强化措施细节10/864:需测量该步骤的耗时、错误类型以及令牌消耗情况,然后依据固定的评估标准而非个人经验来决定是否保留该变更。

将强化措施的第11阶段视为可度量的对象来处理效果最佳。在扩大范围之前,先记录一份标准操作流程、一个失败案例以及回滚说明。 应将配置信息与应用程序代码分开。环境文件、密钥存储和功能开关应集中存放于一处,以便操作人员无需查看全部代码结构即可进行审计。

强化措施细节11/864:测量该任务的执行时间、错误类型以及代币消耗情况,然后依据固定的问题集而非个人经验来判断是否保留该变更。

在强化措施的第12阶段,应在修改代码之前明确输入参数、该步骤的负责人以及完成标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。相比复杂的脚本,更应采用小型且可测试的单元。当某一步骤失败时,故障原因应能指向单一责任方,而非混乱的整个流程。

强化措施细节12/864:测量该任务的执行时间、错误类型以及代币消耗情况,然后依据固定的问题集而非个人经验来判断是否保留该变更。

在处理强化措施第13阶段时,首先写下相关约定:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 在功能结果旁记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在系统从演示环境过渡到共享环境时出现意外费用。

强化措施细节13/864:为该阶段测量实际执行时间、错误类型以及令牌消耗情况,然后依据固定的评估标准而非个人经验来决定是否保留相关修改。

将强化措施第14阶段视为可度量的对象来处理效果最佳。在扩大范围之前,先记录一个理想运行案例、一个失败案例以及回滚说明。 同时记录正常流程和故障恢复流程。重试机制、人工审核环节以及死信处理都是产品本身的组成部分,而非后续需要补充的内容。

强化细节14/864:测量该笔记的处理时间、错误类型以及令牌消耗情况,然后依据固定的问题集而非个别案例来决定是否保留该变更。