This article is published in English.
Practical notes: Vector Databases for Production RAG: Indexing, Hybrid Search
Operable walkthrough of Practical notes: Vector Databases for Production RAG: Indexing, Hybrid Search: contracts, checks, and drop-in code slots for teams shipping this pattern.
The following notes reconstruct a practical path around “Vector Databases for Production RAG: Indexing, Hybrid Search, and Scaling Retrieval”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
The Vector Database Is the Search Algorithm
The The Vector Database Is stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Exact Nearest Neighbour Search: The Baseline You Cannot Use at Scale
The Exact Nearest Neighbour Search stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
import faiss
import numpy as np
from typing import Tuple
def build_exact_index(
embeddings: np.ndarray,
use_cosine: bool = True
) -> faiss.IndexFlatIP:
"""
Build a FAISS flat index for exact nearest neighbour search.
embeddings: (N, D) float32 array.
use_cosine: If True, normalises a copy of the embeddings and uses inner
product (equivalent to cosine similarity). The caller's array
is not mutated.
Returns a FAISS flat index. Benchmark latency against your corpus and
latency SLO before deciding whether ANN indexing is necessary.
"""
dimension = embeddings.shape[1]
if use_cosine:
# Copy before normalising to avoid mutating the caller's array.
embeddings_copy = embeddings.astype(np.float32).copy()
faiss.normalize_L2(embeddings_copy)
index = faiss.IndexFlatIP(dimension)
index.add(embeddings_copy)
return index
else:
index = faiss.IndexFlatL2(dimension)
index.add(embeddings.astype(np.float32).copy())
return index
def search_exact(
index: faiss.IndexFlatIP,
query_vector: np.ndarray,
top_k: int = 10
) -> Tuple[np.ndarray, np.ndarray]:
"""
Search the flat index. Returns (distances, indices).
query_vector must already be normalised if the index was built with
normalised embeddings.
"""
query = query_vector.reshape(1, -1).astype(np.float32)
faiss.normalize_L2(query)
distances, indices = index.search(query, top_k)
return distances[0], indices[0]
HNSW: Why It Is So Common in Production Vector Search
The HNSW Why It Is stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The HNSW Why It Is stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
How the Graph Is Built
For the How the Graph Is stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
The Parameters That Define Your Recall-Latency Trade-off
For the The Parameters That Define stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
import faiss
import numpy as np
from typing import Tuple
def build_hnsw_index(
embeddings: np.ndarray,
m: int = 32,
ef_construction: int = 200,
ef_search: int = 100,
use_cosine: bool = True
) -> faiss.IndexHNSWFlat:
"""
Build a FAISS HNSW index for approximate nearest neighbour search.
m: Graph connectivity parameter. Higher = better recall potential, more memory.
Starting range for banking policy corpora: 16 to 32. Benchmark your corpus.
ef_construction: Candidates explored during index build. Higher = better graph quality.
One-time cost at index build; does not affect query latency.
ef_search: Candidates explored at query time. Controls recall-latency trade-off.
Can be changed without rebuilding. Starting range: 50 to 200.
use_cosine: Normalise embeddings and use inner product (cosine similarity).
Note: FAISS HNSW does not support GPU acceleration. For GPU-accelerated ANN,
use IndexIVFPQ variants.
"""
dimension = embeddings.shape[1]
# Copy before normalising to avoid mutating the caller's array.
embeddings_to_index = embeddings.astype(np.float32).copy()
if use_cosine:
faiss.normalize_L2(embeddings_to_index)
index = faiss.IndexHNSWFlat(dimension, m, faiss.METRIC_INNER_PRODUCT)
else:
index = faiss.IndexHNSWFlat(dimension, m, faiss.METRIC_L2)
index.hnsw.efConstruction = ef_construction
index.hnsw.efSearch = ef_search
index.add(embeddings_to_index)
return index
def search_hnsw(
index: faiss.IndexHNSWFlat,
query_vector: np.ndarray,
top_k: int = 10
) -> Tuple[np.ndarray, np.ndarray]:
"""
Search the HNSW index. Returns (scores, indices).
query_vector must be normalised if the index was built with normalised embeddings.
"""
query = query_vector.reshape(1, -1).astype(np.float32)
faiss.normalize_L2(query)
scores, indices = index.search(query, top_k)
return scores[0], indices[0]
HNSW Memory Requirements
For the HNSW Memory Requirements stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the HNSW Memory Requirements stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
IVF: Inverted File Indexing for Memory-Constrained Deployments
When working through the IVF Inverted File Indexing stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
import faiss
import numpy as np
from typing import Tuple
def build_ivf_index(
embeddings: np.ndarray,
nlist: int = 1024,
nprobe: int = 64,
use_cosine: bool = True
) -> faiss.IndexIVFFlat:
"""
Build a FAISS IVF flat index.
nlist: Number of Voronoi cells. A common starting heuristic is sqrt(N),
where N is corpus size. For 100K vectors: 300-1000. For 1M: 1024-4096.
Validate empirically.
nprobe: Number of cells searched at query time. Higher = better recall, slower.
Set based on your recall benchmark results.
use_cosine: Use inner product on normalised vectors.
Requires training on a representative sample before adding vectors.
"""
dimension = embeddings.shape[1]
embeddings_to_index = embeddings.astype(np.float32).copy()
if use_cosine:
faiss.normalize_L2(embeddings_to_index)
quantiser = faiss.IndexFlatIP(dimension)
index = faiss.IndexIVFFlat(quantiser, dimension, nlist, faiss.METRIC_INNER_PRODUCT)
else:
quantiser = faiss.IndexFlatL2(dimension)
index = faiss.IndexIVFFlat(quantiser, dimension, nlist)
# Use a random representative sample for training. Using the first N records
# risks training on a non-representative slice if the corpus is ordered by
# date, jurisdiction, or document type.
n_available = len(embeddings_to_index)
desired_training_size = min(n_available, 40 * nlist)
rng = np.random.default_rng(seed=42)
training_indices = rng.choice(n_available, size=desired_training_size, replace=False)
training_sample = embeddings_to_index[training_indices]
index.train(training_sample)
index.nprobe = nprobe
index.add(embeddings_to_index)
return index
IVF with Product Quantisation
When working through the IVF with Product Quantisation stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
import faiss
import numpy as np
from typing import Tuple
def build_ivfpq_index(
embeddings: np.ndarray,
nlist: int = 1024,
m_subvectors: int = 8,
bits_per_code: int = 8,
nprobe: int = 64
) -> Tuple[faiss.IndexIVFPQ, faiss.IndexFlatIP]:
"""
Build a FAISS IVF-PQ index paired with a flat index for exact re-scoring.
m_subvectors: Number of sub-vectors. Must divide dimension evenly.
For 1536 dimensions: m=8 (192 dims each), m=16 (96 dims each).
Select based on the storage-recall trade-off for your corpus.
bits_per_code: Bits per sub-vector code. 8 bits = 256 centroids per sub-vector.
Lower bits = smaller code, larger recall degradation.
Returns (pq_index, flat_index).
Use pq_index to retrieve top-N candidates cheaply; use flat_index to re-score
those candidates with full float32 precision.
"""
dimension = embeddings.shape[1]
assert dimension % m_subvectors == 0, (
f"Dimension {dimension} must be divisible by m_subvectors {m_subvectors}"
)
norm_embeddings = embeddings.astype(np.float32).copy()
faiss.normalize_L2(norm_embeddings)
# Compressed IVF-PQ index for broad retrieval
quantiser = faiss.IndexFlatIP(dimension)
pq_index = faiss.IndexIVFPQ(
quantiser, dimension, nlist, m_subvectors, bits_per_code,
faiss.METRIC_INNER_PRODUCT
)
training_size = min(len(norm_embeddings), 50 * nlist)
rng = np.random.default_rng(seed=42)
training_indices = rng.choice(len(norm_embeddings), size=training_size, replace=False)
pq_index.train(norm_embeddings[training_indices])
pq_index.nprobe = nprobe
pq_index.add(norm_embeddings)
# Flat index for exact re-scoring of PQ candidates
flat_index = faiss.IndexFlatIP(dimension)
flat_index.add(norm_embeddings)
return pq_index, flat_index
def two_stage_search(
pq_index: faiss.IndexIVFPQ,
flat_index: faiss.IndexFlatIP,
query_vector: np.ndarray,
top_k: int = 10,
candidate_multiplier: int = 10
) -> Tuple[np.ndarray, np.ndarray]:
"""
Two-stage retrieval: broad PQ candidate recall followed by exact flat re-scoring.
Stage 1: IVF-PQ retrieves top_k * candidate_multiplier candidates cheaply.
Stage 2: The flat index re-scores those candidates with full float32 precision.
The flat index must have been built with the same normalised embeddings added
in the same corpus order so that IVF-PQ indices align to flat index positions.
candidate_multiplier: Higher values improve recall at higher latency cost.
"""
query = query_vector.reshape(1, -1).astype(np.float32)
faiss.normalize_L2(query)
n_candidates = top_k * candidate_multiplier
_, candidate_indices = pq_index.search(query, n_candidates)
valid_mask = candidate_indices[0] >= 0
valid_candidates = candidate_indices[0][valid_mask]
if len(valid_candidates) == 0:
return np.array([]), np.array([])
# Reconstruct candidate vectors from the flat index and score them exactly.
candidate_vectors = np.zeros(
(len(valid_candidates), flat_index.d), dtype=np.float32
)
for i, idx in enumerate(valid_candidates):
flat_index.reconstruct(int(idx), candidate_vectors[i])
exact_scores = (candidate_vectors @ query.T).flatten()
reranked_order = np.argsort(exact_scores)[::-1][:top_k]
final_indices = valid_candidates[reranked_order]
final_scores = exact_scores[reranked_order]
return final_scores, final_indices
Vector Database Comparison for 2026
When working through the Vector Database Comparison for stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Vector Database Comparison for stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
FAISS
The FAISS stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
pgvector
The pgvector stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
-- Enable the pgvector extension
CREATE EXTENSION IF NOT EXISTS vector;
-- Policy chunk table with vector and structured metadata
CREATE TABLE policy_chunks (
chunk_id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
document_id TEXT NOT NULL,
document_version TEXT NOT NULL,
policy_id TEXT,
jurisdiction TEXT,
effective_date DATE,
section TEXT,
content_type TEXT NOT NULL,
chunk_text TEXT NOT NULL,
classification TEXT NOT NULL DEFAULT 'INTERNAL',
permitted_roles TEXT[] NOT NULL DEFAULT '{}',
embedding_model TEXT NOT NULL,
embedding vector(1536),
indexed_at TIMESTAMPTZ DEFAULT NOW()
);
-- HNSW index for cosine similarity retrieval
CREATE INDEX ON policy_chunks
USING hnsw (embedding vector_cosine_ops)
WITH (m = 16, ef_construction = 200);
-- Partial index for jurisdiction-scoped retrieval (common query pattern)
CREATE INDEX ON policy_chunks
USING hnsw (embedding vector_cosine_ops)
WHERE jurisdiction = 'EU';
-- Standard indexes for metadata filter columns
CREATE INDEX ON policy_chunks (policy_id);
CREATE INDEX ON policy_chunks (jurisdiction);
CREATE INDEX ON policy_chunks (classification);
CREATE INDEX ON policy_chunks (effective_date);
import psycopg2
import numpy as np
from typing import List, Dict, Optional
def search_policy_chunks(
query_embedding: List[float],
jurisdiction: Optional[str] = None,
classification_ceiling: str = "INTERNAL",
permitted_role: Optional[str] = None,
top_k: int = 10,
ef_search: int = 100,
connection_string: str = "postgresql://user:password@localhost:5432/rag_db"
) -> List[Dict]:
"""
Retrieve policy chunks from pgvector with jurisdiction and access filtering.
ef_search: Controls the HNSW recall-latency trade-off for this session.
Set per-session; does not require index rebuild.
"""
conn = psycopg2.connect(connection_string)
cur = conn.cursor()
cur.execute(f"SET hnsw.ef_search = {ef_search};")
classification_levels = {"PUBLIC": 0, "INTERNAL": 1, "CONFIDENTIAL": 2}
max_level = classification_levels.get(classification_ceiling, 1)
permitted_classifications = [
k for k, v in classification_levels.items() if v <= max_level
]
filters = ["classification = ANY(%s)"]
params: List = [permitted_classifications]
if jurisdiction:
filters.append("jurisdiction = %s")
params.append(jurisdiction)
if permitted_role:
filters.append("%s = ANY(permitted_roles) OR cardinality(permitted_roles) = 0")
params.append(permitted_role)
where_clause = " AND ".join(filters)
embedding_str = "[" + ",".join(str(x) for x in query_embedding) + "]"
query = f"""
SELECT
chunk_id,
document_id,
document_version,
policy_id,
jurisdiction,
effective_date,
section,
content_type,
chunk_text,
1 - (embedding <=> %s::vector) AS cosine_similarity
FROM policy_chunks
WHERE {where_clause}
ORDER BY embedding <=> %s::vector
LIMIT %s;
"""
params_with_embedding = [embedding_str] + params + [embedding_str, top_k]
cur.execute(query, params_with_embedding)
rows = cur.fetchall()
columns = [
"chunk_id", "document_id", "document_version", "policy_id",
"jurisdiction", "effective_date", "section", "content_type",
"chunk_text", "cosine_similarity"
]
results = [dict(zip(columns, row)) for row in rows]
cur.close()
conn.close()
return results
Qdrant
The Qdrant stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Qdrant stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
from qdrant_client import QdrantClient
from qdrant_client.models import (
VectorParams, Distance, HnswConfigDiff,
PointStruct, Filter, FieldCondition, MatchValue, MatchAny,
SparseVectorParams, SparseIndexParams, SparseVector
)
from typing import List, Dict, Optional
client = QdrantClient(host="localhost", port=6333)
COLLECTION_NAME = "banking_policy"
DENSE_VECTOR_NAME = "dense"
SPARSE_VECTOR_NAME = "sparse"
def create_policy_collection(
dimension: int = 1536,
m: int = 16,
ef_construction: int = 200
) -> None:
"""
Create a Qdrant collection configured for both dense and sparse vectors.
"""
client.recreate_collection(
collection_name=COLLECTION_NAME,
vectors_config={
DENSE_VECTOR_NAME: VectorParams(
size=dimension,
distance=Distance.COSINE,
hnsw_config=HnswConfigDiff(
m=m,
ef_construct=ef_construction,
full_scan_threshold=10000
)
)
},
sparse_vectors_config={
SPARSE_VECTOR_NAME: SparseVectorParams(
index=SparseIndexParams(on_disk=False)
)
}
)
def upsert_policy_chunks(chunks: List[Dict]) -> None:
"""
Index policy chunks with dense vectors, sparse vectors, and metadata payloads.
Each chunk dict must contain:
chunk_id, dense_vector, sparse_indices, sparse_values,
chunk_text, document_id, document_version, policy_id,
jurisdiction, effective_date, content_type, classification,
permitted_roles
"""
points = [
PointStruct(
id=chunk["chunk_id"],
vector={
DENSE_VECTOR_NAME: chunk["dense_vector"],
SPARSE_VECTOR_NAME: SparseVector(
indices=chunk["sparse_indices"],
values=chunk["sparse_values"]
)
},
payload={
"chunk_text": chunk["chunk_text"],
"document_id": chunk["document_id"],
"document_version": chunk["document_version"],
"policy_id": chunk.get("policy_id"),
"jurisdiction": chunk.get("jurisdiction"),
"effective_date": chunk.get("effective_date"),
"content_type": chunk["content_type"],
"classification": chunk["classification"],
"permitted_roles": chunk.get("permitted_roles", []),
"status": "active"
}
)
for chunk in chunks
]
client.upsert(collection_name=COLLECTION_NAME, points=points)
def search_dense_filtered(
dense_query: List[float],
jurisdiction: Optional[str] = None,
permitted_classifications: List[str] = None,
top_k: int = 10,
score_threshold: float = 0.3
) -> List[Dict]:
"""
Dense vector search with integrated payload filtering.
Filtering is applied inside the HNSW graph traversal, not as a post-filter.
"""
if permitted_classifications is None:
permitted_classifications = ["PUBLIC", "INTERNAL"]
must_conditions = [
FieldCondition(
key="classification",
match=MatchAny(any=permitted_classifications)
),
FieldCondition(key="status", match=MatchValue(value="active"))
]
if jurisdiction:
must_conditions.append(
FieldCondition(key="jurisdiction", match=MatchValue(value=jurisdiction))
)
search_filter = Filter(must=must_conditions)
results = client.search(
collection_name=COLLECTION_NAME,
query_vector=(DENSE_VECTOR_NAME, dense_query),
query_filter=search_filter,
limit=top_k,
score_threshold=score_threshold,
with_payload=True
)
return [
{"chunk_id": hit.id, "score": hit.score, **hit.payload}
for hit in results
]
def search_hybrid_qdrant(
dense_query: List[float],
sparse_query_indices: List[int],
sparse_query_values: List[float],
jurisdiction: Optional[str] = None,
permitted_classifications: List[str] = None,
top_k: int = 10
) -> List[Dict]:
"""
Hybrid search using both dense and sparse vectors with access-control filtering.
This uses Qdrant's native prefetch-and-fuse API. Both the dense and sparse
signals contribute to retrieval. The Fusion.RRF strategy applies Reciprocal
Rank Fusion internally.
"""
from qdrant_client.models import Prefetch, FusionQuery, Fusion
if permitted_classifications is None:
permitted_classifications = ["PUBLIC", "INTERNAL"]
must_conditions = [
FieldCondition(
key="classification",
match=MatchAny(any=permitted_classifications)
),
FieldCondition(key="status", match=MatchValue(value="active"))
]
if jurisdiction:
must_conditions.append(
FieldCondition(key="jurisdiction", match=MatchValue(value=jurisdiction))
)
search_filter = Filter(must=must_conditions)
results = client.query_points(
collection_name=COLLECTION_NAME,
prefetch=[
Prefetch(
query=dense_query,
using=DENSE_VECTOR_NAME,
limit=top_k * 5,
filter=search_filter
),
Prefetch(
query=SparseVector(
indices=sparse_query_indices,
values=sparse_query_values
),
using=SPARSE_VECTOR_NAME,
limit=top_k * 5,
filter=search_filter
),
],
query=FusionQuery(fusion=Fusion.RRF),
limit=top_k,
with_payload=True
)
return [
{"chunk_id": hit.id, "score": hit.score, **hit.payload}
for hit in results.points
]
Weaviate
For the Weaviate stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Milvus
For the Milvus stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
ChromaDB
For the ChromaDB stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the ChromaDB stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Pinecone
When working through the Pinecone stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Hybrid Search: Combining Dense and Sparse Retrieval
When working through the Hybrid Search Combining Dense stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Weighted Reciprocal Rank Fusion
When working through the Weighted Reciprocal Rank Fusion stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Weighted Reciprocal Rank Fusion stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
from typing import List, Dict, Tuple
from collections import defaultdict
def weighted_reciprocal_rank_fusion(
dense_results: List[Tuple[str, float]],
sparse_results: List[Tuple[str, float]],
k: int = 60,
dense_weight: float = 0.6,
sparse_weight: float = 0.4
) -> List[Tuple[str, float]]:
"""
Fuse dense vector search results with sparse BM25 results using weighted RRF.
dense_results: List of (chunk_id, dense_score) sorted by dense score descending.
sparse_results: List of (chunk_id, sparse_score) sorted by sparse score descending.
k: RRF constant. Higher k reduces the influence of top-ranked documents.
Conventional default: 60.
dense_weight / sparse_weight: Relative weights. Tune against your evaluation set.
For corpora with high-precision identifier queries, increase sparse_weight.
Returns fused list of (chunk_id, rrf_score) sorted by rrf_score descending.
"""
rrf_scores: Dict[str, float] = defaultdict(float)
for rank, (chunk_id, _) in enumerate(dense_results, start=1):
rrf_scores[chunk_id] += dense_weight * (1.0 / (k + rank))
for rank, (chunk_id, _) in enumerate(sparse_results, start=1):
rrf_scores[chunk_id] += sparse_weight * (1.0 / (k + rank))
return sorted(rrf_scores.items(), key=lambda x: x[1], reverse=True)
Tuning the Dense-Sparse Weight
The Tuning the Dense-Sparse Weight stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Metadata Filtering: Retrieval Scope and the Authorisation Boundary
The Metadata Filtering Retrieval Scope stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
from typing import List, Dict, Optional
from enum import Enum
class ClassificationLevel(Enum):
PUBLIC = 0
INTERNAL = 1
CONFIDENTIAL = 2
def build_access_filter(
user_classification_ceiling: str,
user_jurisdiction: Optional[str] = None,
user_roles: Optional[List[str]] = None
) -> Dict:
"""
Build a Qdrant-compatible filter dict enforcing access control rules.
user_classification_ceiling: Highest classification the user can see.
user_jurisdiction: If set, restrict to chunks applicable to that jurisdiction.
user_roles: If set, restrict to chunks permitted for those roles.
Integrate with your identity provider at request time, not at index time.
This filter represents one layer of the authorisation model; it does not
replace identity verification, audit logging, tenant isolation, or
downstream response controls.
"""
ceiling = ClassificationLevel[user_classification_ceiling].value
permitted = [
level.name
for level in ClassificationLevel
if level.value <= ceiling
]
must_conditions = [
{"key": "classification", "match": {"any": permitted}},
{"key": "status", "match": {"value": "active"}}
]
if user_jurisdiction:
must_conditions.append(
{"key": "jurisdiction", "match": {"value": user_jurisdiction}}
)
if user_roles:
# Chunks with empty permitted_roles are accessible to all roles.
must_conditions.append({
"should": [
{"key": "permitted_roles", "match": {"any": user_roles}},
{"is_empty": {"key": "permitted_roles"}}
]
})
return {"must": must_conditions}
Incremental Indexing: Adding New Documents Without Rebuilding
The Incremental Indexing Adding New stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Incremental Indexing Adding New stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
import logging
from typing import List, Dict
from datetime import datetime
logger = logging.getLogger(__name__)
class IncrementalIndexManager:
"""
Manages incremental updates to a Qdrant collection using soft deletion.
Production pattern:
1. New chunks are inserted immediately with status='active'.
2. Superseded chunks are marked status='deleted' (soft delete).
3. Retrieval filters exclude deleted chunks without graph rebuild.
4. Full rebuild is triggered on schedule or when deleted fraction exceeds threshold.
"""
def __init__(self, qdrant_client, collection_name: str):
self.client = qdrant_client
self.collection = collection_name
self.deleted_threshold = 0.15 # Rebuild when 15% of index is soft-deleted
def upsert_policy_version(
self,
new_chunks: List[Dict],
superseded_chunk_ids: List[str],
policy_id: str,
new_version: str
) -> Dict:
"""
Insert new policy version chunks and soft-delete superseded ones.
"""
if superseded_chunk_ids:
self.client.set_payload(
collection_name=self.collection,
payload={
"status": "deleted",
"deleted_at": datetime.now().isoformat(),
"superseded_by_version": new_version
},
points=superseded_chunk_ids
)
logger.info(
f"Soft-deleted {len(superseded_chunk_ids)} chunks "
f"from policy {policy_id}, superseded by version {new_version}"
)
from qdrant_client.models import PointStruct
points = [
PointStruct(
id=chunk["chunk_id"],
vector={"dense": chunk["dense_vector"]},
payload={
**{k: v for k, v in chunk.items()
if k not in ("chunk_id", "dense_vector")},
"status": "active",
"indexed_at": datetime.now().isoformat()
}
)
for chunk in new_chunks
]
self.client.upsert(collection_name=self.collection, points=points)
logger.info(
f"Inserted {len(new_chunks)} chunks for policy {policy_id} version {new_version}"
)
return {
"inserted": len(new_chunks),
"soft_deleted": len(superseded_chunk_ids),
"policy_id": policy_id,
"new_version": new_version
}
def should_rebuild(self) -> bool:
"""Check whether the fraction of soft-deleted vectors justifies a full rebuild."""
from qdrant_client.models import Filter, FieldCondition, MatchValue
total = self.client.get_collection(self.collection).vectors_count
deleted_filter = Filter(must=[
FieldCondition(key="status", match=MatchValue(value="deleted"))
])
deleted_count = self.client.count(
collection_name=self.collection,
count_filter=deleted_filter
).count
fraction = deleted_count / total if total > 0 else 0
logger.info(
f"Index health: {deleted_count}/{total} soft-deleted ({fraction:.1%})"
)
return fraction >= self.deleted_threshold
Embedding Model Version Changes
For the Embedding Model Version Changes stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
Sharding and Replication: Scaling Beyond a Single Node
For the Sharding and Replication Scaling stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Sharding Strategies
For the Sharding Strategies stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Sharding Strategies stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Capacity Planning
When working through the Capacity Planning stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
from dataclasses import dataclass
@dataclass
class VectorIndexCapacityPlan:
"""
Illustrative capacity model for HNSW vector indexes.
All figures are approximations for planning purposes.
Benchmark against your actual implementation and workload.
"""
n_vectors: int
dimension: int
hnsw_m: int = 16
replication_factor: int = 2
avg_payload_bytes: int = 2048
memory_headroom_factor: float = 1.5
def vector_storage_gb(self) -> float:
return (self.n_vectors * self.dimension * 4) / (1024 ** 3)
def hnsw_graph_gb(self) -> float:
# Approximate; actual graph overhead varies by implementation and configuration.
return (self.n_vectors * self.hnsw_m * 2 * 8) / (1024 ** 3)
def payload_storage_gb(self) -> float:
return (self.n_vectors * self.avg_payload_bytes) / (1024 ** 3)
def total_index_gb(self) -> float:
return self.vector_storage_gb() + self.hnsw_graph_gb() + self.payload_storage_gb()
def memory_per_replica_gb(self) -> float:
return self.total_index_gb() * self.memory_headroom_factor
def total_cluster_memory_gb(self) -> float:
# Each replica holds a full copy of the index.
return self.memory_per_replica_gb() * self.replication_factor
def report(self) -> str:
return (
f"Illustrative capacity model — {self.n_vectors:,} vectors at {self.dimension}d:\n"
f" Vector storage (approx): {self.vector_storage_gb():.2f} GB\n"
f" HNSW graph estimate: {self.hnsw_graph_gb():.2f} GB\n"
f" Payload storage (approx): {self.payload_storage_gb():.2f} GB\n"
f" Index footprint before overhead: {self.total_index_gb():.2f} GB\n"
f" Per-replica memory + headroom: {self.memory_per_replica_gb():.2f} GB\n"
f" Total cluster memory (approx): {self.total_cluster_memory_gb():.2f} GB\n"
f" ({self.replication_factor} replicas, each holding a full copy)\n"
f" Treat these as planning estimates, not deployment guarantees.\n"
f" Benchmark against your implementation before provisioning."
)
# Illustrative example: banking policy corpus
plan = VectorIndexCapacityPlan(
n_vectors=500_000,
dimension=1536,
hnsw_m=16,
replication_factor=2,
avg_payload_bytes=2048,
memory_headroom_factor=1.5
)
print(plan.report())
The Complete Retrieval Pipeline
When working through the The Complete Retrieval Pipeline stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
User Query
↓
Query Embedding (dense + sparse)
↓
Metadata / Authorisation Constraints
↓
Dense Vector Retrieval (filtered HNSW)
+
Sparse Retrieval (BM25)
↓
Weighted RRF Fusion
↓
Candidate Documents with Provenance Metadata
↓
[Part 7: Reranking and Context Assembly]
↓
LLM Generation
import logging
import re
from typing import List, Dict, Optional
from dataclasses import dataclass
from collections import defaultdict
logger = logging.getLogger(__name__)
@dataclass
class RetrievalConfig:
dense_candidate_pool: int = 50
sparse_candidate_pool: int = 50
rrf_k: int = 60
dense_weight: float = 0.6
sparse_weight: float = 0.4
final_top_k: int = 10
score_threshold: float = 0.2
class BankingPolicyRetriever:
"""
Production retrieval pipeline for a regulated banking policy corpus.
Combines dense vector search, BM25 sparse retrieval, access filtering,
and weighted RRF score fusion.
Retrieval ends at the fused candidate list. Reranking and context assembly
are handled in Part 7.
"""
def __init__(
self,
qdrant_client,
collection_name: str,
embedding_pipeline,
bm25_index,
chunk_store: Dict[str, Dict],
config: Optional[RetrievalConfig] = None
):
self.client = qdrant_client
self.collection = collection_name
self.embedder = embedding_pipeline
self.bm25 = bm25_index
self.chunk_store = chunk_store
self.config = config or RetrievalConfig()
def retrieve(
self,
query: str,
user_classification_ceiling: str = "INTERNAL",
user_jurisdiction: Optional[str] = None,
user_roles: Optional[List[str]] = None
) -> List[Dict]:
"""
Full hybrid retrieval with access control.
Returns top-k chunks with provenance metadata, access-filtered
for the requesting user's classification ceiling and jurisdiction.
"""
from qdrant_client.models import Filter, FieldCondition, MatchValue, MatchAny
query_embedding = self.embedder.embed_query(query)
classification_levels = {"PUBLIC": 0, "INTERNAL": 1, "CONFIDENTIAL": 2}
ceiling = classification_levels.get(user_classification_ceiling, 1)
permitted_classifications = [
k for k, v in classification_levels.items() if v <= ceiling
]
must_conditions = [
FieldCondition(
key="classification",
match=MatchAny(any=permitted_classifications)
),
FieldCondition(key="status", match=MatchValue(value="active"))
]
if user_jurisdiction:
must_conditions.append(
FieldCondition(
key="jurisdiction",
match=MatchValue(value=user_jurisdiction)
)
)
access_filter = Filter(must=must_conditions)
# Dense vector search with integrated access filtering
dense_hits = self.client.search(
collection_name=self.collection,
query_vector=("dense", query_embedding),
query_filter=access_filter,
limit=self.config.dense_candidate_pool,
score_threshold=self.config.score_threshold,
with_payload=True
)
dense_results = [(hit.id, hit.score) for hit in dense_hits]
# Sparse BM25 retrieval with post-retrieval access filtering
tokens = re.findall(r'\b\w+\b', query.lower())
bm25_scores = self.bm25.get_scores(tokens)
sparse_ranked = sorted(enumerate(bm25_scores), key=lambda x: x[1], reverse=True)
chunk_ids = list(self.chunk_store.keys())
sparse_results = []
for corpus_idx, score in sparse_ranked:
if score <= 0 or len(sparse_results) >= self.config.sparse_candidate_pool:
break
chunk_id = chunk_ids[corpus_idx]
chunk_meta = self.chunk_store.get(chunk_id, {})
if chunk_meta.get("classification") not in permitted_classifications:
continue
if user_jurisdiction and chunk_meta.get("jurisdiction") != user_jurisdiction:
continue
if chunk_meta.get("status") != "active":
continue
sparse_results.append((chunk_id, score))
# Weighted RRF fusion
rrf_scores: Dict[str, float] = defaultdict(float)
k = self.config.rrf_k
for rank, (chunk_id, _) in enumerate(dense_results, start=1):
rrf_scores[chunk_id] += self.config.dense_weight * (1.0 / (k + rank))
for rank, (chunk_id, _) in enumerate(sparse_results, start=1):
rrf_scores[chunk_id] += self.config.sparse_weight * (1.0 / (k + rank))
fused = sorted(rrf_scores.items(), key=lambda x: x[1], reverse=True)
top_chunk_ids = [cid for cid, _ in fused[:self.config.final_top_k]]
# Assemble results with provenance metadata
results = []
for chunk_id in top_chunk_ids:
chunk_data = self.chunk_store.get(chunk_id, {})
results.append({
"chunk_id": chunk_id,
"rrf_score": rrf_scores[chunk_id],
**chunk_data
})
logger.info(
f"Retrieval complete: query={query[:60]!r}, "
f"dense_candidates={len(dense_results)}, "
f"sparse_candidates={len(sparse_results)}, "
f"final_results={len(results)}"
)
return results
Observability: What to Measure in Production
When working through the Observability What to Measure stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Observability What to Measure stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
import time
import logging
from typing import Callable, TypeVar, Any
from functools import wraps
logger = logging.getLogger(__name__)
F = TypeVar("F", bound=Callable[..., Any])
def retrieval_instrumented(func: F) -> F:
"""
Decorator that adds structured latency logging and empty-result alerting
to retrieval functions. Wrap your primary retrieve() method in production.
"""
@wraps(func)
def wrapper(*args, **kwargs):
start = time.perf_counter()
result = None
error = None
try:
result = func(*args, **kwargs)
return result
except Exception as e:
error = str(e)
raise
finally:
elapsed_ms = (time.perf_counter() - start) * 1000
n_results = len(result) if result is not None else 0
log_payload = {
"function": func.__name__,
"latency_ms": round(elapsed_ms, 2),
"n_results": n_results,
"error": error
}
query = kwargs.get("query", args[1] if len(args) > 1 else None)
if query:
log_payload["query_prefix"] = str(query)[:80]
if error:
logger.error("retrieval_error", extra=log_payload)
elif n_results == 0:
logger.warning("retrieval_empty_result", extra=log_payload)
elif elapsed_ms > 500:
logger.warning("retrieval_high_latency", extra=log_payload)
else:
logger.info("retrieval_success", extra=log_payload)
return wrapper # type: ignore
Return to the EUR 12 Million Credit Proposal
The Return to the EUR stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Before You Move to Reranking: A Production Checklist
The Before You Move to stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
What Comes Next
The What Comes Next stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The What Comes Next stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Operational checklist
For the Operational checklist stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Track cost and latency beside quality. A slightly worse answer that costs 10x less may be the right production trade.
Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.
Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for fa68a70d815a: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.
For the hardening note 0 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 0/864: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 1 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 1/864: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 2 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 2/864: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 3 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 3/864: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 4 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 4/864: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 5 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 5/864: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 6 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 6/864: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 7 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 7/864: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 8 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 8/864: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 9 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 9/864: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 10 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 10/864: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 11 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 11/864: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 12 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 12/864: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 13 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 13/864: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 14 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 14/864: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.