This article is published in English.
Practical notes: The AI Series: The Hardest Part of RAG, Part 2 — Why Perfect
Operable walkthrough of Practical notes: The AI Series: The Hardest Part of RAG, Part 2 — Why Perfect: contracts, checks, and drop-in code slots for teams shipping this pattern.
This walkthrough rebuilds the path from raw materials to a working system for: The AI Series: The Hardest Part of RAG, Part 2 — Why Perfect Chunks Still Give You Bad Retrieval. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
First, What Is an Embedding, Actually?
When working through the First What Is an stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
"How do I reset my password?" → [0.81, 0.12, -0.44, 0.09]
"Steps to change your password" → [0.79, 0.15, -0.41, 0.11]
"How do I cancel my subscription?"→ [0.22, 0.68, 0.05, -0.39]
from numpy import dot
from numpy.linalg import norm
def cosine_similarity(a, b):
return dot(a, b) / (norm(a) * norm(b))
query = [0.80, 0.13, -0.42, 0.10] # "reset my password"
print(cosine_similarity(query, [0.81, 0.12, -0.44, 0.09])) # 0.998 - very close
print(cosine_similarity(query, [0.22, 0.68, 0.05, -0.39])) # 0.310 - far apart
Why Retrieval Is Genuinely Difficult
When working through the Why Retrieval Is Genuinely stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
What a Real Retrieval Pipeline Looks Like
When working through the What a Real Retrieval stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the What a Real Retrieval stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hybrid Search: Why Vector Search Alone Isn’t Enough
The Hybrid Search Why Vector stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
from collections import defaultdict
def reciprocal_rank_fusion(result_lists, k=60):
scores = defaultdict(float)
for result_list in result_lists:
for rank, doc_id in enumerate(result_list):
scores[doc_id] += 1 / (k + rank + 1)
return sorted(scores, key=scores.get, reverse=True)
vector_ids = [d.metadata["chunk_id"] for d in vector_store.similarity_search(query, k=20)]
bm25_ids = [d.metadata["chunk_id"] for d in bm25_index.search(query, k=20)]
final_ranking = reciprocal_rank_fusion([vector_ids, bm25_ids])
Reranking Is the Step Nobody Skips Twice
The Reranking Is the Step stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
from sentence_transformers import CrossEncoder
reranker = CrossEncoder("BAAI/bge-reranker-base")
def rerank(query, candidates, top_n=5):
pairs = [(query, c["text"]) for c in candidates]
scores = reranker.predict(pairs)
for c, s in zip(candidates, scores):
c["rerank_score"] = float(s)
return sorted(candidates, key=lambda c: c["rerank_score"], reverse=True)[:top_n]
Fixing the Question, Not Just the Index
The Fixing the Question Not stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Fixing the Question Not stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
def hyde_retrieve(query, llm, vector_store, k=5):
hypothetical = llm.invoke(f"Write a short passage answering: {query}")
return vector_store.similarity_search(hypothetical, k=k)
Measuring It Instead of Eyeballing It
For the Measuring It Instead of stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
def recall_at_k(retrieved_ids, relevant_ids, k):
top_k = set(retrieved_ids[:k])
return len(top_k & relevant_ids) / len(relevant_ids) if relevant_ids else 0.0
What you’d Tell Someone Starting This Today
For the What you d Tell stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
The Uncomfortable Truth About RAG
For the The Uncomfortable Truth About stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the The Uncomfortable Truth About stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Operational checklist
The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.
Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Score single-turn answers and multi-turn trajectories separately. Aggregate chat scores bury tool-loop failures.
Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.
Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for da3fc7392f6d: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.