This article is published in English.
Practical notes: Beyond Semantic Search: The Complete Guide to Advanced RAG
Operable walkthrough of Practical notes: Beyond Semantic Search: The Complete Guide to Advanced RAG: contracts, checks, and drop-in code slots for teams shipping this pattern.
This walkthrough rebuilds the path from raw materials to a working system for: Beyond Semantic Search: The Complete Guide to Advanced RAG with Milvus | the author. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
What Is RAG and Why Does It Exist?
When working through the What Is RAG and stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
User Question
│
▼
[Embed the question] → query vector
│
▼
[Search Vector DB] → top-K relevant document chunks
│
▼
[LLM prompt: "Given these passages, answer: {question}"]
│
▼
Accurate, Grounded Answer
The Role of a Vector Database
When working through the The Role of a stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Understanding Embeddings: Dense and Sparse
When working through the Understanding Embeddings Dense and stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Understanding Embeddings Dense and stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Dense Embeddings
The Dense Embeddings stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
"sick leave policy" → [0.12, -0.87, 0.34, 0.56, ...] (1024 numbers)
"medical absence entitlement" → [0.13, -0.85, 0.31, 0.54, ...] ← very close
"quarterly revenue target" → [0.91, 0.23, -0.67, 0.02, ...] ← far away
Sparse Embeddings (BM25)
The Sparse Embeddings BM25 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
"sick leave policy" → {word_index_for_"sick": 0.82, word_index_for_"leave": 0.91, ...}
Why You Need Both
The Why You Need Both stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Why You Need Both stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Project Setup and Dependencies
For the Project Setup and Dependencies stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
pip install --upgrade pymilvus
pip install "pymilvus[model]"
pip install sentence-transformers
pip install langchain-text-splitters
pip install langchain-openai
pip install langchain-community
pip install scipy
pip install nltk
import uuid
from tqdm import tqdm
from pymilvus import (
MilvusClient, DataType,
AnnSearchRequest, RRFRanker
)
from pymilvus.model.sparse import BM25EmbeddingFunction
from pymilvus.model.sparse.bm25.tokenizers import build_default_analyzer
from langchain_community.document_loaders import PyPDFLoader
from langchain_text_splitters import RecursiveCharacterTextSplitter
from langchain_openai import OpenAIEmbeddings, ChatOpenAI
import scipy.sparse as sp
import re, json
import nltk
nltk.download('stopwords')
Configuration
For the Configuration stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
PDF_PATH = "./data/sample_employee_handbook.pdf" # path of you document
COLLECTION_NAME = "rag_documents_hybrid"
MILVUS_DB_PATH = "./db/milvus_demo.db"
API_KEY = "sk-..."
EMBEDDING_MODEL = "text-embedding-3-large"
EMBEDDING_DIM = 1024
CHUNK_SIZE = 500
CHUNK_OVERLAP = 100
TOP_K = 5
Building the Indexing Pipeline
For the Building the Indexing Pipeline stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Building the Indexing Pipeline stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
PDF → Pages → Chunks → Dense Embeddings
→ Sparse Embeddings
→ Milvus Collection
Step 1 & 2: Initialize Models and Connect
When working through the Step 1 2 Initialize stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
# Dense embedding model here we'll be using OpenAI's embedding model
embedding_obj = OpenAIEmbeddings(
model=EMBEDDING_MODEL,
api_key=API_KEY,
dimensions=EMBEDDING_DIM
)
# Milvus Lite - single file, no server needed
client = MilvusClient(MILVUS_DB_PATH)
print("Models and DB connection ready.")
Step 3 & 4: Load and Chunk the Document
When working through the Step 3 4 Load stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
# Load PDF — one Document object per page
loader = PyPDFLoader(PDF_PATH)
documents = loader.load()
print(f"Loaded {len(documents)} pages.")
# Split into overlapping chunks
text_splitter = RecursiveCharacterTextSplitter(
chunk_size=CHUNK_SIZE,
chunk_overlap=CHUNK_OVERLAP,
separators=["\n\n", "\n", ".", " ", ""]
)
chunks = text_splitter.split_documents(documents)
print(f"Created {len(chunks)} chunks.")
Step 5: Generate Both Types of Embeddings
When working through the Step 5 Generate Both stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
texts = [doc.page_content for doc in chunks]
# Dense embeddings - one API call for the entire corpus
print("Generating dense embeddings...")
dense_embeddings = embedding_obj.embed_documents(texts)
print(f"Dense dimension: {len(dense_embeddings[0])}")
# Sparse embeddings - BM25 must be fit on YOUR corpus first
print("Fitting BM25 on corpus...")
analyzer = build_default_analyzer(language="en") # for this you will require nltk-stopwords
bm25_ef = BM25EmbeddingFunction(analyzer)
bm25_ef.fit(texts) # Builds vocabulary from your documents
sparse_embeddings = bm25_ef.encode_documents(texts)
print("Sparse embeddings generated.")
When working through the Step 5 Generate Both stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Step 6: Create the Collection with Schema and Indexes
The Step 6 Create the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
# Drop and recreate for a clean state
if COLLECTION_NAME in client.list_collections():
client.drop_collection(COLLECTION_NAME)
# Define schema
schema = client.create_schema()
schema.add_field("id", DataType.VARCHAR, is_primary=True, max_length=100)
schema.add_field("vector", DataType.FLOAT_VECTOR, dim=EMBEDDING_DIM)
schema.add_field("sparse_vector", DataType.SPARSE_FLOAT_VECTOR)
schema.add_field("text", DataType.VARCHAR, max_length=65535)
schema.add_field("page_number", DataType.INT64)
schema.add_field("source", DataType.VARCHAR, max_length=500)
schema.add_field("chunk_id", DataType.INT64)
# Create the collection
client.create_collection(collection_name=COLLECTION_NAME, schema=schema)
# Build indexes separately
index_params = client.prepare_index_params()
index_params.add_index(
field_name="vector",
index_type="FLAT", # Exact search - swap to HNSW for production
metric_type="COSINE"
)
index_params.add_index(
field_name="sparse_vector",
index_type="SPARSE_INVERTED_INDEX",
metric_type="IP" # Inner Product is the only valid metric for sparse
)
client.create_index(collection_name=COLLECTION_NAME, index_params=index_params)
print("Collection and indexes created.")
Step 7 & 8: Prepare Records and Insert
The Step 7 8 Prepare stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
def sparse_to_dict(s_emb) -> dict:
"""Convert a scipy sparse row into Milvus-compatible {index: value} dict."""
if sp.issparse(s_emb):
coo = s_emb.tocoo()
return {int(col): float(val) for col, val in zip(coo.col, coo.data)}
elif isinstance(s_emb, dict):
return s_emb
else:
return {int(i): float(v) for i, v in enumerate(s_emb) if v != 0.0}
# Build the records list
data = []
for idx, (chunk, d_emb) in enumerate(tqdm(zip(chunks, dense_embeddings), total=len(chunks))):
sparse_dict = sparse_to_dict(sparse_embeddings[idx])
if not sparse_dict:
print(f"Warning: empty sparse vector at chunk {idx}, skipping.")
continue
data.append({
"id": str(uuid.uuid4()),
"vector": d_emb,
"sparse_vector": sparse_dict,
"text": chunk.page_content,
"page_number": int(chunk.metadata.get("page", -1)),
"source": PDF_PATH,
"chunk_id": idx
})
# Insert into Milvus
res = client.insert(collection_name=COLLECTION_NAME, data=data)
print(f"Inserted {res['insert_count']} records.")
# Load into memory - required before any search operation
client.load_collection(COLLECTION_NAME)
print(f"Load state: {client.get_load_state(COLLECTION_NAME)}")
Basic RAG: Dense Vector Search
The Basic RAG Dense Vector stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Basic RAG Dense Vector stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
# ════════════════════════════════════════════════════════════
# Dense Vector Search
# ════════════════════════════════════════════════════════════
query = "What is the leave policy?"
# Step 1: Embed the query using the same model used at index time
query_dense_embedding = embedding_obj.embed_query(query)
# Step 2: Search
results = client.search(
collection_name=COLLECTION_NAME,
data=[query_dense_embedding],
anns_field="vector",
search_param={"metric_type": "COSINE"},
limit=TOP_K,
output_fields=["text", "page_number", "source"]
)
# Step 3: Display results
for idx, hit in enumerate(results[0], start=1):
entity = hit["entity"]
print(f"Rank {idx} | Cosine Score: {hit['distance']:.4f} | Page: {entity['page_number']}")
print(f" {entity['text'][:300]}\n")
Better RAG: Hybrid Search (Dense + Sparse)
For the Better RAG Hybrid Search stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
# ════════════════════════════════════════════════════════════
# Hybrid Search (Dense + Sparse)
# ════════════════════════════════════════════════════════════
query = "leave policy?"
# Dense query vector
query_dense = embedding_obj.embed_query(query)
# Sparse query vector - uses the same BM25 model fitted on the corpus
sparse_raw = bm25_ef.encode_queries([query])
sparse_dict = sparse_to_dict(sparse_raw[0])
print(f"Sparse query terms: {len(sparse_dict)}") # Should be > 0
# Build two separate ANN search requests
dense_req = AnnSearchRequest(
data=[query_dense],
anns_field="vector",
param={"metric_type": "COSINE"},
limit=TOP_K
)
sparse_req = AnnSearchRequest(
data=[sparse_dict],
anns_field="sparse_vector",
param={"metric_type": "IP"},
limit=TOP_K
)
# Execute hybrid search with RRF fusion
results = client.hybrid_search(
collection_name=COLLECTION_NAME,
reqs=[dense_req, sparse_req],
ranker=RRFRanker(k=60),
limit=TOP_K,
output_fields=["text", "page_number", "source"]
)
for idx, hit in enumerate(results[0], start=1):
entity = hit["entity"]
print(f"Rank {idx} | RRF Score: {hit['distance']:.4f} | Page: {entity['page_number']}")
print(f" {entity['text'][:300]}\n")
Advanced RAG — Four Retrieval Techniques
For the Advanced RAG Four Retrieval stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Metadata Filtering
For the Metadata Filtering stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Metadata Filtering stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
# ════════════════════════════════════════════════════════════
# METADATA FILTERING
# ════════════════════════════════════════════════════════════
def search_with_metadata_filter(
client, collection_name, embedding_obj, bm25_ef,
query: str,
page_range: tuple = None,
source_file: str = None,
top_k: int = 5
):
filter_parts = []
if page_range:
lo, hi = page_range
filter_parts.append(f"page_number >= {lo} && page_number <= {hi}")
if source_file:
filter_parts.append(f'source == "{source_file}"')
filter_expr = " && ".join(filter_parts) if filter_parts else None
print(f"\n[Metadata Filter] Query : '{query}'")
print(f"[Metadata Filter] Filter: {filter_expr or 'None (unfiltered)'}")
results = hybrid_search(
client, collection_name, embedding_obj, bm25_ef,
query_text=query,
top_k=top_k,
filters=filter_expr
)
return results
# ── Run ──────────────────────────────────────────────────────
meta_results = search_with_metadata_filter(
client, COLLECTION_NAME, embedding_obj, bm25_ef,
query = "What is the leave policy?",
page_range = (1, 30),
source_file= None,
top_k = 5
)
# ── Print Results ─────────────────────────────────────────────
print("\nMETADATA-FILTERED RESULTS")
print("=" * 55)
if not meta_results or not meta_results[0]:
print("No results returned.")
else:
for idx, hit in enumerate(meta_results[0], start=1):
entity = hit["entity"]
print(f"\nRank : {idx}")
print(f"Score : {hit['distance']:.4f}")
print(f"Page : {entity['page_number']}")
print(f"Text :\n{entity['text'][:400]}")
'page_number >= 1 && page_number <= 30' # page range
'source == "hr_policy.pdf"' # exact source
'category in ["leave", "performance"]' # in a list
'source like "hr%"' # prefix match
Query Rewriting
When working through the Query Rewriting stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
# ════════════════════════════════════════════════════════════
# QUERY REWRITING
# ════════════════════════════════════════════════════════════
import re, json
REWRITE_PROMPT = """You are an expert at reformulating search queries to improve document retrieval.
Given a user query, produce {n} alternative search queries that:
- Use formal, document-style language
- Include relevant keywords and synonyms
- Cover different angles of the same question
User query: {query}
Respond ONLY with a JSON array of strings. Example:
["rewritten query 1", "rewritten query 2", "rewritten query 3"]"""
def rewrite_query(query: str, n: int = 3) -> list[str]:
prompt = REWRITE_PROMPT.format(query=query, n=n)
response = llm.invoke(prompt)
raw = re.sub(r"^```json|^```|```quot;, "", response.content.strip(), flags=re.MULTILINE).strip()
try:
variants = json.loads(raw)
return [query] + variants # always keep the original
except json.JSONDecodeError:
print("Warning: Could not parse rewrites, using original query only.")
return [query]
def search_with_query_rewriting(
client, collection_name, embedding_obj, bm25_ef,
query: str,
n_rewrites: int = 3,
top_k: int = 5
):
variants = rewrite_query(query, n=n_rewrites)
print(f"\n[Query Rewriting] Original : '{query}'")
for i, v in enumerate(variants[1:], 1):
print(f"[Query Rewriting] Variant {i} : '{v}'")
seen_ids = {}
rank_scores = {}
for variant in variants:
results = hybrid_search(
client, collection_name, embedding_obj, bm25_ef,
query_text=variant,
top_k=top_k
)
if not results or not results[0]:
continue
for rank, hit in enumerate(results[0], start=1):
hit_id = hit["id"]
rank_scores[hit_id] = rank_scores.get(hit_id, 0) + 1.0 / (60 + rank)
if hit_id not in seen_ids:
seen_ids[hit_id] = hit
merged = sorted(seen_ids.values(), key=lambda h: rank_scores[h["id"]], reverse=True)[:top_k]
return [merged]
# ── Run ──────────────────────────────────────────────────────
rewrite_results = search_with_query_rewriting(
client, COLLECTION_NAME, embedding_obj, bm25_ef,
query = "What is the leave policy?",
n_rewrites = 3,
top_k = 5
)
# ── Print Results ─────────────────────────────────────────────
print("\nQUERY-REWRITTEN RESULTS")
print("=" * 55)
if not rewrite_results or not rewrite_results[0]:
print("No results returned.")
else:
for idx, hit in enumerate(rewrite_results[0], start=1):
entity = hit["entity"]
print(f"\nRank : {idx}")
print(f"Score : {hit['distance']:.4f}")
print(f"Page : {entity['page_number']}")
print(f"Text :\n{entity['text'][:400]}")
Input: "how many days off do I get?"
Output variants:
1. "annual leave entitlement number of days employee handbook"
2. "vacation days accrual policy full-time employee"
3. "paid time off PTO allowance per calendar year"
HyDE — Hypothetical Document Embeddings
When working through the HyDE Hypothetical Document Embeddings stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
# ════════════════════════════════════════════════════════════
# HyDE (Hypothetical Document Embeddings)
# ════════════════════════════════════════════════════════════
HYDE_PROMPT = """You are a corporate policy document writer.
Write a 2-3 paragraph excerpt from an official HR policy or company document
that would DIRECTLY ANSWER the following question.
Write in formal document style. Do not mention the question itself.
Question: {query}
Document excerpt:"""
def generate_hypothetical_document(query: str) -> str:
response = llm.invoke(HYDE_PROMPT.format(query=query))
return response.content.strip()
def search_with_hyde(
client, collection_name, embedding_obj, bm25_ef,
query: str,
top_k: int = 5
):
hypothetical_doc = generate_hypothetical_document(query)
print(f"\n[HyDE] Query : '{query}'")
print(f"[HyDE] Hypothetical doc :\n {hypothetical_doc[:300]}...\n")
# Search using the hypothetical document's embedding
hyde_results = hybrid_search(
client, collection_name, embedding_obj, bm25_ef,
query_text=hypothetical_doc, # embed the answer, not the question
top_k=top_k
)
# Also search with the original query and merge both via RRF
original_results = hybrid_search(
client, collection_name, embedding_obj, bm25_ef,
query_text=query,
top_k=top_k
)
seen_ids = {}
rank_scores = {}
for result_set in [hyde_results, original_results]:
if not result_set or not result_set[0]:
continue
for rank, hit in enumerate(result_set[0], start=1):
hit_id = hit["id"]
rank_scores[hit_id] = rank_scores.get(hit_id, 0) + 1.0 / (60 + rank)
if hit_id not in seen_ids:
seen_ids[hit_id] = hit
merged = sorted(seen_ids.values(), key=lambda h: rank_scores[h["id"]], reverse=True)[:top_k]
return [merged]
# ── Run ──────────────────────────────────────────────────────
hyde_results = search_with_hyde(
client, COLLECTION_NAME, embedding_obj, bm25_ef,
query = "What is the leave policy?",
top_k = 5
)
# ── Print Results ─────────────────────────────────────────────
print("\nHyDE RESULTS")
print("=" * 55)
if not hyde_results or not hyde_results[0]:
print("No results returned.")
else:
for idx, hit in enumerate(hyde_results[0], start=1):
entity = hit["entity"]
print(f"\nRank : {idx}")
print(f"Score : {hit['distance']:.4f}")
print(f"Page : {entity['page_number']}")
print(f"Text :\n{entity['text'][:400]}")
Query Decomposition
When working through the Query Decomposition stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Query Decomposition stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
# ════════════════════════════════════════════════════════════
# QUERY DECOMPOSITION
# ════════════════════════════════════════════════════════════
DECOMPOSE_PROMPT = """You are an expert at breaking down complex questions for document retrieval.
Decompose the following question into 2-4 simple, self-contained sub-questions.
Each sub-question should target a single distinct piece of information.
Complex question: {query}
Respond ONLY with a JSON array of strings. Example:
["sub-question 1", "sub-question 2", "sub-question 3"]"""
def decompose_query(query: str) -> list[str]:
response = llm.invoke(DECOMPOSE_PROMPT.format(query=query))
raw = re.sub(r"^```json|^```|```quot;, "", response.content.strip(), flags=re.MULTILINE).strip()
try:
return json.loads(raw)
except json.JSONDecodeError:
print("Warning: Could not parse decomposition, using original query.")
return [query]
def search_with_decomposition(
client, collection_name, embedding_obj, bm25_ef,
query: str,
top_k: int = 5
):
sub_questions = decompose_query(query)
print(f"\n[Decomposition] Original query : '{query}'")
for i, sq in enumerate(sub_questions, 1):
print(f"[Decomposition] Sub-question {i} : '{sq}'")
per_subquery_results = {}
seen_ids = {}
rank_scores = {}
for sq in sub_questions:
results = hybrid_search(
client, collection_name, embedding_obj, bm25_ef,
query_text=sq,
top_k=top_k
)
per_subquery_results[sq] = results
if not results or not results[0]:
continue
for rank, hit in enumerate(results[0], start=1):
hit_id = hit["id"]
rank_scores[hit_id] = rank_scores.get(hit_id, 0) + 1.0 / (60 + rank)
if hit_id not in seen_ids:
seen_ids[hit_id] = hit
merged = sorted(seen_ids.values(), key=lambda h: rank_scores[h["id"]], reverse=True)[:top_k]
# Per sub-question breakdown
print("\n── Per Sub-question Results ──")
for sq, res in per_subquery_results.items():
print(f"\n SUB-QUERY: '{sq[:60]}'")
if res and res[0]:
for i, hit in enumerate(res[0], start=1):
print(f" {i}. Page {hit['entity']['page_number']} | Score {hit['distance']:.4f} | {hit['entity']['text'][:150]}")
return {"per_subquery": per_subquery_results, "merged": [merged]}
# ── Run ──────────────────────────────────────────────────────
decomp_results = search_with_decomposition(
client, COLLECTION_NAME, embedding_obj, bm25_ef,
query = "What is the leave policy and how does it affect salary deductions?",
top_k = 5
)
# ── Print Merged Results ──────────────────────────────────────
print("\nDECOMPOSED — MERGED FINAL RESULTS")
print("=" * 55)
merged_hits = decomp_results["merged"]
if not merged_hits or not merged_hits[0]:
print("No results returned.")
else:
for idx, hit in enumerate(merged_hits[0], start=1):
entity = hit["entity"]
print(f"\nRank : {idx}")
print(f"Score : {hit['distance']:.4f}")
print(f"Page : {entity['page_number']}")
print(f"Text :\n{entity['text'][:400]}")
Input: "What is the leave policy and how does performance review affect salary?"
Sub-questions:
1. "What is the annual leave policy?"
2. "How many sick days are employees entitled to?"
3. "How does performance review affect salary?"
4. "What is the performance review schedule?"
Cross-Encoder Reranking
The Cross-Encoder Reranking stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
# ============================================================
# RERANKING WITH CROSS-ENCODER
# ============================================================
from sentence_transformers import CrossEncoder
# Huggingface: cross-encoder/ms-marco-MiniLM-L12-v2
cross_encoder = CrossEncoder("cross-encoder/ms-marco-MiniLM-L12-v2")
query = "What is the leave policy?"
RETRIEVAL_K = 20 # fetch more than you need
FINAL_K = 5 # rerank down to this
# Step 1: Broad retrieval - fetch 20 candidates
query_dense = embedding_obj.embed_query(query)
results = client.search(
collection_name=COLLECTION_NAME,
data=[query_dense],
anns_field="vector",
search_param={"metric_type": "COSINE"},
limit=RETRIEVAL_K,
output_fields=["text", "page_number", "source"]
)
hits = results[0]
print(f"Retrieved {len(hits)} candidates for reranking.")
# Step 2: Score each (query, chunk) pair with the cross-encoder
pairs = [[query, hit["entity"]["text"]] for hit in hits]
rerank_scores = cross_encoder.predict(pairs)
# Step 3: Sort by cross-encoder score
for hit, score in zip(hits, rerank_scores):
hit["rerank_score"] = float(score)
reranked = sorted(hits, key=lambda x: x["rerank_score"], reverse=True)[:FINAL_K]
# Step 4: Display
for idx, hit in enumerate(reranked, start=1):
entity = hit["entity"]
print(f"Rank {idx} | Rerank: {hit['rerank_score']:.4f} | Vector: {hit['distance']:.4f}")
print(f" Page {entity['page_number']}: {entity['text'][:300]}\n")
How the Techniques Compare
The How the Techniques Compare stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Conclusion
The Conclusion stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Conclusion stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Operational checklist
The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.
Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for c9664ffe2213: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.
Deployment note 1 (c9664ffe2213): pin images, set request budgets, and verify tenant isolation on a canary before wider rollouts.
Deployment note 2 (c9664ffe2213): pin images, set request budgets, and verify tenant isolation on a canary before wider rollouts.
Deployment note 3 (c9664ffe2213): pin images, set request budgets, and verify tenant isolation on a canary before wider rollouts.
Deployment note 4 (c9664ffe2213): pin images, set request budgets, and verify tenant isolation on a canary before wider rollouts.
Deployment note 5 (c9664ffe2213): pin images, set request budgets, and verify tenant isolation on a canary before wider rollouts.
Deployment note 6 (c9664ffe2213): pin images, set request budgets, and verify tenant isolation on a canary before wider rollouts.
Deployment note 7 (c9664ffe2213): pin images, set request budgets, and verify tenant isolation on a canary before wider rollouts.
Deployment note 8 (c9664ffe2213): pin images, set request budgets, and verify tenant isolation on a canary before wider rollouts.
Deployment note 9 (c9664ffe2213): pin images, set request budgets, and verify tenant isolation on a canary before wider rollouts.
Deployment note 10 (c9664ffe2213): pin images, set request budgets, and verify tenant isolation on a canary before wider rollouts.