This article is published in English.
Practical notes: I Deleted My Vector Database and My RAG System Got Better
Operable walkthrough of Practical notes: I Deleted My Vector Database and My RAG System Got Better: contracts, checks, and drop-in code slots for teams shipping this pattern.
Use this as an operator-facing rebuild of the ideas in “I Deleted My Vector Database and My RAG System Got Better”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
What Is RAG, and Why Should You Care?
For the What Is RAG and stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
The Obvious First Idea (And Why It Fails)
For the The Obvious First Idea stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
from openai import OpenAI
import PyPDF2
client = OpenAI(api_key="your-api-key")
# Extract all text from a PDF
def extract_pdf_text(pdf_path):
reader = PyPDF2.PdfReader(pdf_path)
full_text = ""
for page in reader.pages:
full_text += page.extract_text() + "\n"
return full_text
document_text = extract_pdf_text("annual_report.pdf")
user_question = "What was the total revenue in 2024?"
response = client.chat.completions.create(
model="gpt-4.1",
messages=[
{"role": "system", "content": "Answer based on the provided document."},
{"role": "user", "content": f"Document:\n{document_text}\n\nQuestion: {user_question}"}
]
)
print(response.choices[0].message.content)
Traditional Vector RAG: The Current Industry Standard
For the Traditional Vector RAG The stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Traditional Vector RAG The stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
from openai import OpenAI
import chromadb
import PyPDF2
client = OpenAI(api_key="your-api-key")
# Step 1: Extract and chunk the document
def extract_and_chunk(pdf_path, chunk_size=500):
reader = PyPDF2.PdfReader(pdf_path)
full_text = ""
for page in reader.pages:
full_text += page.extract_text() + "\n"
words = full_text.split()
chunks = []
for i in range(0, len(words), chunk_size):
chunk = " ".join(words[i:i + chunk_size])
chunks.append(chunk)
return chunks
chunks = extract_and_chunk("annual_report.pdf")
# Step 2 and 3: Embed and store in ChromaDB
chroma_client = chromadb.Client()
collection = chroma_client.create_collection("my_documents")
collection.add(
documents=chunks,
ids=[f"chunk_{i}" for i in range(len(chunks))]
)
# Step 4: Search for relevant chunks
user_question = "What was the total revenue in 2024?"
results = collection.query(
query_texts=[user_question],
n_results=5
)
relevant_chunks = "\n\n".join(results["documents"][0])
# Step 5: Generate answer with focused context
response = client.chat.completions.create(
model="gpt-4.1",
messages=[
{"role": "system", "content": "Answer based only on the provided context."},
{"role": "user", "content": f"Context:\n{relevant_chunks}\n\nQuestion: {user_question}"}
]
)
print(response.choices[0].message.content)
Where Vector RAG Breaks Down
When working through the Where Vector RAG Breaks stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Problem 1: Chunking Destroys Context
When working through the Problem 1 Chunking Destroys stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
paragraph = """The company's Q3 operating income was $4.2 billion,
representing a 12% increase over the prior year period. This growth
was primarily driven by the expansion of cloud services, which
contributed $2.8 billion in recurring revenue as detailed in the
segment breakdown in Appendix C."""
# Simulating a chunk boundary at word 20
words = paragraph.split()
chunk_1 = " ".join(words[:20])
chunk_2 = " ".join(words[20:])
print("Chunk 1:", chunk_1)
print("---")
print("Chunk 2:", chunk_2)
Problem 2: Cross-References Get Lost
When working through the Problem 2 Cross-References Get stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Problem 2 Cross-References Get stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
# Simulating a cross-reference failure
documents = {
"page_12": "The company's debt-to-equity ratio improved significantly in 2024. For a full breakdown of long-term obligations, see Appendix G on page 87.",
"page_45": "Marketing expenses increased by 15% due to new campaigns.",
"page_87": "Appendix G: Long-term debt stands at $12.4B. Senior notes: $8.1B. Credit facility: $4.3B. Maturity schedule: 2026-2034."
}
# Vector RAG would likely return page_12 for a debt question
# but miss page_87 where the actual numbers live
# because "debt-to-equity ratio" is more similar to the query
# than "Senior notes" and "Credit facility"
Problem 3: User Wording Matters Too Much
The Problem 3 User Wording stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Enter Vectorless RAG: Teaching AI to Read Like a Human
The Enter Vectorless RAG Teaching stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Step 1: Build a Hierarchical Tree Index
The Step 1 Build a stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Step 1 Build a stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Root: Annual Report 2024
|-- Executive Summary (pages 1-3)
| Summary: "Overview of company performance, key metrics..."
|-- Financial Statements (pages 15-45)
| |-- Income Statement (pages 15-20)
| |-- Balance Sheet (pages 21-30)
| +-- Cash Flow Statement (pages 31-45)
|-- Risk Factors (pages 46-60)
+-- Appendices (pages 80-120)
|-- Appendix A: Segment Data (pages 80-95)
+-- Appendix G: Detailed Tables (pages 96-120)
Step 2: Reasoning-Based Tree Search
For the Step 2 Reasoning-Based Tree stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
import json
from openai import OpenAI
import PyPDF2
client = OpenAI(api_key="your-api-key")
# Extract text page by page
def extract_pages(pdf_path):
reader = PyPDF2.PdfReader(pdf_path)
pages = {}
for i, page in enumerate(reader.pages):
pages[i + 1] = page.extract_text()
return pages
pages = extract_pages("annual_report.pdf")
all_text = "\n".join([f"--- Page {k} ---\n{v}" for k, v in pages.items()])
# Step 1: Build the tree index using AI reasoning
tree_prompt = f"""You are a document analyst. Read this document and create
a hierarchical table of contents as a JSON tree. Each node should have:
- "title": section name
- "summary": 2-3 sentence description of what this section covers
- "pages": [start_page, end_page]
- "children": array of child nodes (or empty array)
Document:
{all_text[:80000]}
Return ONLY valid JSON. No other text."""
tree_response = client.chat.completions.create(
model="gpt-4.1",
messages=[{"role": "user", "content": tree_prompt}],
response_format={"type": "json_object"}
)
document_tree = json.loads(tree_response.choices[0].message.content)
print("Document tree built successfully!")
print(json.dumps(document_tree, indent=2)[:500])
# Step 2: Reasoning-based search over the tree
user_question = "What was the year-over-year change in operating margin?"
search_prompt = f"""You are a retrieval expert. Given a user question and
a document's table of contents tree, identify which sections are most
likely to contain the answer.
Think step by step:
1. What kind of information does the question ask for?
2. Which sections would a human expert check first?
3. Are there sections that might cross-reference each other?
Document tree:
{json.dumps(document_tree, indent=2)}
Question: {user_question}
Return JSON with:
- "reasoning": your step-by-step thought process
- "relevant_pages": list of page numbers to retrieve"""
search_response = client.chat.completions.create(
model="gpt-4.1",
messages=[{"role": "user", "content": search_prompt}],
response_format={"type": "json_object"}
)
search_result = json.loads(search_response.choices[0].message.content)
print("Reasoning:", search_result["reasoning"])
# Step 3: Fetch only relevant pages and generate answer
relevant_text = ""
for page_num in search_result["relevant_pages"]:
if page_num in pages:
relevant_text += f"\n--- Page {page_num} ---\n{pages[page_num]}"
answer_response = client.chat.completions.create(
model="gpt-4.1",
messages=[
{"role": "system", "content": "Answer precisely based on the document context provided."},
{"role": "user", "content": f"Context:{relevant_text}\n\nQuestion: {user_question}"}
]
)
print("\nAnswer:", answer_response.choices[0].message.content)
Going Production-Ready with PageIndex SDK
For the Going Production-Ready with PageIndex stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
from pageindex import PageIndexClient
import time
# Initialize the client
pi_client = PageIndexClient(api_key="your-pageindex-api-key")
# Upload and process your document
result = pi_client.submit_document("./annual_report.pdf")
doc_id = result["doc_id"]
# Wait for processing to complete
while True:
status = pi_client.get_document(doc_id)["status"]
if status == "completed":
print("Document processed!")
break
time.sleep(5)
# Inspect the generated tree structure
tree_result = pi_client.get_tree(doc_id)
if tree_result.get("status") == "completed":
tree = tree_result["result"]
for node in tree:
print(f"[{node['node_id']}] {node['title']} - Page {node['page_index']}")
# Ask questions using the Chat API
response = pi_client.chat_completions(
messages=[{"role": "user", "content": "What was total revenue in FY2024 vs FY2023?"}],
doc_id=doc_id
)
print(response["choices"][0]["message"]["content"])
{
"title": "Financial Stability",
"node_id": "0006",
"page_index": 21,
"text": "The Federal Reserve maintains financial stability...",
"nodes": [
{
"title": "Monitoring Financial Vulnerabilities",
"node_id": "0007",
"page_index": 22,
"text": "The Federal Reserve's monitoring focuses on..."
},
{
"title": "Domestic and International Cooperation",
"node_id": "0008",
"page_index": 28,
"text": "In 2023, the Federal Reserve collaborated..."
}
]
}
response = pi_client.chat_completions(
messages=[{"role": "user", "content": "Compare the results across these two reports."}],
doc_id=["pi-abc123def456", "pi-abc123ghi789"]
)
The Numbers Tell the Story
For the The Numbers Tell the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the The Numbers Tell the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
When to Use Which Approach
When working through the When to Use Which stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
And What’s Next?
When working through the And What s Next stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
ReCap:
When working through the ReCap stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the ReCap stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Operational checklist
When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.
Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 61253a21aab9: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.
When working through the hardening note 0 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 0/802: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 1 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 1/802: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 2 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 2/802: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 3 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 3/802: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 4 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 4/802: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 5 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 5/802: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.