This article is published in English.
Practical notes: I Stopped Using Vector Databases for RAG : PageIndex
Operable walkthrough of Practical notes: I Stopped Using Vector Databases for RAG : PageIndex: contracts, checks, and drop-in code slots for teams shipping this pattern.
The following notes reconstruct a practical path around “I Stopped Using Vector Databases for RAG : PageIndex Vectorless RAG”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
What Even Is PageIndex?
The What Even Is PageIndex stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
The Problem with Vector RAG (That Nobody Talks About Enough)
The The Problem with Vector stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Enter PageIndex- Retrieval by Reasoning
The Enter PageIndex- Retrieval by stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices. The Enter PageIndex- Retrieval by stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
How It Actually Works
For the How It Actually Works stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Annual Report 2023
├── Business Overview
│ ├── Products and Services
│ └── Market Position
├── Risk Factors
│ ├── Financial Risks
│ └── Operational Risks
├── Financial Statements
│ ├── Balance Sheet
│ │ ├── Assets
│ │ └── Liabilities
│ └── Income Statement
└── Notes to Financial Statements
├── Note 1: Accounting Policies
└── Note 12: Long-term Debt
Architecture of PageIndex
For the Architecture of PageIndex stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Let’s Code
For the Let s Code stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Step 1: Parse the Document
For the Step 1 Parse the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
import fitz # pip install pymupdf
def parse_pdf(pdf_path: str) -> list[dict]:
doc = fitz.open(pdf_path)
pages = []
for i, page in enumerate(doc):
text = page.get_text().strip()
if text:
pages.append({"page_num": i + 1, "text": text})
doc.close()
return pages
def group_pages_into_sections(pages, per_section=3):
sections = []
for i in range(0, len(pages), per_section):
batch = pages[i : i + per_section]
section_id = f"S{str(i // per_section + 1).zfill(3)}"
combined_text = "\n\n".join(p["text"] for p in batch)
sections.append({
"section_id": section_id,
"start_page": batch[0]["page_num"],
"end_page": batch[-1]["page_num"],
"text": combined_text,
})
return sections
Step 2: Build the Tree Index (LLM-Powered)
For the Step 2 Build the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
from google import genai
from google.genai import types
client = genai.Client(api_key="YOUR_GEMINI_API_KEY")
def index_section(section: dict) -> dict:
preview = section["text"][:1500]
prompt = f"""Read this section from a document and summarize it.
Section pages: {section['start_page']} to {section['end_page']}
Text:
{preview}
Respond with ONLY valid JSON:
{{
"title": "short descriptive title (5-8 words)",
"summary": "2-3 sentence summary of what this section covers",
"key_topics": ["topic1", "topic2", "topic3"]
}}"""
response = client.models.generate_content(
model="gemini-2.0-flash",
contents=prompt,
config=types.GenerateContentConfig(temperature=0.0),
)
parsed = json.loads(response.text.strip())
return {
"node_id": section["section_id"],
"title": parsed["title"],
"pages": f"{section['start_page']}-{section['end_page']}",
"summary": parsed["summary"],
"key_topics": parsed["key_topics"],
}
def build_tree_index(sections):
nodes = [index_section(s) for s in sections]
return {
"title": "Your Document Title",
"total_sections": len(nodes),
"nodes": nodes,
}
# Save for reuse
with open("tree.json", "w") as f:
json.dump(tree, f, indent=2)
Step 3: Tree Search (Reasoning, Not Similarity)
For the Step 3 Tree Search stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
def retrieve_sections(tree: dict, query: str) -> dict:
# Build a compact text representation of the tree
tree_text = f"Document: {tree['title']}\n\n"
for node in tree["nodes"]:
tree_text += f"[{node['node_id']}] Pages {node['pages']} | {node['title']}\n"
tree_text += f" Summary: {node['summary']}\n"
tree_text += f" Topics: {', '.join(node['key_topics'])}\n\n"
prompt = f"""You are a document retrieval expert.
Given this document tree, identify which sections most likely answer the question.
Think step by step about where a human expert would look.
{tree_text}
QUESTION: {query}
Respond with ONLY valid JSON:
{{
"reasoning": "your step-by-step reasoning about where to look",
"selected_ids": ["S001", "S004"],
"confidence": "high/medium/low"
}}"""
response = client.models.generate_content(
model="gemini-2.0-flash",
contents=prompt,
config=types.GenerateContentConfig(temperature=0.0),
)
return json.loads(response.text.strip())
Step 4: Content Retrieval
For the Step 4 Content Retrieval stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
def retrieve_content(selected_ids: list, sections: list) -> str:
section_map = {s["section_id"]: s for s in sections}
context_parts = []
for sid in selected_ids:
if sid in section_map:
sec = section_map[sid]
context_parts.append(
f"--- Pages {sec['start_page']}-{sec['end_page']} ---\n"
+ sec["text"][:3000]
)
return "\n\n".join(context_parts)
Step 5: Answer Generation
For the Step 5 Answer Generation stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Step 5 Answer Generation stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
def generate_answer(query: str, context: str) -> str:
prompt = f"""Answer the question using only the provided context.
Be specific. Include exact numbers, technical terms, and cite page numbers.CONTEXT:
{context}
QUESTION: {query}
ANSWER:"""
response = client.models.generate_content(
model="gemini-2.0-flash",
contents=prompt,
config=types.GenerateContentConfig(temperature=0.1),
)
return response.text.strip()
PageIndex vs Traditional Vector RAG
When working through the PageIndex vs Traditional Vector stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
When Should You Use PageIndex?
When working through the When Should You Use stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Final Thought
When working through the Final Thought stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Final Thought stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Getting Started
The Getting Started stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Operational checklist
For the Operational checklist stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.
Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for e54dedbe364e: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.