This article is published in English.
Practical notes: I Threw Out My Vector Database. RAG Got Way Better With
Operable walkthrough of Practical notes: I Threw Out My Vector Database. RAG Got Way Better With: contracts, checks, and drop-in code slots for teams shipping this pattern.
The following notes reconstruct a practical path around “I Threw Out My Vector Database. RAG Got Way Better With PageIndex”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
The Core Lie of Vector RAG
The The Core Lie of stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
PageIndex: RAG Without the Vector Database
The PageIndex RAG Without the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
How PageIndex Works: The Two-Step Process
The How PageIndex Works The stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The How PageIndex Works The stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Getting Started: Running PageIndex Locally
For the Getting Started Running PageIndex stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
git clone https://github.com/VectifyAI/PageIndex.git
cd PageIndex
pip3 install --upgrade -r requirements.txt
CHATGPT_API_KEY=your_openai_api_key_here
python3 run_pageindex.py --pdf_path /path/to/annual_report.pdf
python3 run_pageindex.py --md_path /path/to/technical_spec.md
What the Index Actually Looks Like
For the What the Index Actually stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
{
"document": "Apple Inc. Annual Report 2023",
"index": {
"title": "Apple Inc. Annual Report 2023",
"summary": "Comprehensive financial and operational report covering revenue, product segments, risks, and strategic outlook",
"children": [
{
"title": "Business Overview",
"summary": "Company description, product lines, and market position",
"pages": [1, 8],
"children": [...]
},
{
"title": "Financial Results",
"summary": "Revenue, operating income, EPS, and segment performance for fiscal 2023",
"pages": [45, 72],
"children": [
{
"title": "Revenue by Product Category",
"summary": "iPhone, Mac, iPad, Wearables, and Services revenue breakdown",
"pages": [46, 52]
},
{
"title": "Geographic Revenue Distribution",
"summary": "Americas, Europe, Greater China, Japan, Rest of Asia Pacific",
"pages": [53, 58]
}
]
},
{
"title": "Risk Factors",
"summary": "Operational, market, regulatory, and competitive risks",
"pages": [89, 110]
}
]
}
}
Querying the Index: Under the Hood
For the Querying the Index Under stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Querying the Index Under stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
import json
from openai import OpenAI
client = OpenAI()
def navigate_index(query: str, index_node: dict, depth: int = 0) -> list[dict]:
"""
Recursively navigate the document index using LLM reasoning.
Returns list of relevant leaf nodes with page references.
"""
children = index_node.get("children", [])
if not children:
# Leaf node: return this section as relevant
return [index_node]
# Ask the LLM which branches are relevant to the query
children_summary = "\n".join([
f"[{i}] {child['title']}: {child['summary']}"
for i, child in enumerate(children)
])
response = client.chat.completions.create(
model="gpt-4o",
messages=[
{
"role": "system",
"content": (
"You are navigating a document index to find sections relevant "
"to a query. Select the index numbers of sections that are likely "
"to contain the answer. Return a JSON array of selected indices."
)
},
{
"role": "user",
"content": (
f"Query: {query}\n\n"
f"Available sections:\n{children_summary}\n\n"
f"Which sections should I look into? Return JSON array of indices only."
)
}
],
temperature=0,
response_format={"type": "json_object"}
)
selected = json.loads(response.choices[0].message.content).get("indices", [])
relevant_nodes = []
for idx in selected:
if idx # Recurse into selected branches
relevant_nodes.extend(
navigate_index(query, children[idx], depth + 1)
)
return relevant_nodes
def answer_with_pageindex(query: str, index: dict, document_pages: dict) -> str:
"""
Full PageIndex retrieval and answer generation.
"""
# Navigate the index to find relevant sections
relevant_nodes = navigate_index(query, index)
# Retrieve full text from identified pages
context_parts = []
citations = []
for node in relevant_nodes:
pages = node.get("pages", [])
if pages:
page_start, page_end = pages[0], pages[1]
for page_num in range(page_start, page_end + 1):
if page_num in document_pages:
context_parts.append(document_pages[page_num])
citations.append(f"p.{page_num}")
context = "\n\n".join(context_parts)
# Generate answer with full, unchunked context
answer_response = client.chat.completions.create(
model="gpt-4o",
messages=[
{
"role": "system",
"content": (
"Answer the question based on the provided document sections. "
"Be precise. If the answer involves numbers or dates, quote them exactly."
)
},
{
"role": "user",
"content": f"Document sections:\n{context}\n\nQuestion: {query}"
}
],
temperature=0
)
answer = answer_response.choices[0].message.content
citation_str = ", ".join(set(citations))
return f"{answer}\n\n**Source:** {citation_str}"
The FinanceBench Result That Made Me Pay Attention
When working through the The FinanceBench Result That stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
When to Use PageIndex vs. Traditional RAG
When working through the When to Use PageIndex stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Using the PageIndex Cloud API
When working through the Using the PageIndex Cloud stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Using the PageIndex Cloud stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
import requests
PAGEINDEX_API_KEY = "your_api_key"
BASE_URL = "https://api.pageindex.ai/v1"
def upload_document(file_path: str) -> str:
"""Upload a document and get back a document_id."""
with open(file_path, "rb") as f:
response = requests.post(
f"{BASE_URL}/documents",
headers={"Authorization": f"Bearer {PAGEINDEX_API_KEY}"},
files={"file": f}
)
return response.json()["document_id"]
def query_document(document_id: str, question: str) -> dict:
"""Query an indexed document and get a cited answer."""
response = requests.post(
f"{BASE_URL}/query",
headers={
"Authorization": f"Bearer {PAGEINDEX_API_KEY}",
"Content-Type": "application/json"
},
json={
"document_id": document_id,
"question": question
}
)
return response.json()
# Example usage
doc_id = upload_document("q3_earnings_report.pdf")
result = query_document(doc_id, "What was total revenue in Q3?")
print(result["answer"])
print(f"Sources: {result['citations']}")
The Deeper Shift This Represents
The The Deeper Shift This stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
What This Means If You Are Building Document AI Today
The What This Means If stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Let’s Keep Learning Together
The Let s Keep Learning stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Let s Keep Learning stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Resources
For the Resources stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Operational checklist
For the Operational checklist stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.
Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 888b75aac33b: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.