This article is published in English.
Practical notes: PageIndex: The RAG Framework That Threw Out Vector Databases
Operable walkthrough of Practical notes: PageIndex: The RAG Framework That Threw Out Vector Databases: contracts, checks, and drop-in code slots for teams shipping this pattern.
Use this as an operator-facing rebuild of the ideas in “PageIndex: The RAG Framework That Threw Out Vector Databases and Still Hit 98.7% Accuracy”: clear stages, ordered code slots, and recovery notes that survive a handoff.
How VectifyAI’s reasoning-based retrieval is quietly dismantling the most deep-rooted assumption in production RAG
The How VectifyAI s reasoning-based stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
The Problem We Keep Papering Over
The The Problem We Keep stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
What PageIndex Actually Is
The What PageIndex Actually Is stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Step 1: Build a Hierarchical Tree Index
The Step 1 Build a stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
{
"node_id": "0006",
"title": "Financial Stability",
"start_index": 21,
"end_index": 22,
"summary": "Covers the Federal Reserve's financial stability oversight...",
"sub_nodes": [
{
"node_id": "0007",
"title": "Monitoring Financial Vulnerabilities",
"start_index": 22,
"end_index": 28,
"summary": "Describes the Fed's vulnerability monitoring framework..."
},
{
"node_id": "0008",
"title": "Domestic and International Cooperation",
"start_index": 28,
"end_index": 31,
"summary": "Federal Reserve collaboration with international bodies..."
}
]
}
Step 2: Reasoning-Based Tree Search
The Step 2 Reasoning-Based Tree stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices. The Step 2 Reasoning-Based Tree stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Why This Actually Works: The Appendix G Example
For the Why This Actually Works stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Python Implementation: End-to-End Vectorless RAG
For the Python Implementation End-to-End Vectorless stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.
Installation
For the Installation stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
pip install pageindex openai
Setup
For the Setup stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
import os
import json
import asyncio
from pageindex import PageIndexClient
from openai import AsyncOpenAI
# Grab an API key from https://dash.pageindex.ai/api-keys
PAGEINDEX_API_KEY = os.environ["PAGEINDEX_API_KEY"]
OPENAI_API_KEY = os.environ["OPENAI_API_KEY"]pi_client = PageIndexClient(api_key=PAGEINDEX_API_KEY)
openai_client = AsyncOpenAI(api_key=OPENAI_API_KEY)
Ingest a Document and Build the Tree
For the Ingest a Document and stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
import pageindex.utils as utils
# Upload a PDF; PageIndex handles the tree generation
doc = pi_client.upload("annual_report_2024.pdf")
doc_id = doc["doc_id"]# Tree generation takes a bit, so we poll
while not pi_client.is_retrieval_ready(doc_id):
print("Still indexing...")
import time; time.sleep(5)# Grab the tree and take a look
tree = pi_client.get_tree(doc_id, node_summary=True)["result"]
print("Document Tree:")
utils.print_tree(tree)
The Core: LLM-Driven Tree Search
For the The Core LLM-Driven Tree stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
async def find_relevant_nodes(tree: dict, query: str) -> list:
"""LLM reasons over tree structure to identify relevant nodes."""
# Strip raw text to save tokens; the LLM only needs titles and summaries
tree_without_text = utils.remove_fields(
tree.copy(), fields=["text"]
) search_prompt = f"""
You are a document retrieval expert. Given a question and
a hierarchical tree structure of a document, identify all
nodes likely to contain the answer. Each node has a node_id, title, and summary.
Follow cross-references if a section mentions another. Question: {query} Document tree structure:
{json.dumps(tree_without_text, indent=2)} Reply in this JSON format only:
{{
"thinking": "<reasoning about which nodes are relevant>",
"node_list": ["node_id_1", "node_id_2"]
}}
""" response = await openai_client.chat.completions.create(
model="gpt-4.1",
messages=[{"role": "user", "content": search_prompt}],
temperature=0,
response_format={"type": "json_object"},
) result = json.loads(response.choices[0].message.content)
print(f"LLM reasoning: {result['thinking']}")
return result["node_list"]
Retrieve Content and Generate the Answer
For the Retrieve Content and Generate stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
def collect_node_content(tree: dict, node_ids: list) -> str:
"""Pull raw text from the nodes the LLM selected."""
all_nodes = utils.flatten_tree(tree)
context_parts = []
for node in all_nodes:
if node["node_id"] in node_ids:
title = node.get("title", "Untitled")
pages = f"pages {node.get('start_index', '?')}-{node.get('end_index', '?')}"
text = node.get("text", "")
context_parts.append(
f"[{title} | {pages}]\n{text}"
)
return "\n\n---\n\n".join(context_parts)
async def answer_query(tree: dict, query: str) -> dict:
"""Full vectorless RAG pipeline: tree search + answer generation.""" # Step 1: LLM picks the nodes
node_ids = await find_relevant_nodes(tree, query) # Step 2: Fetch content from those nodes
context = collect_node_content(tree, node_ids) # Step 3: Generate answer with citations
answer_prompt = f"""
Answer the question using only the provided context.
Cite specific pages and sections in your answer. Context:
{context} Question: {query}
""" response = await openai_client.chat.completions.create(
model="gpt-4.1",
messages=[{"role": "user", "content": answer_prompt}],
temperature=0,
) return {
"answer": response.choices[0].message.content,
"retrieved_nodes": node_ids,
"context_length": len(context),
}
# Run it
query = "What was the total value of deferred assets in 2023?"
result = asyncio.run(answer_query(tree, query))
print(result["answer"])
print(f"Nodes used: {result['retrieved_nodes']}")
Bonus: MCP Integration
For the Bonus MCP Integration stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Bonus MCP Integration stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
{
"mcpServers": {
"pageindex": {
"type": "http",
"url": "https://api.pageindex.ai/mcp",
"headers": {
"Authorization": "Bearer your_api_key"
}
}
}
}
{
"mcpServers": {
"pageindex": {
"command": "npx",
"args": ["-y", "@pageindex/mcp"]
}
}
}
The Benchmark Numbers (With Context)
When working through the The Benchmark Numbers With stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Where PageIndex Falls Short (And It Does)
When working through the Where PageIndex Falls Short stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
So When Should You Actually Use This?
When working through the So When Should You stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the So When Should You stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
What’s Happened Since Launch (Recent Developments)
The What s Happened Since stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
pip install openai-agents
python3 examples/agentic_vectorless_rag_demo.py
The Bigger Picture
The The Bigger Picture stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Getting Started
The Getting Started stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Getting Started stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Operational checklist
When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.
Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for d194e0549478: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.
When working through the hardening note 0 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 0/751: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 1 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 1/751: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.