This article is published in English.
Practical notes: LlamaIndex RAG: A Practical Guide to Building Smarter AI
Operable walkthrough of Practical notes: LlamaIndex RAG: A Practical Guide to Building Smarter AI: contracts, checks, and drop-in code slots for teams shipping this pattern.
Use this as an operator-facing rebuild of the ideas in “LlamaIndex RAG: A Practical Guide to Building Smarter AI Applications”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
What Exactly Is RAG?
For the What Exactly Is RAG stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
User
↓
Question
↓
LLM
↓
Answer
User Question
↓
Retrieval
↓
Relevant Documents
↓
LLM Prompt
↓
LLM
↓
Answer
Where LlamaIndex Fits
For the Where LlamaIndex Fits stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Your Data
│
┌────────────┼────────────┐
↓ ↓ ↓
PDFs Websites Databases
│ │ │
└────────────┼────────────┘
↓
LlamaIndex
↓
Data Ingestion
↓
Chunks
↓
Embeddings
↓
Vector Store
↓
Retriever
↓
Reranker
↓
LLM
↓
Answer
The LlamaIndex RAG Pipeline
For the The LlamaIndex RAG Pipeline stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the The LlamaIndex RAG Pipeline stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Documents
↓
Loading
↓
Parsing
↓
Chunking
↓
Indexing
↓
Retrieval
↓
Context Selection
↓
Generation
1. Load Your Data
When working through the 1 Load Your Data stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
PDFs
Markdown
Web pages
Notion
Google Drive
SQL databases
APIs
CSV files
documents = load_documents("data/")
Offline / ingestion time
↓
Prepare the knowledge
Online / query time
↓
Retrieve the knowledge
2. Split Documents Into Chunks
When working through the 2 Split Documents Into stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Document
↓
Chapter
↓
Section
↓
Paragraph
↓
Chunk
chunks = split_document(
document,
chunk_size=512,
)
Huge chunk
↓
Lots of irrelevant information
↓
Large prompt
↓
Higher latency
Tiny chunk
↓
Missing context
↓
Poor retrieval
3. Create Embeddings
When working through the 3 Create Embeddings stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the 3 Create Embeddings stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
"How can I reset my password?"
"What should I do if I forgot my login credentials?"
Text
↓
Embedding Model
↓
[0.12, -0.42, 0.81, ...]
4. Store the Vectors
The 4 Store the Vectors stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Document
↓
Chunk
↓
Embedding
↓
Vector Store
Vector Store
ID Vector Metadata
--------------------------------
001 [....] product=api
002 [....] product=web
003 [....] product=mobile
document_id
page_number
department
product
version
created_at
tenant_id
access_level
5. Retrieve Relevant Information
The 5 Retrieve Relevant Information stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Question
↓
Query Embedding
↓
Vector Search
Top 5 Results
1. API Authentication Guide
2. OAuth Configuration
3. API Token Documentation
4. Authentication Troubleshooting
5. Security Configuration
6. Turn Retrieved Documents Into Context
The 6 Turn Retrieved Documents stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The 6 Turn Retrieved Documents stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
User Question
+
Retrieved Context
↓
Prompt
↓
LLM
System:
Answer using the supplied context.
Context:
[Relevant document 1]
[Relevant document 2]
[Relevant document 3]
Question:
How do I configure API authentication?
7. Generate the Answer
For the 7 Generate the Answer stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Question
+
Relevant Context
User
↓
Query
↓
Query Embedding
↓
Vector Retrieval
↓
Relevant Chunks
↓
Context Assembly
↓
LLM
↓
Answer
LlamaIndex Is More Than “Vector Search + LLM”
For the LlamaIndex Is More Than stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
Query
↓
Vector Search
↓
Top 5 Documents
↓
LLM
Query
↓
Query Processing
↓
┌─────────┴─────────┐
↓ ↓
Dense Search Keyword Search
↓ ↓
└─────────┬─────────┘
↓
Fusion
↓
Rerank
↓
Context Selection
↓
LLM
Query Engines: Turning Retrieval Into a Question-Answering System
For the Query Engines Turning Retrieval stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Query Engines Turning Retrieval stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
query_embedding = embed(query)
documents = search(query_embedding)
context = build_context(documents)
answer = llm.generate(
query=query,
context=context,
)
query_engine = index.as_query_engine()
response = query_engine.query(
"How does authentication work?"
)
Retrieval Is Where Many RAG Systems Fail
When working through the Retrieval Is Where Many stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
User Question
↓
Bad Retrieval
↓
Wrong Context
↓
LLM
↓
Bad Answer
Improve Retrieval With Metadata
When working through the Improve Retrieval With Metadata stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Product A
Product B
Product C
product = Product B
version = 3
document_type = documentation
Entire Knowledge Base
↓
Metadata Filter
↓
Relevant Subset
↓
Semantic Search
Hybrid Search Can Be Better Than Vector Search Alone
When working through the Hybrid Search Can Be stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Hybrid Search Can Be stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
ERR_CONNECTION_RESET_502
Dense Retrieval
+
Sparse Retrieval
↓
Result Fusion
↓
Reranking
Reranking: Spend More Compute Only on the Best Candidates
The Reranking Spend More Compute stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Top 20 documents
Vector Search
↓
20 candidates
↓
Reranker
↓
Top 5
↓
LLM
Retriever
→ Find potentially relevant documents
Reranker
→ Determine which are actually relevant
LLM
→ Use those documents to answer
Context Is a Limited Resource
The Context Is a Limited stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
20 chunks
×
500 tokens
=
10,000 tokens
Retrieve 20
↓
Rerank
↓
Keep 5
↓
Compress
↓
Send 2,500 tokens
LlamaIndex RAG for PDFs
The LlamaIndex RAG for PDFs stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The LlamaIndex RAG for PDFs stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
PDF Files
↓
Document Loading
↓
Text Extraction
↓
Chunking
↓
Embeddings
↓
Vector Store
↓
Retriever
↓
LLM
Company Handbook
↓
Employee Documentation
↓
HR Policies
↓
Benefits
↓
Leave Policies
LlamaIndex RAG for AI Applications
For the LlamaIndex RAG for AI stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
User
↓
RAG
↓
Answer
User
↓
Agent
↓
┌─────────┼─────────┐
↓ ↓ ↓
RAG Database API
↓ ↓ ↓
└─────────┼─────────┘
↓
LLM
↓
Answer
RAG Latency Matters
For the RAG Latency Matters stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Query
↓
Embedding API
↓
Vector Database
↓
Reranker
↓
LLM
Retrieve fewer documents
For the Retrieve fewer documents stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Retrieve fewer documents stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Use metadata filters
When working through the Use metadata filters stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Parallelize independent retrieval
When working through the Parallelize independent retrieval stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Dense ──────┐
├──→ Fusion
Sparse ─────┘
Dense
↓
Sparse
↓
Fusion
Reduce prompt size
When working through the Reduce prompt size stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the Reduce prompt size stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Stream the response
The Stream the response stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
RAG Is Not Just About Accuracy
The RAG Is Not Just stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
RAG Quality
│
┌────────────┼────────────┐
↓ ↓ ↓
Retrieval Generation System
Quality Quality Performance
│ │ │
Recall Faithfulness Latency
Precision Relevance Cost
Ranking Completeness Reliability
Common Mistakes When Building LlamaIndex RAG
The Common Mistakes When Building stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Mistake 1: Treating the LLM as the whole system
The Mistake 1 Treating the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
Mistake 2: Using giant chunks
The Mistake 2 Using giant stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Mistake 3: Retrieving too much
Mistake 4: Ignoring metadata
Mistake 5: Skipping evaluation
Mistake 6: Assuming vector search is enough
A Production-Oriented LlamaIndex RAG Architecture
User Query
│
▼
Query Processing
│
▼
Query Router
│
┌──────────────┴──────────────┐
│ │
Direct Answer Retrieval Needed
│
▼
Metadata Filtering
│
┌──────────────────┴──────────────────┐
▼ ▼
Dense Search Sparse Search
│ │
└──────────────────┬──────────────────┘
▼
Fusion
│
▼
Rerank
│
▼
Context Selection
│
▼
LLM
│
▼
Response
Start With the Simplest Possible RAG
Documents
↓
Chunk
↓
Embed
↓
Vector Store
↓
Retrieve
↓
LLM
Add metadata
Add hybrid retrieval
Add reranking
Add context compression
Cache + parallelize + reduce retrieval
The Real Power of LlamaIndex
Documents
Databases
APIs
Knowledge Bases
Search Systems
Structured Data
Unstructured Data
LLM
AI Application
│
┌────────────────┼────────────────┐
↓ ↓ ↓
LLM Tools Data
│ │
└───────┬────────┘
↓
Retrieval
↓
Context
↓
LLM
Final Thoughts
Data
↓
Ingestion
↓
Indexing
↓
Retrieval
↓
Context
↓
LLM
↓
Answer