Home / Articles / Practical notes: LlamaIndex RAG: A Practical Guide to Building Smarter AI

This article is published in English.

Practical notes: LlamaIndex RAG: A Practical Guide to Building Smarter AI

Operable walkthrough of Practical notes: LlamaIndex RAG: A Practical Guide to Building Smarter AI: contracts, checks, and drop-in code slots for teams shipping this pattern.

3766 words

Use this as an operator-facing rebuild of the ideas in “LlamaIndex RAG: A Practical Guide to Building Smarter AI Applications”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

What Exactly Is RAG?

For the What Exactly Is RAG stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

User
  ↓
Question
  ↓
LLM
  ↓
Answer
User Question
                         ↓
                    Retrieval
                         ↓
                 Relevant Documents
                         ↓
                    LLM Prompt
                         ↓
                       LLM
                         ↓
                      Answer

Where LlamaIndex Fits

For the Where LlamaIndex Fits stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

Your Data
                       │
          ┌────────────┼────────────┐
          ↓            ↓            ↓
       PDFs         Websites     Databases
          │            │            │
          └────────────┼────────────┘
                       ↓
                   LlamaIndex
                       ↓
                 Data Ingestion
                       ↓
                    Chunks
                       ↓
                  Embeddings
                       ↓
                 Vector Store
                       ↓
                    Retriever
                       ↓
                   Reranker
                       ↓
                     LLM
                       ↓
                    Answer

The LlamaIndex RAG Pipeline

For the The LlamaIndex RAG Pipeline stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the The LlamaIndex RAG Pipeline stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Documents
    ↓
Loading
    ↓
Parsing
    ↓
Chunking
    ↓
Indexing
    ↓
Retrieval
    ↓
Context Selection
    ↓
Generation

1. Load Your Data

When working through the 1 Load Your Data stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

PDFs
Markdown
Web pages
Notion
Google Drive
SQL databases
APIs
CSV files
documents = load_documents("data/")
Offline / ingestion time
        ↓
Prepare the knowledge
Online / query time
        ↓
Retrieve the knowledge

2. Split Documents Into Chunks

When working through the 2 Split Documents Into stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Document
   ↓
Chapter
   ↓
Section
   ↓
Paragraph
   ↓
Chunk
chunks = split_document(
    document,
    chunk_size=512,
)
Huge chunk
    ↓
Lots of irrelevant information
    ↓
Large prompt
    ↓
Higher latency
Tiny chunk
    ↓
Missing context
    ↓
Poor retrieval

3. Create Embeddings

When working through the 3 Create Embeddings stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the 3 Create Embeddings stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

"How can I reset my password?"

"What should I do if I forgot my login credentials?"
Text
 ↓
Embedding Model
 ↓
[0.12, -0.42, 0.81, ...]

4. Store the Vectors

The 4 Store the Vectors stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

Document
   ↓
Chunk
   ↓
Embedding
   ↓
Vector Store
Vector Store

ID     Vector        Metadata
--------------------------------
001    [....]        product=api
002    [....]        product=web
003    [....]        product=mobile
document_id
page_number
department
product
version
created_at
tenant_id
access_level

5. Retrieve Relevant Information

The 5 Retrieve Relevant Information stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

Question
   ↓
Query Embedding
   ↓
Vector Search
Top 5 Results

1. API Authentication Guide
2. OAuth Configuration
3. API Token Documentation
4. Authentication Troubleshooting
5. Security Configuration

6. Turn Retrieved Documents Into Context

The 6 Turn Retrieved Documents stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The 6 Turn Retrieved Documents stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

User Question
      +
Retrieved Context
      ↓
Prompt
      ↓
LLM
System:
Answer using the supplied context.

Context:
[Relevant document 1]

[Relevant document 2]

[Relevant document 3]

Question:
How do I configure API authentication?

7. Generate the Answer

For the 7 Generate the Answer stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

Question
+
Relevant Context
                  User
                    ↓
                  Query
                    ↓
              Query Embedding
                    ↓
              Vector Retrieval
                    ↓
             Relevant Chunks
                    ↓
             Context Assembly
                    ↓
                   LLM
                    ↓
                 Answer

LlamaIndex Is More Than “Vector Search + LLM”

For the LlamaIndex Is More Than stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

Query
 ↓
Vector Search
 ↓
Top 5 Documents
 ↓
LLM
                      Query
                         ↓
                  Query Processing
                         ↓
               ┌─────────┴─────────┐
               ↓                   ↓
          Dense Search        Keyword Search
               ↓                   ↓
               └─────────┬─────────┘
                         ↓
                      Fusion
                         ↓
                      Rerank
                         ↓
                  Context Selection
                         ↓
                       LLM

Query Engines: Turning Retrieval Into a Question-Answering System

For the Query Engines Turning Retrieval stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Query Engines Turning Retrieval stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

query_embedding = embed(query)

documents = search(query_embedding)

context = build_context(documents)

answer = llm.generate(
    query=query,
    context=context,
)
query_engine = index.as_query_engine()

response = query_engine.query(
    "How does authentication work?"
)

Retrieval Is Where Many RAG Systems Fail

When working through the Retrieval Is Where Many stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

User Question
      ↓
Bad Retrieval
      ↓
Wrong Context
      ↓
LLM
      ↓
Bad Answer

Improve Retrieval With Metadata

When working through the Improve Retrieval With Metadata stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Product A
Product B
Product C
product = Product B
version = 3
document_type = documentation
Entire Knowledge Base
        ↓
Metadata Filter
        ↓
Relevant Subset
        ↓
Semantic Search

Hybrid Search Can Be Better Than Vector Search Alone

When working through the Hybrid Search Can Be stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Hybrid Search Can Be stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

ERR_CONNECTION_RESET_502
Dense Retrieval
      +
Sparse Retrieval
      ↓
Result Fusion
      ↓
Reranking

Reranking: Spend More Compute Only on the Best Candidates

The Reranking Spend More Compute stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

Top 20 documents
Vector Search
     ↓
20 candidates
     ↓
Reranker
     ↓
Top 5
     ↓
LLM
Retriever
→ Find potentially relevant documents

Reranker
→ Determine which are actually relevant

LLM
→ Use those documents to answer

Context Is a Limited Resource

The Context Is a Limited stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

20 chunks
×
500 tokens
=
10,000 tokens
Retrieve 20
     ↓
Rerank
     ↓
Keep 5
     ↓
Compress
     ↓
Send 2,500 tokens

LlamaIndex RAG for PDFs

The LlamaIndex RAG for PDFs stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The LlamaIndex RAG for PDFs stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

PDF Files
    ↓
Document Loading
    ↓
Text Extraction
    ↓
Chunking
    ↓
Embeddings
    ↓
Vector Store
    ↓
Retriever
    ↓
LLM
Company Handbook
        ↓
Employee Documentation
        ↓
HR Policies
        ↓
Benefits
        ↓
Leave Policies

LlamaIndex RAG for AI Applications

For the LlamaIndex RAG for AI stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

User
 ↓
RAG
 ↓
Answer
                      User
                        ↓
                      Agent
                        ↓
              ┌─────────┼─────────┐
              ↓         ↓         ↓
             RAG     Database    API
              ↓         ↓         ↓
              └─────────┼─────────┘
                        ↓
                       LLM
                        ↓
                     Answer

RAG Latency Matters

For the RAG Latency Matters stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

Query
 ↓
Embedding API
 ↓
Vector Database
 ↓
Reranker
 ↓
LLM

Retrieve fewer documents

For the Retrieve fewer documents stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Retrieve fewer documents stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Use metadata filters

When working through the Use metadata filters stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Parallelize independent retrieval

When working through the Parallelize independent retrieval stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Dense ──────┐
            ├──→ Fusion
Sparse ─────┘
Dense
 ↓
Sparse
 ↓
Fusion

Reduce prompt size

When working through the Reduce prompt size stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the Reduce prompt size stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Stream the response

The Stream the response stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

RAG Is Not Just About Accuracy

The RAG Is Not Just stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

                RAG Quality
                     │
        ┌────────────┼────────────┐
        ↓            ↓            ↓
   Retrieval      Generation    System
     Quality        Quality     Performance
        │            │            │
    Recall        Faithfulness  Latency
    Precision     Relevance     Cost
    Ranking       Completeness  Reliability

Common Mistakes When Building LlamaIndex RAG

The Common Mistakes When Building stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

Mistake 1: Treating the LLM as the whole system

The Mistake 1 Treating the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

Mistake 2: Using giant chunks

The Mistake 2 Using giant stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Mistake 3: Retrieving too much

Mistake 4: Ignoring metadata

Mistake 5: Skipping evaluation

Mistake 6: Assuming vector search is enough

A Production-Oriented LlamaIndex RAG Architecture

                         User Query
                              │
                              ▼
                       Query Processing
                              │
                              ▼
                       Query Router
                              │
               ┌──────────────┴──────────────┐
               │                             │
          Direct Answer                 Retrieval Needed
                                             │
                                             ▼
                                    Metadata Filtering
                                             │
                          ┌──────────────────┴──────────────────┐
                          ▼                                     ▼
                    Dense Search                         Sparse Search
                          │                                     │
                          └──────────────────┬──────────────────┘
                                             ▼
                                           Fusion
                                             │
                                             ▼
                                          Rerank
                                             │
                                             ▼
                                      Context Selection
                                             │
                                             ▼
                                            LLM
                                             │
                                             ▼
                                         Response

Start With the Simplest Possible RAG

Documents
   ↓
Chunk
   ↓
Embed
   ↓
Vector Store
   ↓
Retrieve
   ↓
LLM
Add metadata
Add hybrid retrieval
Add reranking
Add context compression
Cache + parallelize + reduce retrieval

The Real Power of LlamaIndex

Documents
Databases
APIs
Knowledge Bases
Search Systems
Structured Data
Unstructured Data
LLM
                  AI Application
                         │
        ┌────────────────┼────────────────┐
        ↓                ↓                ↓
      LLM              Tools            Data
                         │                │
                         └───────┬────────┘
                                 ↓
                              Retrieval
                                 ↓
                              Context
                                 ↓
                                LLM

Final Thoughts

Data
 ↓
Ingestion
 ↓
Indexing
 ↓
Retrieval
 ↓
Context
 ↓
LLM
 ↓
Answer

Operational checklist