This article is published in English.
RAG Explained: Stop Chatbots From Inventing Company Facts
A practical walkthrough of the RAG pipeline—loaders, chunking, embeddings, vector stores, re-ranking, hybrid search, and RRF—plus when not to use retrieval at all.
A common early chatbot assignment goes like this: answer questions from company files. A quick prototype often sounds polished—and invents facts. It may claim “24/7 support in 12 countries” when support exists in one country, business hours only, and only when a particular IT person is available.
That failure mode is the whole point of RAG (Retrieval-Augmented Generation): a model that can produce an answer is not the same as a model that has the organization’s real data. Without grounding, fluent models behave like a confident relative inventing family history at a wedding—about forty percent correct and one hundred percent sure.
Part 1: What is RAG, actually?
RAG means Retrieval-Augmented Generation. The idea is plain.
Rather than asking the model to answer from training memory alone, the system first fetches relevant material, then asks for an answer constrained to that material.
Treat the text model as a sharp, overconfident intern: excellent at writing, weak at recalling the 2023 HR policy. RAG’s job is to put the right page in the intern’s hands before the reply starts.
Without RAG: replies lean on memory (stale, generic, or company-blind). With RAG: replies lean on memory plus files fetched for this question.
A compact formula that sticks:
Right Information × Right Context × Right Prompt = Right Answer
Part 2: The RAG Pipeline
Retrieval augmentation is not one magic call. It is a multi-stop route: if any stop breaks, answers arrive late or wrong.
Stop 1 — Knowledge Sources
Where does material live? PDFs, Word files, Notion pages, SQL tables, Slack exports, and forgotten spreadsheets all count as knowledge sources.
Stop 2 — Document Loaders (Getting the data IN)
Ingestion turns PDF/HTML/CSV/etc. into a clean, uniform shape—often something like:
Document(
page_content="The quick brown fox...",
metadata={"source": "annual_report.pdf", "page": 12, "author": "Finance Team"}
)
Skipping careful loading and dumping raw PDF bytes into the flow tends to promote headers and footers into “facts” (“According to page 47 of 92…”). Nobody asked for the page chrome.
Lesson: dirty ingestion produces dirty retrieval and dirty answers. Clean at the door.
Stop 3 — Chunking (a.k.a. slicing the pizza)
A 200-page PDF cannot be dropped into a prompt. Context windows are finite, and flooding the model with an entire cookbook when one recipe was requested hurts precision.
Documents are therefore split into segments.
Common strategies:
- Fixed-size chunking — cut every N tokens. Fast, but mid-sentence cuts happen.
- Fixed-size with overlap — same cuts with slight overlap so boundary context survives. A common default.
- Hierarchical/recursive chunking — honor sections and paragraphs before cutting.
- Semantic chunking — cluster same-topic sentences via embeddings. Smarter, more compute.
- LLM-based / Agentic chunking — ask a model where cuts belong. Costly, useful for legal/medical text.
A practical rule after watching a contract chunk split “shall NOT be liable” after “shall”: start with fixed size plus overlap; add complexity only when retrieval quality is actually poor.
Stop 4 — Embeddings (turning words into GPS coordinates)
An embedding turns a sentence into a number list that captures meaning, not spelling. Near-meaning sentences land near each other even with zero shared words.
“How can a password be reset?” and “Login credentials were forgotten” share almost no tokens yet mean nearly the same thing. Keyword search misses that link; an embedder catches it by comparing meaning.
History moved from Bag-of-Words (counts only) → TF-IDF → Word2Vec → BERT → modern embedding APIs and open models (OpenAI, Cohere, BGE, E5, and peers) that respect context.
Analogy: Bag-of-Words is word-by-word machine translation. Modern embeddings are closer to someone who lived in both cultures and hears the intent.
Stop 5 — Vector Stores (the library where you keep all those number-lists)
After segments become vectors, they need fast storage and search: Pinecone, Qdrant, Weaviate, Milvus, ChromaDB, FAISS, and similar systems.
Given a query, return the top-K segments whose vectors sit closest to the query vector. Indexing options include:
- Brute force — compare everything. Exact, slow at scale.
- ANN (Approximate Nearest Neighbor) — slightly less exact, much faster. Solid default.
- IVF — cluster first; search only promising clusters.
- HNSW — walk a neighbor graph quickly. Common in production RAG.
Brute force on millions of vectors “for perfect accuracy” can turn millisecond queries into tea-break latency. Match the index to data size, not pride.
Stop 6 — Retrieval (Similarity Search)
When a question arrives, embed it with the same model used for files—mixing models is a classic foot-gun—then fetch top-K similar segments.
Cosine similarity is the usual measure: how aligned two vectors are in direction, ignoring length. It is a stable default across varying text lengths.
Stop 7 — Augmentation (feeding the intern the right file)
Retrieved segments are inserted into the prompt with the user question. That is the “Augmented” step:
System: You are a helpful assistant. Only answer using the context below.
Context: [retrieved chunk 1] [retrieved chunk 2] [retrieved chunk 3]
Question: What is our refund policy?
Stop 8 — Generation (the LLM finally speaks)
Only then does the model write the final answer—anchored in fetched context rather than vibes.
Part 3: The stuff that separates “it works” from “it works well”
Tutorials often stop at the basic flow. Live systems quality usually needs more.
Re-ranking — because your first search result isn’t always your best result
Bi-encoder retrieval is fast but coarse: query and document are encoded separately, then compared. It is excellent for narrowing millions of items to ~50 candidates.
“Fast and roughly right” is not always “actually right,” so a slower cross-encoder re-ranker scores query and document together, catching nuance, negation, and context. It runs only on the shortlist and lifts the true top results.
Analogy: bi-encoders skim resumes for a shortlist; cross-encoders run the interview.
In a medical FAQ bot without re-ranking, a question about alcohol interactions can surface a different medicine that merely shares vocabulary. A cross-encoder re-ranker often fixes that immediately. When precision matters (legal, medical, finance), re-ranking is mandatory, not optional.
Hybrid Search — because lexical search and vector search are both a little blind, alone
Vector search captures meaning but can miss exact tokens—SKU codes, error IDs, proper names. BM25 lexical search nails exact strings but misses synonyms.
Hybrid search blends BM25 and vector retrieval: keyword precision plus semantic recall. When unsure, hybrid is a strong default—rarely worse, often better.
RAG Fusion — combining multiple opinions like a group project (but it actually works this time)
Multiple retrievers (dense, keyword, domain-specific) produce multiple ranked lists. RAG Fusion merges them with Reciprocal Rank Fusion (RRF).
The idea is simple even when the formula looks formal: a document that ranks near the top across several independent methods is probably relevant. RRF rewards cross-list consistency rather than raw scores that are not comparable across retrievers.
RRF(d) = Σ 1 / (k + rank_i(d))
Here k is commonly 60—an industry convention more than a derived constant.
Metadata — the labels that save your life later
Metadata is data about data: title, author, date, source, access level. It feels dull until an access rule appears—“never show internal HR docs to external users”—and access_level: internal tags suddenly matter.
Metadata enables filters, re-rank boosts, access control, and debugging (“why did a 2019 doc answer a 2026 question?” → missing date filters).
Memory and Caching — because nobody wants to pay for the same computation twice
Two ideas often confused:
- Memory = longer-term recall of preferences, prior turns, lasting facts.
- Caching = short-term reuse of expensive results (model replies, search hits) for repeated asks.
Memory makes assistants feel consistent. Caching makes them fast and cheap. Mixing them—for example “caching” preferences with a ten-minute TTL—can erase a user’s name mid-chat. Keep the concepts separate.
Part 4: When should you NOT use RAG?
Retrieval augmentation is not a universal hammer.
Skip RAG when:
- The question is general knowledge the model already handles (“capital of France” needs no vector index).
- Facts change continuously (prices, live scores)—prefer an API.
- The task is purely creative (poetry, brainstorming)—retrieval rarely helps.
- The corpus is small enough to place directly in the prompt.
- Ultra-low latency cannot tolerate an extra retrieval hop.
Sometimes the fix is prompting or fine-tuning, not a full vector flow. Choose the tool before building plumbing.
Part 5: Best practices learned the hard way
- Chunk smart, not tiny. Tiny segments lose context; huge segments lose precision. Prefer overlap.
- Use one embedder for files and queries. Mixing models is measuring in mixed units.
- Add re-ranking when accuracy matters. Small latency for large quality gains.
- Tag metadata from day one. Future filters will need it.
- Evaluate constantly. Track Recall@k / Precision@k and generation faithfulness/relevance. Vibes are not a metric.
- Cache expensive stable work; refresh what changes. Do not freeze fast-moving facts; do not recompute static ones.
- Keep guardrails. Moderation, citations, and hallucination checks prevent confident lies in demos.
Teams that treat retrieval as an afterthought usually rediscover the same failure modes: silent schema drift in loaders, chunk boundaries that bisect negations, embedding model mismatches between offline indexing and online queries, and dashboards that only measure “answer latency” while ignoring faithfulness. A durable RAG practice treats each stop on the route as a product surface with owners, tests, and rollback plans—especially when the assistant sits in front of customers or regulated workflows.
When evaluating changes, prefer paired experiments: same question set, same judge rubric, compare Recall@k and groundedness before and after a chunking or re-ranker tweak. Small retrieval gains often beat larger prompt-only tweaks because the generator cannot cite what never entered the context window.
The RAG Mantra
Right Information → Right Context → Right Prompt → Right Answer.
RAG is plumbing, not magic. Clean intake, sensible retrieval, honest augmentation, and a tight generation prompt turn an overconfident cousin into a briefed specialist who read the file first—and stop inventing global 24/7 support that never existed.