This article is published in English.
RAG Explained: Giving LLMs Access to External Knowledge
A beginner-friendly RAG walkthrough—chunking, embeddings, vector search, hybrid retrieval, and when RAG still hallucinates—with clear architecture diagrams.
Large language models answer from training memory. That memory is powerful and incomplete: it can miss internal handbooks, outdated policies, and facts that never appeared in public text. Retrieval-Augmented Generation (RAG) closes the gap by fetching relevant external material first, then asking the model to answer from that material.
What is RAG?
Without retrieval, a question goes straight to the model:
User Question
↓
LLM
↓
Answer
With RAG, the system finds relevant information before generation:
User Question
↓
Find Relevant Information
↓
Give Information to the LLM
↓
LLM
↓
Answer
The model still writes the final wording; the new ingredient is grounded context from a knowledge store.
Why Do We Need RAG?
Training cutoffs, private corpora, and fast-changing policies all break “memory-only” answers. An employee handbook updated last week will not be inside a frontier model’s weights. RAG lets the assistant consult that handbook at ask-time without retraining.
A Simple RAG Example
An employee asks about vacation carryover. The flow looks like:
Employee asks a question
↓
Search the employee handbook
↓
Find the relevant section
↓
Give that section to the LLM
↓
LLM generates the answer
The assistant should quote the handbook, not invent a policy that “sounds right.”
How Does RAG Work?
Two phases matter: prepare knowledge offline, then answer online.
1. Prepare the knowledge
Ingest files, split them, embed chunks, and store vectors for search.
2. Answer the question
Embed the question, fetch top chunks, pack them into a prompt, generate.
Part 1: Preparing the Knowledge
Typical sources for an HR assistant:
Employee Handbook
Vacation Policy
Benefits Guide
Leave Policy
The preparation pipeline:
Documents
↓
Extract Text
↓
Break into Smaller Pieces
↓
Create Embeddings
↓
Store for Search
Step 1: Get the Information
Extract clean text from each source:
Employee Handbook
↓
Extract text
↓
"Employees receive..."
"Vacation requests..."
"Leave policy..."
Headers, footers, and navigation chrome should be stripped early so they never become “facts.”
Step 2: Break the Document into Chunks
Long handbooks exceed context limits and bury the relevant paragraph. Chunking produces searchable units:
Employee Handbook
↓
┌───────────────┐
│ Chunk 1 │
├───────────────┤
│ Chunk 2 │
├───────────────┤
│ Chunk 3 │
├───────────────┤
│ ... │
└───────────────┘
Why does chunking matter?
Chunks that are too large dilute relevance; chunks that are too small lose surrounding rules (especially negations and exceptions). Overlap between neighbors preserves boundary context. Start simple—fixed size with overlap—then tune when evaluations show misses.
Step 3: Create Embeddings
Embeddings map text to vectors so meaning-similar phrases sit nearby even without shared keywords.
Question text:
"What is my vacation allowance?"
A paraphrase with the same intent:
"How many annual leave days do I get?"
Both can land near the same vacation-policy chunk after embedding:
"What is my vacation allowance?"
↓
Embedding
↓
Numerical representation
Use the same embedder for indexing and querying; mixing models silently breaks nearest-neighbor search.
Step 4: Store the Information for Search
Each chunk plus its vector goes into a vector index (or hybrid store):
Document Chunk
↓
Embedding
↓
Vector Database
Metadata—source title, page, access level, effective date—should travel with the chunk for filters and citations later.
Part 2: Answering the User’s Question
Online path:
User Question
↓
Understand the question
↓
Search the stored information
↓
Find relevant chunks
↓
Give those chunks to the LLM
↓
Generate an answer
Step 5: Retrieve Relevant Information
The question embedding retrieves top chunks, for example:
Chunk 147 → Vacation carryover policy
Chunk 148 → Vacation request process
Chunk 62 → Employee benefits
Chunk 300 → Security policy
Rank quality here dominates final answer quality. Bad retrieval cannot be fixed by a prettier prompt.
Step 6: Give the Information to the LLM
Fetched text becomes prompt context:
Context:Employees may carry over up to 5 unused
vacation days into the following year.
Question:How many vacation days can I carry over?
Instructions should require answering from context and admitting gaps when context is missing.
Step 7: Generate the Answer
Retrieved Information
+
User Question
↓
LLM
↓
Answer
The model composes a reply grounded in retrieved lines rather than generic HR folklore.
The Complete RAG Architecture
Offline preparation label:
KNOWLEDGE PREPARATION
End-to-end diagram:
Documents
↓
Extract Text
↓
Chunking
↓
Embeddings
↓
Vector Database
│
│
│
▼
USER QUESTION
↓
Query Embedding
↓
Retrieval
↓
Relevant Information
↓
Question + Context
↓
LLM
↓
Answer
Before the question
Documents
↓
Chunks
↓
Embeddings
↓
Vector Database
When the question arrives
Question
↓
Retrieval
↓
Relevant Context
↓
LLM
↓
Answer
Does RAG Only Use Vector Search?
No. Production systems often combine methods.
Keyword Search
Lexical matchers (BM25 and friends) excel at exact tokens: policy codes, SKUs, error IDs, proper names.
Semantic Search
Vector search catches paraphrases and synonyms that keywords miss.
Hybrid Search
Blend both, then merge rankings:
Keyword Search
+
Semantic Search
↓
Hybrid Search
Hybrid is a strong default when traffic mixes exact identifiers with natural-language asks.
RAG Does Not Eliminate Hallucinations
Retrieval reduces unsupported invention; it does not remove it. Failure modes include:
Wrong chunk fetched:
User Question
↓
Wrong information retrieved
↓
LLM
↓
Wrong answer
Nothing relevant found, yet the model still answers:
User Question
↓
No relevant information found
↓
LLM
↓
Unsupported answer
Mitigations: tighter retrieval, re-ranking, refusal instructions, citations, and evaluation on faithfulness—not only fluency.
RAG vs Fine-Tuning
RAG
Best when knowledge changes often, must be cited, or stays private outside training data. Updates mean re-indexing, not re-training.
Fine-tuning
Best when behavior, style, or task format must change, or when knowledge is stable and compact enough to bake in. Costly to refresh when policies churn.
Many products use both: fine-tune for skills, RAG for facts.
Where Is RAG Useful?
Customer Support
Assistants grounded in product docs and ticket macros.
Healthcare
Protocol and guideline lookup with strict citation and access controls (domain rules still apply).
Finance
Policy, disclosure, and product-rule answers that must track the latest approved text.
Human Resources
Handbooks, benefits, leave rules—exactly the employee-question pattern above.
Software Development
Internal ADRs, runbooks, and API references beside public docs.
When Should You Use RAG?
Use RAG when answers must reflect a specific corpus that is larger than a prompt, changes faster than fine-tuning cycles, or requires provenance. Skip RAG for pure trivia the base model already knows, pure creativity, or ultra-low-latency paths that cannot afford a retrieval hop.
What Can Make RAG Difficult?
Chunk boundaries, embedding drift, stale indexes, access-control leaks, and evaluation gaps. A concrete misshape:
User asks:
User asks:
"What is the vacation carryover policy?"
System retrieves the wrong policy family:
↓System retrieves:
"Health insurance policy" ↓LLM receives wrong context ↓Poor answer
The model then sounds confident while explaining health insurance as if it were vacation carryover. Fix retrieval and metadata filters before blaming the generator.
What Do You Need to Learn to Build RAG?
A practical skill ladder:
RAG Fundamentals
↓
Document Processing
↓
Chunking
↓
Embeddings
↓
Vector Databases
↓
Retrieval
↓
Prompt + Context
↓
LLM
↓
Evaluation
Document processing, chunking strategy, embeddings, vector stores, prompt assembly, evaluation, and ops (rebuilds, access, monitoring) all matter.
The Big Picture
Knowledge
↓
Find relevant
information
↓
Give it to LLM
↓
Generate
answer
Knowledge lives outside the model; fetching bridges the gap; generation stays the last mile.
Chunking choices interact with embedding dimensionality and index type: dense-only HNSW setups behave differently from hybrid BM25+vector stores when queries contain both prose and identifiers. Keep a written policy for rebuild triggers—new handbook versions, deleted pages, permission changes—and verify that deleted materials actually disappear from the index rather than lingering as orphan vectors. Citation formatting belongs in the generation contract too: if the UI promises page-level provenance, the prompt and post-processor must emit stable source ids, not decorative footnotes that point nowhere.
Re-ranking deserves a short callout even in a beginner path. A bi-encoder shortlist is cheap; a cross-encoder pass over the top fifty candidates often fixes the “almost right document, wrong section” errors that make demos look broken. Pair that with simple refusal rules when similarity scores fall below a calibrated threshold. Teams that skip evaluation usually discover these issues in front of stakeholders instead of in a notebook.
Finally, remember the economic shape of RAG: indexing cost is paid continuously as corpora grow, while fine-tuning cost is paid in bursts. For policy-heavy domains the continuous index bill is usually cheaper than weekly fine-tunes—and it preserves the ability to quote the exact paragraph an auditor asked for. That auditability is often the real product requirement hiding behind “build a chatbot.”
When rolling RAG into an existing support stack, start with one corpus and one question family rather than boiling the ocean. Measure deflection, escalation, and wrong-answer complaints on that slice before adding Confluence spaces wholesale. Narrow wins teach which chunk sizes and hybrid weights transfer; broad launches mostly teach that dashboards without faithfulness metrics hide regressions. Keep a human review queue for contested answers during the first weeks so editors can mark “good retrieval / bad generation” versus “bad retrieval,” which are different repair paths. Over time those labels become training signal for re-rankers and for deciding whether a topic should leave RAG and become a deterministic workflow instead.
Final Takeaway
Store external knowledge
↓
Find relevant information
↓
Give that information to an LLM
↓
Generate a grounded response
Store external knowledge, find the relevant pieces for each question, and only then generate. That three-beat loop is RAG—simple to sketch, demanding to operate well, and still the most practical way to keep assistants honest about private and changing facts.
Operational discipline separates demos from durable assistants: freeze a question set, measure retrieval hit rates beside answer faithfulness, and rebuild indexes on a schedule that matches how often source materials change. When hybrid search, re-ranking, or metadata filters land, keep the same scorecard so improvements are visible rather than anecdotal. Treat access labels as part of the chunk payload from day one; bolting permissions on later is how private HR notes leak into public chats. Finally, prefer refusal plus a citation request over a fluent guess whenever top chunks look weak—users trust calibrated silence more than polished invention.
Keep runbooks for index rebuilds, access reviews, and known failure modes beside the happy-path diagrams so operators inherit more than a slide deck.
Keep runbooks for index rebuilds, access reviews, and known failure modes beside the happy-path diagrams so operators inherit more than a slide deck.
Keep runbooks for index rebuilds, access reviews, and known failure modes beside the happy-path diagrams so operators inherit more than a slide deck.
Keep runbooks for index rebuilds, access reviews, and known failure modes beside the happy-path diagrams so operators inherit more than a slide deck.
Keep runbooks for index rebuilds, access reviews, and known failure modes beside the happy-path diagrams so operators inherit more than a slide deck.
Keep runbooks for index rebuilds, access reviews, and known failure modes beside the happy-path diagrams so operators inherit more than a slide deck.
Keep runbooks for index rebuilds, access reviews, and known failure modes beside the happy-path diagrams so operators inherit more than a slide deck.