This article is published in English.
Practical notes: RAG, Embeddings, and Vector Databases — Explained From Scratch
Operable walkthrough of Practical notes: RAG, Embeddings, and Vector Databases — Explained From Scratch: contracts, checks, and drop-in code slots for teams shipping this pattern.
Use this as an operator-facing rebuild of the ideas in “RAG, Embeddings, and Vector Databases — Explained From Scratch”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
What is RAG?
For the What is RAG stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
[ Documents ] --> [ Chunk ] --> [ Embed ] --> [ Store in Vector DB ]
|
User Question --> [ Embed Question ] --> [ Find Similar Chunks ] --> [ LLM ] --> Answer
What is an Embedding?
For the What is an Embedding stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
"king" → [0.21, -0.05, 0.87, ..., 0.33] (384 numbers)
"queen" → [0.19, -0.03, 0.85, ..., 0.31] (384 numbers)
"banana" → [-0.72, 0.44, 0.01, ..., -0.56] (384 numbers)
How Text Becomes Numbers
For the How Text Becomes Numbers stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the How Text Becomes Numbers stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Step 1: Tokenization
When working through the Step 1 Tokenization stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
"backend engineer jobs in bangalore"
Tokens: ["backend", "engineer", "jobs", "in", "bang", "##alore"]
Step 2: Transformer Layers
When working through the Step 2 Transformer Layers stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
"backend" → [0.52, -0.18, 0.83, 0.45, ...] (384 numbers)
"engineer" → [0.48, -0.15, 0.79, 0.41, ...] (384 numbers)
"jobs" → [0.31, -0.05, 0.67, 0.22, ...] (384 numbers)
"in" → [0.07, 0.02, 0.11, 0.04, ...] (384 numbers)
"bang" → [0.39, 0.27, 0.55, 0.30, ...] (384 numbers)
"##alore" → [0.14, 0.10, 0.21, 0.11, ...] (384 numbers)
Step 3: Pooling
When working through the Step 3 Pooling stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Step 3 Pooling stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Dimension 1: (0.52 + 0.48 + 0.31 + 0.07 + 0.39 + 0.14) / 6 = 0.318
Dimension 2: (-0.18 + -0.15 + -0.05 + 0.02 + 0.27 + 0.10) / 6 = 0.002
... same for all 384 dimensions ...
Final: [0.318, 0.002, ...] ← ONE vector for the entire sentence
Quick clarification: Vectors vs Dimensions
The Quick clarification Vectors vs stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
2D point (on a map): [x, y] → 1 vector, 2 dimensions
3D point (in a room): [x, y, z] → 1 vector, 3 dimensions
Embedding: [n1, n2, ..n384] → 1 vector, 384 dimensions
What is a Vector Database?
The What is a Vector stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
┌──────────────────────────────────────────────────┐
│ Vector DB │
│ │
│ Entry 1: │
│ text: "Senior Backend Engineer Wipro..." │
│ vector: [0.21, -0.05, 0.87, ...] │
│ │
│ Entry 2: │
│ text: "Frontend dev needed Chennai..." │
│ vector: [0.71, 0.55, -0.30, ...] │
│ │
│ Entry 3: │
│ text: "Backend Engineer Flipkart..." │
│ vector: [0.19, -0.03, 0.85, ...] │
│ │
└──────────────────────────────────────────────────┘
How Many Vectors Get Stored?
The How Many Vectors Get stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The How Many Vectors Get stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Page 1 (naukri.com) → 3000 chars → split into 6 chunks → 6 vectors
Page 2 (linkedin.com) → 5000 chars → split into 10 chunks → 10 vectors
Page 3 (indeed.com) → 2000 chars → split into 4 chunks → 4 vectors
Page 4 (glassdoor.com) → 4500 chars → split into 9 chunks → 9 vectors
Page 5 (some blog) → 1500 chars → split into 3 chunks → 3 vectors
─────────
32 vectors in the database
How Retrieval Works
For the How Retrieval Works stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Step 1: Embed the query with the SAME model
For the Step 1 Embed the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
"backend engineering jobs in Bengaluru" → [0.20, -0.04, 0.86, ...]
Step 2: Compare against every stored vector
For the Step 2 Compare against stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Step 2 Compare against stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Query vector vs Chunk 1 ("Senior Backend Engineer Wipro, Bengaluru...")
→ similarity: 0.95 ✅
Query vector vs Chunk 2 ("Frontend dev needed Chennai...")
→ similarity: 0.23 ❌Query vector vs Chunk 3 ("Backend Engineer Flipkart Bengaluru...")
→ similarity: 0.97 ✅
Step 3: Return the top k chunks
When working through the Step 3 Return the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
The Chunking Problem
When working through the The Chunking Problem stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Original text:
"...looking for a talented software engineer with 5 years of backend experience..."
Chunk 1 (chars 0-500): "...looking for a talented software engi"
Chunk 2 (chars 500-999): "neer with 5 years of backend experience..."
"Home About Careers Login Contact Us Senior Backend Eng"
Better approaches:
When working through the Better approaches stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Better approaches stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
A Complete Walkthrough With Real Numbers
The A Complete Walkthrough With stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
"Home Jobs Companies Senior Backend Engineer Wipro, Bengaluru
Build scalable APIs using Java Spring Boot. 5+ years experience required."
Tokenization (29 tokens):
The Tokenization 29 tokens stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
["[CLS]", "home", "jobs", "companies", "senior", "backend", "engineer",
"wi", "##pro", ",", "bengal", "##uru", "build", "scala", "##ble",
"api", "##s", "using", "java", "spring", "boot", ".", "5", "+",
"years", "experience", "required", ".", "[SEP]"]
Per-token vectors (showing 8 dims instead of 384 for readability):
The Per-token vectors showing 8 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
"[CLS]" → [ 0.01, 0.03, -0.02, 0.05, 0.01, -0.04, 0.02, 0.06]
"home" → [ 0.44, 0.12, -0.33, 0.08, -0.21, 0.55, 0.09, -0.17]
"jobs" → [ 0.31, -0.05, 0.67, 0.22, 0.14, -0.08, 0.43, 0.11]
"companies" → [ 0.28, 0.09, 0.51, 0.18, 0.07, -0.12, 0.38, 0.05]
"senior" → [ 0.15, -0.22, 0.71, 0.33, 0.41, 0.09, 0.55, 0.27]
"backend" → [ 0.52, -0.18, 0.83, 0.45, 0.38, 0.21, 0.61, 0.34]
"engineer" → [ 0.48, -0.15, 0.79, 0.41, 0.35, 0.18, 0.58, 0.31]
"wi" → [ 0.11, 0.04, 0.22, 0.09, 0.03, 0.07, 0.14, 0.02]
"##pro" → [ 0.08, 0.02, 0.19, 0.06, 0.01, 0.05, 0.11, -0.01]
"," → [ 0.00, 0.01, 0.00, 0.01, -0.01, 0.00, 0.01, 0.00]
"bengal" → [ 0.39, 0.27, 0.55, 0.30, 0.22, 0.33, 0.41, 0.19]
"##uru" → [ 0.14, 0.10, 0.21, 0.11, 0.08, 0.12, 0.15, 0.07]
"build" → [ 0.35, -0.11, 0.62, 0.28, 0.19, 0.15, 0.47, 0.22]
"scala" → [ 0.29, -0.09, 0.54, 0.24, 0.16, 0.11, 0.40, 0.18]
"##ble" → [ 0.10, 0.03, 0.18, 0.07, 0.05, 0.04, 0.13, 0.06]
"api" → [ 0.41, -0.14, 0.73, 0.37, 0.30, 0.19, 0.53, 0.28]
"##s" → [ 0.03, 0.01, 0.05, 0.02, 0.01, 0.01, 0.04, 0.01]
"using" → [ 0.07, 0.02, 0.11, 0.04, 0.03, 0.02, 0.08, 0.03]
"java" → [ 0.46, -0.20, 0.77, 0.39, 0.32, 0.17, 0.56, 0.29]
"spring" → [ 0.42, -0.16, 0.70, 0.35, 0.28, 0.14, 0.51, 0.25]
"boot" → [ 0.38, -0.13, 0.65, 0.31, 0.25, 0.12, 0.48, 0.23]
"." → [ 0.00, 0.01, 0.00, 0.01, -0.01, 0.00, 0.01, 0.00]
"5" → [ 0.05, 0.01, 0.09, 0.03, 0.02, 0.01, 0.06, 0.02]
"+" → [ 0.01, 0.00, 0.02, 0.01, 0.00, 0.00, 0.01, 0.00]
"years" → [ 0.20, -0.07, 0.44, 0.19, 0.13, 0.08, 0.33, 0.14]
"experience" → [ 0.33, -0.10, 0.61, 0.27, 0.20, 0.13, 0.45, 0.21]
"required" → [ 0.18, -0.06, 0.39, 0.16, 0.11, 0.07, 0.29, 0.12]
"." → [ 0.00, 0.01, 0.00, 0.01, -0.01, 0.00, 0.01, 0.00]
"[SEP]" → [ 0.02, 0.01, -0.01, 0.03, 0.00, -0.02, 0.01, 0.04]
The Per-token vectors showing 8 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Pooling — 29 vectors become 1:
For the Pooling 29 vectors become stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Dim 1: (0.01 + 0.44 + 0.31 + 0.28 + 0.15 + 0.52 + 0.48 + 0.11 + 0.08 + 0.00
+ 0.39 + 0.14 + 0.35 + 0.29 + 0.10 + 0.41 + 0.03 + 0.07 + 0.46 + 0.42
+ 0.38 + 0.00 + 0.05 + 0.01 + 0.20 + 0.33 + 0.18 + 0.00 + 0.02) / 29
= 0.231
Dim 2: (0.03 + 0.12 + -0.05 + 0.09 + -0.22 + -0.18 + -0.15 + 0.04 + 0.02 + 0.01
+ 0.27 + 0.10 + -0.11 + -0.09 + 0.03 + -0.14 + 0.01 + 0.02 + -0.20 + -0.16
+ -0.13 + 0.01 + 0.01 + 0.00 + -0.07 + -0.10 + -0.06 + 0.01 + 0.01) / 29
= -0.030... same for all 384 dimensions ...
Stored in the vector database:
For the Stored in the vector stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
text: "Home Jobs Companies Senior Backend Engineer Wipro, Bengaluru..."
vector: [0.231, -0.030, 0.421, 0.189, ...]
Now a user searches: “senior backend engineer at Wipro”
For the Now a user searches stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Now a user searches stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Query vector: [0.228, -0.028, 0.418, 0.185, ...]
Chunk vector: [0.231, -0.030, 0.421, 0.189, ...]
Two Models, Two Jobs
When working through the Two Models Two Jobs stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
Why RAG Matters
When working through the Why RAG Matters stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
What Makes a Good Embedding Model?
When working through the What Makes a Good stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the What Makes a Good stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
A: "Senior Backend Engineer at Wipro"
B: "Software Developer role in Bengaluru"
C: "Best chocolate cake recipe"
Operational checklist
When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.
Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.