Home / Articles / Practical notes: Build a RAG System From Scratch — Hands-On, No API Costs

This article is published in English.

Practical notes: Build a RAG System From Scratch — Hands-On, No API Costs

Operable walkthrough of Practical notes: Build a RAG System From Scratch — Hands-On, No API Costs: contracts, checks, and drop-in code slots for teams shipping this pattern.

2697 words

The following notes reconstruct a practical path around “Build a RAG System From Scratch — Hands-On, No API Costs”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

What RAG actually is (60 seconds)

The What RAG actually is stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

question ──► [embed] ──► [search your docs] ──► top chunks ──┐
                                                             ▼
                                          [LLM: "answer using this context"] ──► answer

Step 0 — Setup

The Step 0 Setup stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

pip install sentence-transformers transformers torch numpy

Step 1 — A knowledge base the model has never seen

The Step 1 A knowledge stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices. The Step 1 A knowledge stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

# rag.py
DOCUMENTS = [
    """Nimbus is a fictional note-taking app launched in 2023. The free plan,
    called Nimbus Lite, allows up to 50 notes and 1 GB of storage. There are no
    collaboration features on the free plan.""",
    """Nimbus Pro costs 8 dollars per month billed annually, or 10 dollars billed
    monthly. Pro removes the note limit, gives 50 GB of storage, and unlocks
    real-time collaboration with up to 5 people per note.""",    """Nimbus stores all notes encrypted at rest using AES-256. End-to-end
    encryption is only available on the Pro plan and must be enabled manually in
    Settings > Security. Once enabled it cannot be turned off for that note.""",    """The Nimbus mobile app supports offline editing. Changes made offline are
    queued and sync automatically the next time the device is online. If two
    devices edit the same note offline, Nimbus keeps both versions and flags a
    conflict for the user to resolve.""",    """Nimbus offers a 30-day refund policy on all paid plans, no questions asked.
    Refunds are processed to the original payment method within 5 business days.
    Annual plans cancelled after 30 days are not refundable but stay active until
    the end of the billing period.""",    """Nimbus support is available via email at help@nimbus.example and live chat.
    Live chat is only staffed for Pro customers, Monday to Friday, 9am to 6pm UTC.
    Free-plan users receive email support with a typical 48-hour response time.""",
]
from transformers import pipeline
gen = pipeline("text2text-generation", model="google/flan-t5-base")
print(gen("How much does Nimbus Pro cost?", max_new_tokens=50)[0]["generated_text"])

Step 2 — Chunking

For the Step 2 Chunking stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

def chunk_text(text, chunk_size=60, overlap=15):
    """Split text into overlapping chunks of `chunk_size` words."""
    words = text.split()
    chunks = []
    start = 0
    while start < len(words):
        end = start + chunk_size
        chunks.append(" ".join(words[start:end]))
        if end >= len(words):
            break
        start = end - overlap   # step back by `overlap` so context isn't cut
    return chunks
# Build our chunk list, remembering which doc each chunk came from
chunks = []
for doc_id, doc in enumerate(DOCUMENTS):
    for c in chunk_text(doc):
        chunks.append({"doc_id": doc_id, "text": c})print(f"{len(DOCUMENTS)} documents -> {len(chunks)} chunks")
for c in chunks[:3]:
    print("-", c["text"][:70], "...")

Step 3 — Embeddings: turning text into vectors

For the Step 3 Embeddings turning stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

from sentence_transformers import SentenceTransformer
embedder = SentenceTransformer("all-MiniLM-L6-v2")# Embed every chunk. normalize_embeddings=True makes the vectors unit-length,
# which lets us measure similarity with a simple dot product later.
chunk_texts = [c["text"] for c in chunks]
chunk_vectors = embedder.encode(chunk_texts, normalize_embeddings=True)print("vector shape:", chunk_vectors.shape)   # (num_chunks, 384)
import numpy as np
pairs = embedder.encode(
    ["the price of the pro plan", "how much does it cost", "the weather in Paris"],
    normalize_embeddings=True,
)
print("price vs cost :", round(float(pairs[0] @ pairs[1]), 3))   # should be HIGH
print("price vs weather:", round(float(pairs[0] @ pairs[2]), 3)) # should be LOW

Step 4 — Retrieval: find the chunks that answer a question

For the Step 4 Retrieval find stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Step 4 Retrieval find stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

import numpy as np
def retrieve(question, k=3):
    q_vec = embedder.encode([question], normalize_embeddings=True)[0]
    scores = chunk_vectors @ q_vec              # cosine similarity to every chunk
    top_idx = np.argsort(scores)[::-1][:k]      # indices of the k highest scores
    return [(chunks[i]["text"], float(scores[i])) for i in top_idx]for text, score in retrieve("How much does Nimbus Pro cost?"):
    print(f"[{score:.3f}] {text[:80]}...")

Step 5 — Generation: let the model answer from the context

When working through the Step 5 Generation let stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

from transformers import pipeline
generator = pipeline("text2text-generation", model="google/flan-t5-base")def rag_answer(question, k=3):
    retrieved = retrieve(question, k=k)
    context = "\n".join(text for text, _ in retrieved)    prompt = f"""Answer the question using only the context below.
If the answer is not in the context, say you don't know.Context:
{context}Question: {question}
Answer:"""    out = generator(prompt, max_new_tokens=80)[0]["generated_text"]
    return out.strip(), retrievedanswer, sources = rag_answer("How much does Nimbus Pro cost?")
print("ANSWER:", answer)
print("\nBased on:")
for text, score in sources:
    print(f"  [{score:.3f}] {text[:70]}...")

Step 6 — Put it all together

When working through the Step 6 Put it stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

# rag.py — a complete, local, no-API RAG system
import numpy as np
from sentence_transformers import SentenceTransformer
from transformers import pipeline
DOCUMENTS = [
    """Nimbus is a fictional note-taking app launched in 2023. The free plan,
    called Nimbus Lite, allows up to 50 notes and 1 GB of storage. There are no
    collaboration features on the free plan.""",
    """Nimbus Pro costs 8 dollars per month billed annually, or 10 dollars billed
    monthly. Pro removes the note limit, gives 50 GB of storage, and unlocks
    real-time collaboration with up to 5 people per note.""",
    """Nimbus stores all notes encrypted at rest using AES-256. End-to-end
    encryption is only available on the Pro plan and must be enabled manually in
    Settings > Security. Once enabled it cannot be turned off for that note.""",
    """The Nimbus mobile app supports offline editing. Changes made offline are
    queued and sync automatically the next time the device is online. If two
    devices edit the same note offline, Nimbus keeps both versions and flags a
    conflict for the user to resolve.""",
    """Nimbus offers a 30-day refund policy on all paid plans, no questions asked.
    Refunds are processed to the original payment method within 5 business days.
    Annual plans cancelled after 30 days are not refundable but stay active until
    the end of the billing period.""",
    """Nimbus support is available via email at help@nimbus.example and live chat.
    Live chat is only staffed for Pro customers, Monday to Friday, 9am to 6pm UTC.
    Free-plan users receive email support with a typical 48-hour response time.""",
]def chunk_text(text, chunk_size=60, overlap=15):
    words = text.split()
    chunks, start = [], 0
    while start < len(words):
        end = start + chunk_size
        chunks.append(" ".join(words[start:end]))
        if end >= len(words):
            break
        start = end - overlap
    return chunksprint("Loading models (first run downloads them)...")
embedder = SentenceTransformer("all-MiniLM-L6-v2")
generator = pipeline("text2text-generation", model="google/flan-t5-base")# Index the documents once at startup
chunks = []
for doc_id, doc in enumerate(DOCUMENTS):
    for c in chunk_text(doc):
        chunks.append({"doc_id": doc_id, "text": c})
chunk_vectors = embedder.encode(
    [c["text"] for c in chunks], normalize_embeddings=True
)def retrieve(question, k=3):
    q_vec = embedder.encode([question], normalize_embeddings=True)[0]
    scores = chunk_vectors @ q_vec
    top_idx = np.argsort(scores)[::-1][:k]
    return [(chunks[i]["text"], float(scores[i])) for i in top_idx]def rag_answer(question, k=3):
    retrieved = retrieve(question, k=k)
    context = "\n".join(text for text, _ in retrieved)
    prompt = (
        "Answer the question using only the context below. "
        "If the answer is not in the context, say you don't know.\n\n"
        f"Context:\n{context}\n\nQuestion: {question}\nAnswer:"
    )
    out = generator(prompt, max_new_tokens=80)[0]["generated_text"]
    return out.strip()if __name__ == "__main__":
    print("RAG ready. Ask about Nimbus (or type 'quit').\n")
    while True:
        q = input("You: ").strip()
        if q.lower() in {"quit", "exit", ""}:
            break
        print("Nimbus bot:", rag_answer(q), "\n")
python rag.py

Step 7 — Prove RAG is doing the work (A/B test)

When working through the Step 7 Prove RAG stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Step 7 Prove RAG stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

def no_rag(question):
    out = generator(f"Question: {question}\nAnswer:", max_new_tokens=80)
    return out[0]["generated_text"].strip()
q = "Can free-plan Nimbus users use live chat support?"
print("WITHOUT context:", no_rag(q))
print("WITH context   :", rag_answer(q))

Step 8 — Make it better (pick what interests you)

The Step 8 Make it stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

The mental model to keep

The The mental model to stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

Troubleshooting

The Troubleshooting stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Troubleshooting stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Operational checklist

For the Operational checklist stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.

Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.

Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for 5223acdafa84: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.