This article is published in English.
Practical notes: Building a Production-Grade Local RAG Pipeline — 100% Free, No
Operable walkthrough of Practical notes: Building a Production-Grade Local RAG Pipeline — 100% Free, No: contracts, checks, and drop-in code slots for teams shipping this pattern.
This walkthrough rebuilds the path from raw materials to a working system for: Building a Production-Grade Local RAG Pipeline — 100% Free, No Cloud Required. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Why This Matters
When working through the Why This Matters stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
The Complete Tech Stack
When working through the The Complete Tech Stack stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Prerequisites
When working through the Prerequisites stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Prerequisites stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Part 1 — Environment Setup
The Part 1 Environment Setup stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Step 1: Create the project
The Step 1 Create the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
mkdir local-rag
cd local-rag
uv init
Step 2: Create and activate the virtual environment
The Step 2 Create and stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
uv venv
.venv\Scripts\activate
The Step 2 Create and stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Step 3: Install all dependencies
For the Step 3 Install all stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
uv add google-genai pypdf chromadb rich python-dotenv huggingface_hub fpdf2
Step 4: Create the project structure
For the Step 4 Create the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
mkdir pdfs
mkdir pdfs\versions
type nul > local_rag.ipynb
type nul > .env
type nul > .gitignore
local-rag/
├── .venv/ ← virtual environment (never commit)
├── pdfs/ ← drop your PDFs here
│ └── versions/ ← test PDFs for CDC testing
├── chroma_db/ ← auto-created on first ingest
├── memory_checkpoints/ ← auto-created on first memory session
├── staleness_registry.json ← auto-created
├── chunk_registry.json ← auto-created
├── local_rag.ipynb ← your notebook
├── .env ← API keys (never commit)
├── .gitignore
└── pyproject.toml
Step 5: Configure API keys
For the Step 5 Configure API stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Step 5 Configure API stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
GEMINI_API_KEY=your_gemini_key_here
HF_API_KEY=your_huggingface_token_here
.env
chroma_db/
memory_checkpoints/
staleness_registry.json
chunk_registry.json
__pycache__/
.venv/
*.pyc
Step 6: Set up VS Code
When working through the Step 6 Set up stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Ctrl+Shift+P → Python: Select Interpreter → .venv\Scripts\python.exe
Part 2 — Core Pipeline Walkthrough
When working through the Part 2 Core Pipeline stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Cell 1 — Dependencies (inside notebook)
When working through the Cell 1 Dependencies inside stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
# Run once inside the notebook if uv add was not used externally
# %pip install google-genai pypdf chromadb rich python-dotenv huggingface_hub fpdf2
When working through the Cell 1 Dependencies inside stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Cell 2 — Configuration
The Cell 2 Configuration stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
import os
from dotenv import load_dotenv
load_dotenv()
GEMINI_API_KEY = os.environ.get("GEMINI_API_KEY", "")
EMBED_DIM = 768
HF_API_KEY = os.environ.get("HF_API_KEY", "")
GEMINI_EMBED_MODEL = "gemini-embedding-001"
HF_LLM_MODEL = "openai/gpt-oss-20b:groq"
CHROMA_DB_PATH = "./chroma_db"
COLLECTION_NAME = "local_rag"
CHUNK_SIZE = 800
CHUNK_OVERLAP = 120
TOP_K = 5
EMBED_BATCH_SIZE = 50
BATCH_SLEEP_SEC = 0.3
LLM_MAX_NEW_TOKENS = 1024
LLM_TEMPERATURE = 0.1
assert GEMINI_API_KEY, "❌ GEMINI_API_KEY not set"
assert HF_API_KEY, "❌ HF_API_KEY not set"
Cell 3 — PDF Text Extraction
The Cell 3 PDF Text stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
from pypdf import PdfReader
def extract_text_from_pdf(pdf_path: str) -> tuple[str, int]:
reader = PdfReader(pdf_path)
pages = []
for i, page in enumerate(reader.pages):
text = page.extract_text()
if text and text.strip():
pages.append(f"[Page {i + 1}]\n{text.strip()}")
return "\n\n".join(pages), len(reader.pages)
Cell 4 — Sliding Window Chunking
The Cell 4 Sliding Window stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
def chunk_text(text: str) -> list[str]:
chunks, start = [], 0
while start < len(text):
end = start + CHUNK_SIZE
chunk = text[start:end].strip()
if len(chunk) >= 80:
chunks.append(chunk)
start += CHUNK_SIZE - CHUNK_OVERLAP
return chunks
The Cell 4 Sliding Window stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Cell 5 — Gemini Embeddings
For the Cell 5 Gemini Embeddings stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.
from google import genai
from google.genai import types
genai_client = genai.Client(api_key=GEMINI_API_KEY)
def embed_documents_batch(chunks: list[str]) -> list[list[float]]:
all_embeddings = []
for i, chunk in enumerate(chunks, 1):
result = genai_client.models.embed_content(
model=GEMINI_EMBED_MODEL,
contents=chunk,
config=types.EmbedContentConfig(
task_type="RETRIEVAL_DOCUMENT",
output_dimensionality=EMBED_DIM,
),
)
all_embeddings.append(result.embeddings[0].values)
return all_embeddings
def embed_query(text: str) -> list[float]:
result = genai_client.models.embed_content(
model=GEMINI_EMBED_MODEL,
contents=text,
config=types.EmbedContentConfig(
task_type="RETRIEVAL_QUERY",
output_dimensionality=EMBED_DIM,
),
)
return result.embeddings[0].values
Cell 6 — ChromaDB Storage and Retrieval
For the Cell 6 ChromaDB Storage stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
import uuid, chromadb
def get_collection():
client = chromadb.PersistentClient(path=CHROMA_DB_PATH)
return client.get_or_create_collection(
name=COLLECTION_NAME,
metadata={"hnsw:space": "cosine"},
)
def store_in_chroma(chunks, embeddings, doc_name):
collection = get_collection()
ids = [str(uuid.uuid4()) for _ in chunks]
metadatas = [{"source": doc_name, "chunk_index": i}
for i in range(len(chunks))]
collection.add(ids=ids, embeddings=embeddings,
documents=chunks, metadatas=metadatas)
return len(chunks)
def retrieve_context(query: str) -> list[dict]:
collection = get_collection()
query_embedding = embed_query(query)
results = collection.query(
query_embeddings=[query_embedding],
n_results=TOP_K,
include=["documents", "metadatas", "distances"],
)
chunks = []
for doc, meta, dist in zip(results["documents"][0],
results["metadatas"][0],
results["distances"][0]):
chunks.append({
"text": doc,
"source": meta.get("source", "unknown"),
"chunk_index": meta.get("chunk_index", -1),
"score": round(1 - dist, 4),
})
return sorted(chunks, key=lambda x: x["score"], reverse=True)
Cell 7 — Running Ingestion
For the Cell 7 Running Ingestion stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
PDF_PATH = "./pdfs/attention.pdf"
raw_text, page_count = extract_text_from_pdf(PDF_PATH)
chunks = chunk_text(raw_text)
embeddings = embed_documents_batch(chunks)
stored = store_in_chroma(chunks, embeddings,
os.path.basename(PDF_PATH))
For the Cell 7 Running Ingestion stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Cell 8 — LLM: gpt-oss-20b via HuggingFace
When working through the Cell 8 LLM gpt-oss-20b stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
from huggingface_hub import InferenceClient
hf_client = InferenceClient(api_key=HF_API_KEY)
def build_messages(query: str, context_chunks: list[dict]) -> list[dict]:
context_str = "\n\n---\n\n".join([
f"[Source: {c['source']} | Chunk #{c['chunk_index']} | "
f"Relevance: {c['score']}]\n{c['text']}"
for c in context_chunks
])
return [
{"role": "system", "content": SYSTEM_MSG},
{"role": "user", "content":
f"CONTEXT:\n{context_str}\n\nQUESTION:\n{query}"},
]
def generate_answer(messages: list[dict]) -> str:
completion = hf_client.chat.completions.create(
model=HF_LLM_MODEL,
messages=messages,
max_tokens=LLM_MAX_NEW_TOKENS,
temperature=LLM_TEMPERATURE,
)
return completion.choices[0].message.content.strip()
Cell 9–11 — The ask() Pipeline and REPL
When working through the Cell 9 11 The stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
query → embed_query() → ChromaDB cosine search → top-5 chunks
→ build_messages() → generate_answer() → printed answer
1. attention.pdf chunk #34 [██████████████████████░░░░░░░░] 0.7335
Part 3 — Conversation Memory
When working through the Part 3 Conversation Memory stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
The problem
When working through the The problem stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
The solution: ConversationMemory
When working through the The solution ConversationMemory stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
class ConversationMemory:
def __init__(self, session_id: str = None):
self.session_id = session_id or datetime.now().strftime("%Y%m%d_%H%M%S")
self.filepath = os.path.join(MEMORY_DIR, f"{self.session_id}.json")
self.history = []
# Auto-loads if resuming an existing session
if os.path.exists(self.filepath):
self._load()
def add_turn(self, question: str, answer: str, chunks: list[dict]):
# Append user + assistant turns, checkpoint immediately
...
self._save()
def get_messages_with_history(self, query, context_chunks, system_msg):
# Injects last 6 Q&A pairs into the message list before the current turn
...
# New session
memory = ConversationMemory()
# Resume yesterday's session
memory = ConversationMemory("20260413_104959")
Turn 4 question: "How does that compare to what you said about the BLEU score?"
Turn 4 answer: "The context also reports a BLEU score of 28.4 for the
Transformer (big) on WMT 2014 English-to-German. This matches
exactly what I previously stated."
Part 4 — Staleness Tracking, CDC, and Recency Weighting
When working through the Part 4 Staleness Tracking stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
The real-world problem
When working through the The real-world problem stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Staleness Tracking (Cell 14A)
When working through the Staleness Tracking Cell 14A stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Staleness Tracking Cell 14A stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
def compute_file_hash(pdf_path: str) -> str:
sha = hashlib.sha256()
with open(pdf_path, "rb") as f:
for block in iter(lambda: f.read(65536), b""):
sha.update(block)
return sha.hexdigest()
CDC Engine (Cell 14B)
The CDC Engine Cell 14B stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
def compute_chunk_hash(text: str) -> str:
return hashlib.md5(text.encode("utf-8")).hexdigest()
def diff_chunks(old_registry: dict, new_chunks: list[str]) -> dict:
new_hash_map = {compute_chunk_hash(c): c for c in new_chunks}
old_hashes = set(old_registry.keys())
new_hashes = set(new_hash_map.keys())
return {
"added": {h: new_hash_map[h] for h in (new_hashes - old_hashes)},
"removed": {h: old_registry[h] for h in (old_hashes - new_hashes)},
"unchanged": {h: old_registry[h] for h in (old_hashes & new_hashes)},
}
v1 → v2 CDC result:
✅ Unchanged : 8 (kept — zero re-embedding cost)
➕ Added : 6 (embedded + inserted)
➖ Removed : 4 (deleted from ChromaDB)
💰 API calls saved: 8/14 (57% reuse)
v2 → v3 CDC result:
✅ Unchanged : 10 (kept - zero re-embedding cost)
➕ Added : 4 (embedded + inserted)
➖ Removed : 2 (deleted from ChromaDB)
💰 API calls saved: 10/14 (71% reuse)
Recency-Weighted Retrieval (Cell 14C)
The Recency-Weighted Retrieval Cell 14C stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
def recency_decay(ingested_at_str: str, half_life_days: float = 30) -> float:
ingested = datetime.fromisoformat(ingested_at_str)
days_gone = (datetime.now() - ingested).total_seconds() / 86400
λ = math.log(2) / half_life_days
return round(math.exp(-λ * days_gone), 4)
blended = alpha * cosine_score + (1 - alpha) * recency_score
# Default: 0.85 * cosine + 0.15 * recency
Part 5 — Testing with Multi-Version PDFs
The Part 5 Testing with stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Part 5 Testing with stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
TEST 1: Full ingest of v1
TEST 2: Staleness check — same file, correctly skipped
TEST 3: Baseline queries against v1
TEST 4: Copy v2 over active file → CDC kicks in
TEST 5: Same queries now return v2 content, newer chunks visible in recency scores
TEST 6: Copy v3 over active file → second CDC cycle
TEST 7: Recency verification — v3 chunks score highest across the board
TEST 8: Full stack test — weighted retrieval + conversation memory combined
Results and Verified Answers
For the Results and Verified Answers stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Efficiency Summary
For the Efficiency Summary stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Known Limitations
For the Known Limitations stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Known Limitations stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
What You Have Built
When working through the What You Have Built stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
PDF on disk
└► SHA256 hash check (staleness)
├► Unchanged → skip
└► Changed → CDC diff
├► Unchanged chunks → kept in ChromaDB (zero API cost)
├► Removed chunks → deleted from ChromaDB
└► Added chunks → embed (Gemini) → store (ChromaDB)
↓
User question
└► embed_query() [RETRIEVAL_QUERY task type]
└► ChromaDB cosine search (TOP_K × 3 candidates)
└► recency_decay() per chunk
└► blended score re-ranking
└► top-5 chunks as context
└► ConversationMemory.get_messages_with_history()
└► gpt-oss-20b via Groq/HuggingFace
└► grounded answer + checkpoint to disk
Operational checklist
The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 8d172e929623: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.