This article is published in English.
Practical notes: I Built a RAG System That Audits Itself. Here’s How (With
Operable walkthrough of Practical notes: I Built a RAG System That Audits Itself. Here’s How (With: contracts, checks, and drop-in code slots for teams shipping this pattern.
This walkthrough rebuilds the path from raw materials to a working system for: I Built a RAG System That Audits Itself. Here’s How (With Code).. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Why Nobody Monitors Retrieval (And Why That’s About to Bite Them)
When working through the Why Nobody Monitors Retrieval stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Three Things That Go Wrong Quietly
When working through the Three Things That Go stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
1. Chunk Poisoning
When working through the 1 Chunk Poisoning stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the 1 Chunk Poisoning stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
2. Embedding Drift
The 2 Embedding Drift stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
3. Context Window Waste
The 3 Context Window Waste stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
What We’re Building
The What We re Building stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The What We re Building stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
▣ Check #1: Chunk relevance scoring.
For the Check 1 Chunk relevance stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
▣ Check #2: Embedding drift detection.
For the Check 2 Embedding drift stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
▣ Check #3: Context window efficiency.
For the Check 3 Context window stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Check 3 Context window stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Building the Audit: Start With a Golden Query Set
When working through the Building the Audit Start stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
[
{
"query": "What is the refund policy for digital products?",
"expected_chunk_ids": ["faq_doc_chunk_12", "faq_doc_chunk_13"],
"notes": "Customer FAQ, policy updated 2024-Q1"
},
{
"query": "How do I reset my API key?",
"expected_chunk_ids": ["api_docs_chunk_07"],
"notes": "API documentation, stable"
}
]
A few things that matter when you build it:
When working through the A few things that stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
The Audit Script
When working through the The Audit Script stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the The Audit Script stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Github Repository Structure
The Github Repository Structure stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
rag-retrieval-audit/
├── audit/
│ ├── __init__.py # Package exports
│ ├── rag_audit.py # Main audit logic — all three checks
│ └── config.py # All thresholds and model settings
├── examples/
│ ├── golden_queries.json # Sample golden set (10 queries)
│ └── run_audit.py # End-to-end demo — runs without an existing collection
├── tests/
│ └── test_audit.py # Unit tests for all scoring functions
├── requirements.txt # sentence-transformers, chromadb, scipy, numpy
└── README.md # Setup, usage, how to read the report
# rag_audit.py — structure overview
# Full implementation: https://github.com/satyam671/rag-retrieval-audit
# ── CONFIGURATION (tune to your pipeline)
EMBEDDING_MODEL = "all-MiniLM-L6-v2" # must match your index
TOP_K = 5
RELEVANCE_THRESHOLD = 0.70
DRIFT_P_THRESHOLD = 0.05
EFFICIENCY_THRESHOLD = 0.40
# ── CHECK 1: Are the right chunks coming back?
def score_relevance(query_embedding, chunk_embeddings, expected_ids, retrieved_ids):
"""Cosine similarity per chunk + expected chunk hit/miss against golden set."""
...
# ── CHECK 2: Has retrieval quality shifted over time?
def detect_drift(current_sims, baseline_path=None):
"""Two-sample KS test comparing current similarity distribution to baseline."""
...
# ── CHECK 3: How much of the context window is signal?
def score_efficiency(retrieved_docs, answer):
"""Token overlap between retrieved chunks and the LLM answer."""
...
# ── RUNNER
def run_audit(golden_set_path, collection_name, chroma_persist_dir,
baseline_path=None, answers_path=None, output_path="rag_audit_report.json"):
"""Runs all three checks, writes a structured JSON report, saves the baseline."""
...
Walking Through What Each Check Actually Does
The Walking Through What Each stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
▣ Chunk Relevance Scoring
The Chunk Relevance Scoring stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Chunk Relevance Scoring stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
# WITHOUT the relevance gate — standard approach most teams use
def build_context_naive(query, collection, top_k=5):
results = collection.query(
query_embeddings=[embed(query)],
n_results=top_k,
include=["documents"]
)
# Pass everything back, no quality check
return "\n\n".join(results["documents"][0])
# WITH the relevance gate - what the audit tells you to build
def build_context_gated(query, collection, model, top_k=5, threshold=0.70):
q_emb = model.encode([query], normalize_embeddings=True)[0]
results = collection.query(
query_embeddings=[q_emb.tolist()],
n_results=top_k,
include=["documents", "embeddings", "ids"]
)
passed_chunks = []
for i, chunk_emb in enumerate(results["embeddings"][0]):
sim = cosine_sim(q_emb, np.array(chunk_emb))
if sim >= threshold:
passed_chunks.append(results["documents"][0][i])
# Empty context is better than wrong context
return "\n\n".join(passed_chunks)
▣ Embedding Drift Detection
For the Embedding Drift Detection stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
▣ Context Window Efficiency
For the Context Window Efficiency stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Reading the Report
For the Reading the Report stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Reading the Report stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Expected Output upon running run_audit.py:
When working through the Expected Output upon running stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
python -m examples.run_audit
Retrieval Audit Cheat Sheet
When working through the Retrieval Audit Cheat Sheet stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
One Thing you’d Do Differently
When working through the One Thing you d stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the One Thing you d stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
What This Doesn’t Cover
The What This Doesn t stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
References
The References stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Operational checklist
The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.
Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for ffe9673a9f17: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.