This article is published in English.
Practical notes: How to Implement RAG (Retrieval-Augmented Generation) in Your
Operable walkthrough of Practical notes: How to Implement RAG (Retrieval-Augmented Generation) in Your: contracts, checks, and drop-in code slots for teams shipping this pattern.
Use this as an operator-facing rebuild of the ideas in “How to Implement RAG (Retrieval-Augmented Generation) in Your Web Application”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Understanding the Architecture Behind RAG
For the Understanding the Architecture Behind stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
User Question
|
v
Generate Query Embedding
|
v
Metadata Filtering
|
v
Vector Similarity Search
|
v
Top-K Relevant Documents
|
v
Context Construction
|
v
LLM / Gemini
|
v
Grounded Response + Sources
Setting Up the Embedding Infrastructure
For the Setting Up the Embedding stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
const { PredictionServiceClient, helpers } =
require('@google-cloud/aiplatform');
const PROJECT_ID = process.env.PROJECT_ID;
const client = new PredictionServiceClient({
apiEndpoint: 'aiplatform.googleapis.com'
});
async function generateEmbedding(
text,
taskType = 'RETRIEVAL_DOCUMENT'
) {
const endpoint =
`projects/${PROJECT_ID}/locations/global/` +
`publishers/google/models/gemini-embedding-001`;
const instance = {
content: text,
task_type: taskType
};
const request = {
endpoint,
instances: [helpers.toValue(instance)]
};
const [response] = await client.predict(request);
return response.predictions[0].embeddings.values;
}
Why Batching Matters When Generating Embeddings
For the Why Batching Matters When stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Why Batching Matters When stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
const EMBEDDING_CONFIG = {
maxSegmentsPerRequest: 100,
maxTokensPerRequest: 18000,
concurrency: 3,
tokenEstimateDivisor: 3
};
function estimateTokens(text) {
return Math.ceil(
text.length / EMBEDDING_CONFIG.tokenEstimateDivisor
);
}
function packIntoBatches(texts) {
const batches = [];
let currentBatch = [];
let currentTokens = 0;
for (const text of texts) {
const tokens = estimateTokens(text);
const exceedsCount =
currentBatch.length >=
EMBEDDING_CONFIG.maxSegmentsPerRequest;
const exceedsTokens =
currentTokens + tokens >
EMBEDDING_CONFIG.maxTokensPerRequest;
if (exceedsCount || exceedsTokens) {
if (currentBatch.length > 0) {
batches.push(currentBatch);
}
currentBatch = [text];
currentTokens = tokens;
} else {
currentBatch.push(text);
currentTokens += tokens;
}
}
if (currentBatch.length > 0) {
batches.push(currentBatch);
}
return batches;
}
Using Firebase Firestore for Vector Search
When working through the Using Firebase Firestore for stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
const { Firestore } = require('@google-cloud/firestore');
const firestore = new Firestore();
async function storeDocumentWithEmbedding(
collectionPath,
docId,
text,
embedding,
metadata
) {
const docRef =
firestore.doc(`${collectionPath}/${docId}`);
await docRef.set({
text,
embedding,
...metadata,
createdAt: Firestore.FieldValue.serverTimestamp()
});
}
async function findSimilarDocuments(
collectionPath,
queryEmbedding,
limit = 5
) {
const collectionRef =
firestore.collection(collectionPath);
const vectorQuery = collectionRef.findNearest({
vectorField: 'embedding',
queryVector: queryEmbedding,
limit,
distanceMeasure: 'DOT_PRODUCT'
});
const snapshot = await vectorQuery.get();
return snapshot.docs.map(doc => ({
id: doc.id,
data: doc.data()
}));
}
Combining Metadata Filtering With Semantic Retrieval
When working through the Combining Metadata Filtering With stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
async function retrieveContext(
queryText,
selectedTopics = [],
limit = 5
) {
const queryEmbedding =
await generateEmbedding(
queryText,
'RETRIEVAL_QUERY'
);
let collectionRef =
firestore.collection('knowledge_base');
if (selectedTopics.length > 0) {
collectionRef = collectionRef.where(
'topics',
'array-contains-any',
selectedTopics
);
}
const vectorQuery =
collectionRef.findNearest({
vectorField: 'embedding',
queryVector: queryEmbedding,
limit,
distanceMeasure: 'DOT_PRODUCT'
});
const snapshot = await vectorQuery.get();
return snapshot.docs.map(doc => {
const data = doc.data();
return {
text: data.text,
source: data.source,
metadata: data.metadata || {}
};
});
}
Turning Retrieved Documents Into Model Context
When working through the Turning Retrieved Documents Into stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the Turning Retrieved Documents Into stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
async function generateRAGResponse(
userQuery,
contextDocuments
) {
const contextSection =
contextDocuments
.map((doc, index) => {
return `
### Reference ${index + 1}
${doc.text}
Source: ${doc.metadata?.source || 'Unknown'}
`;
})
.join('\n');
const prompt = `
You are an AI assistant with access
to a knowledge base.
Use the provided context to answer
the user's question accurately.
CONTEXT:
${contextSection}
USER QUESTION:
${userQuery}
INSTRUCTIONS:
1. Answer using the provided context.
2. If the context is insufficient, say so clearly.
3. Cite the references used.
4. Do not invent information.
ANSWER:
`;
const result =
await genAI.models.generateContent({
model: 'gemini-2.5-flash-lite',
contents: prompt
});
return result.text.trim();
}
Optimizing RAG With Caching and Retry Logic
The Optimizing RAG With Caching stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
const embeddingCache = new Map();
async function getCachedEmbedding(text, taskType) {
const cacheKey = `${taskType}:${text}`;
if (embeddingCache.has(cacheKey)) {
return embeddingCache.get(cacheKey);
}
const embedding =
await generateEmbedding(text, taskType);
embeddingCache.set(cacheKey, embedding);
return embedding;
}
async function retryWithBackoff(
fn,
maxRetries = 3
) {
for (let attempt = 0; attempt < maxRetries; attempt++) {
try {
return await fn();
} catch (error) {
if (
error.code === 429 ||
error.message.includes('rate limit')
) {
const delay =
Math.pow(2, attempt) * 1000;
await new Promise(resolve =>
setTimeout(resolve, delay)
);
continue;
}
throw error;
}
}
throw new Error('Max retries exceeded');
}
Retrieving Multiple Types of Context
The Retrieving Multiple Types of stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
async function retrieveMultiContext(
queryText,
options = {}
) {
const {
includeDefinitions = true,
includeExamples = true,
includeHistorical = false,
topics = []
} = options;
const queryEmbedding =
await generateEmbedding(
queryText,
'RETRIEVAL_QUERY'
);
const contextPromises = [];
if (includeDefinitions) {
contextPromises.push(
findSimilarDocuments(
'definitions',
queryEmbedding,
3
).then(documents => ({
type: 'definitions',
documents
}))
);
}
if (includeExamples) {
contextPromises.push(
findSimilarDocuments(
'examples',
queryEmbedding,
5
).then(documents => ({
type: 'examples',
documents
}))
);
}
const contexts =
await Promise.all(contextPromises);
return contexts;
}
Monitoring RAG Instead of Guessing About Quality
The Monitoring RAG Instead of stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Monitoring RAG Instead of stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
class RAGMetrics {
constructor() {
this.metrics = {
totalQueries: 0,
averageLatency: 0,
retrievalAccuracy: [],
errors: []
};
}
logQuery(
query,
contextCount,
latency,
sources
) {
this.metrics.totalQueries++;
const previousLatency =
this.metrics.averageLatency *
(this.metrics.totalQueries - 1);
this.metrics.averageLatency =
(previousLatency + latency) /
this.metrics.totalQueries;
console.log({
query: query.substring(0, 100),
contextCount,
latency,
sourceCount: sources.length,
timestamp: new Date().toISOString()
});
}
logRetrievalAccuracy(
retrievedDocs,
relevantDocs
) {
const retrievedIds =
new Set(retrievedDocs.map(d => d.id));
const relevantIds =
new Set(relevantDocs.map(d => d.id));
const intersection =
new Set(
[...retrievedIds]
.filter(id => relevantIds.has(id))
);
const precision =
intersection.size / retrievedIds.size;
const recall =
intersection.size / relevantIds.size;
this.metrics.retrievalAccuracy.push({
precision,
recall
});
}
}
The Real Production RAG Pipeline
For the The Real Production RAG stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
DOCUMENT INGESTION
|
v
Clean + Split Documents
|
v
Generate Embeddings
|
v
Firestore + Metadata Storage
|
|
USER QUERY --------+
|
v
Query Embedding
|
v
Authorization + Filters
|
v
Vector Similarity Search
|
v
Relevant Context
|
v
Prompt Construction
|
v
Gemini / LLM
|
v
Answer + Sources + Metrics
What you Would Focus on as a Senior Engineer
For the What you Would Focus stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Conclusion: Building a Production-Ready RAG System
For the Conclusion Building a Production-Ready stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Conclusion Building a Production-Ready stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Operational checklist
When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.
Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Freeze a golden set before changing prompts or models. Moving both the system and the yardstick hides regressions.
Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 9607363b4f86: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.