Home / Articles / Practical notes: Inside Vector Databases for RAG: From Chunk Storage to HNSW &

This article is published in English.

Practical notes: Inside Vector Databases for RAG: From Chunk Storage to HNSW &

Operable walkthrough of Practical notes: Inside Vector Databases for RAG: From Chunk Storage to HNSW &: contracts, checks, and drop-in code slots for teams shipping this pattern.

2042 words

Use this as an operator-facing rebuild of the ideas in “Inside Vector Databases for RAG: From Chunk Storage to HNSW & IVF Search”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

The Standard Approach (Used in Most Systems)

For the The Standard Approach Used stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

{
        "embedding": [0.123, 0.456, ...],
        "text": "Transformer models are powerful...",
        "metadata": {
        "doc_id": "doc1",
        "page": 5
        }
}

Why Store Chunk + Embedding Together?

For the Why Store Chunk Embedding stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

How Retrieval Works

For the How Retrieval Works stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the How Retrieval Works stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

How Vector Databases Actually Work (Behind the Scenes)

When working through the How Vector Databases Actually stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

HNSW (Hierarchical Navigable Small World)

When working through the HNSW Hierarchical Navigable Small stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

IVF Inverted File Index(Clustering for fast search)

When working through the IVF Inverted File Index stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the IVF Inverted File Index stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Comparison of IVF and HNSW

The Comparison of IVF and stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

The problem with pure Vector Search

The The problem with pure stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

Hybrid Search: Best of Both Worlds

The Hybrid Search Best of stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Hybrid Search Best of stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Re-ranking in RAG Systems: From Good Results to the Best Results

For the Re-ranking in RAG Systems stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

"""Super-simple CrossEncoder reranking example.

Steps:
1. Define a query and a few short documents.
2. Build (query, doc) pairs.
3. Use a CrossEncoder to get a relevance score for each pair.
4. Print raw scores, then print documents sorted by score.
"""

from sentence_transformers import CrossEncoder


def main() -> None:
        # 1. Create model
model_name = "cross-encoder/ms-marco-MiniLM-L-6-v2"
print(f"Loading CrossEncoder model: {model_name}\n")
model = CrossEncoder(model_name)

    # 2. Query and documents
        query = "What is a vector database?"
documents = [
        "A vector database stores embeddings and allows similarity search.",
        "Relational databases store structured data in tables.",
        "FAISS is a library for efficient similarity search of vectors.",
        "Vector databases are used in AI applications like RAG.",
        ]

        # 3. Build (query, doc) pairs
pairs = [(query, doc) for doc in documents]

        # 4. Get scores
scores = model.predict(pairs)

print("Query:\n  " + query + "\n")
print("Raw scores (higher = more relevant):")
    for doc, score in zip(documents, scores):
print(f"  score={score:.4f}  |  doc={doc}")

    # 5. Sort by score (descending)
ranked = sorted(zip(documents, scores), key=lambda x: x[1], reverse=True)

print("\nDocuments sorted by cross-encoder score:\n")
    for rank, (doc, score) in enumerate(ranked, start=1):
print(f"Rank {rank}: score={score:.4f}")
print(f"  {doc}\n")


if __name__ == "__main__":
main()

Output
****************************************************
Query:
  What is a vector database?

Raw scores (higher = more relevant):
  score=8.4591  |  doc=A vector database stores embeddings and allows similarity search.
  score=-7.1590  |  doc=Relational databases store structured data in tables.
  score=1.0106  |  doc=FAISS is a library for efficient similarity search of vectors.
  score=7.0736  |  doc=Vector databases are used in AI applications like RAG.

Documents sorted by cross-encoder score:

Rank 1: score=8.4591
  A vector database stores embeddings and allows similarity search.

Rank 2: score=7.0736
  Vector databases are used in AI applications like RAG.

Rank 3: score=1.0106
  FAISS is a library for efficient similarity search of vectors.

Rank 4: score=-7.1590
  Relational databases store structured data in tables.

Understanding Similarity Measures in Vector Search (Cosine, Dot Product, Euclidean)

For the Understanding Similarity Measures in stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

Conclusion

For the Conclusion stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Conclusion stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Operational checklist

When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.

Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.

Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for 5984e1048405: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.