Home / Articles / Practical notes: RAG Chunking Strategies: The Complete Engineering Guide

This article is published in English.

Practical notes: RAG Chunking Strategies: The Complete Engineering Guide

Operable walkthrough of Practical notes: RAG Chunking Strategies: The Complete Engineering Guide: contracts, checks, and drop-in code slots for teams shipping this pattern.

2654 words

The following notes reconstruct a practical path around “RAG Chunking Strategies: The Complete Engineering Guide”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

The Core Trade-Off: Context vs. Precision

The The Core Trade-Off Context stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

[ TOO SMALL CHUNKS ]                  [ TOO LARGE CHUNKS ]
Loss of Context / Meaning             "Lost in the Middle" Effect
(e.g., "IF request is made...")       (Dense with irrelevant fluff)
            │                                      │
            └───────────────► 🎯 ◄─────────────────┘
                      THE GOLDILOCKS ZONE
                    (High Precision + Context)

Strategy 1: Fixed-Size & Recursive Splitting (The Baselines)

The Strategy 1 Fixed-Size Recursive stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

Fixed-Size Chunking (The Naive Approach)

The Fixed-Size Chunking The Naive stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Fixed-Size Chunking The Naive stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Recursive Character Chunking (The Production Baseline)

For the Recursive Character Chunking The stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

# Example: LangChain Recursive Character Splitter
from langchain.text_splitter import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(
    chunk_size=512,
    chunk_overlap=50,
    separators=["\n\n", "\n", ". ", " ", ""]
)

Strategy 2: Structure-Aware & Semantic Chunking

For the Strategy 2 Structure-Aware Semantic stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

Structure-Aware (Document-Native) Chunking

For the Structure-Aware Document-Native Chunking stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Structure-Aware Document-Native Chunking stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Semantic Chunking (Meaning-Based Boundaries)

When working through the Semantic Chunking Meaning-Based Boundaries stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Sentence A ──► Embed ──┐
Sentence B ──► Embed ──┴── Similarity: 0.89 (Keep together)
Sentence C ──► Embed ──── Similarity: 0.32 (Drop below threshold -> CUT HERE)

Strategy 3: Advanced Context-Preserving Architectures

When working through the Strategy 3 Advanced Context-Preserving stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

1. Small-to-Big (Parent-Child) Chunking

When working through the 1 Small-to-Big Parent-Child Chunking stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the 1 Small-to-Big Parent-Child Chunking stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

2. Late Chunking

The 2 Late Chunking stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

Production Decision Matrix

The Production Decision Matrix stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

Summary Engineering Rules for Success

The Summary Engineering Rules for stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Summary Engineering Rules for stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Time to get the hands dirty..

For the Time to get the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

Prerequisites

For the Prerequisites stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

pip install langchain langchain-community langchain-experimental llama-index llama-index-embeddings-openai openai chromadb
export OPENAI_API_KEY="your-openai-api-key"

1. Recursive Character Chunking (LangChain)

For the 1 Recursive Character Chunking stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the 1 Recursive Character Chunking stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

from langchain_text_splitters import RecursiveCharacterTextSplitter

# Sample document
sample_text = """
# RAG System Design
Retrieval-Augmented Generation (RAG) decouples knowledge storage from reasoning capability.
Instead of forcing the AI to answer strictly from memory, RAG searches an external database first.
## The Ingestion Pipeline
1. Document Extraction: PDFs and web pages are converted into clean text.
2. Chunking: Documents are chopped into smaller, manageable text blocks.
3. Vectorization: Each block is converted into a numeric representation.
"""

# Initialize Recursive Splitter
text_splitter = RecursiveCharacterTextSplitter(
    chunk_size=200,        # Target chunk size in characters/tokens
    chunk_overlap=30,      # Overlap to prevent mid-sentence context loss
    separators=["\n\n", "\n", ". ", " ", ""] # Try double line breaks first
)
chunks = text_splitter.create_documents([sample_text])
print(f"Total chunks created: {len(chunks)}\n")
for i, chunk in enumerate(chunks):
    print(f"--- Chunk {i+1} ---")
    print(chunk.page_content)

2. Parent-Child / Small-to-Big Chunking (LangChain + Chroma)

When working through the 2 Parent-Child Small-to-Big Chunking stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

from langchain_text_splitters import RecursiveCharacterTextSplitter
from langchain.retrievers import ParentDocumentRetriever
from langchain_community.vectorstores import Chroma
from langchain_community.storage import InMemoryStore
from langchain_openai import OpenAIEmbeddings
from langchain_core.documents import Document
# 1. Initialize Vector Store (for small child embeddings) and DocStore (for large parent text)
embeddings = OpenAIEmbeddings()
vectorstore = Chroma(collection_name="parent_child_rag", embedding_function=embeddings)
docstore = InMemoryStore()

# 2. Define Parent (Big) and Child (Small) Splitters
parent_splitter = RecursiveCharacterTextSplitter(chunk_size=600, chunk_overlap=50)
child_splitter = RecursiveCharacterTextSplitter(chunk_size=120, chunk_overlap=20)

# 3. Create the ParentDocumentRetriever
retriever = ParentDocumentRetriever(
    vectorstore=vectorstore,
    docstore=docstore,
    child_splitter=child_splitter,
    parent_splitter=parent_splitter,
)

# 4. Ingest Documents
docs = [
    Document(
        page_content="""
        System Architecture: The Two Pipelines.
        A production RAG framework runs on two main pipelines: Ingestion and Inference.
        The Ingestion pipeline extracts text, creates chunks, embeds them, and stores them in a vector DB.
        The Inference pipeline encodes user queries, performs vector similarity search, injects context into prompts, and generates LLM answers.
        Hybrid search combines keyword and vector retrieval to handle exact SKU IDs alongside general concepts.
        """
    )
]
retriever.add_documents(docs)

# 5. Query the Retriever
query = "What happens during inference in RAG?"
retrieved_parents = retriever.invoke(query)
print(f"Retrieved Parent Context (Full Block):\n")
print(retrieved_parents[0].page_content)

3. Semantic Chunking (LlamaIndex)

When working through the 3 Semantic Chunking LlamaIndex stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

from llama_index.core.node_parser import SemanticSplitterNodeParser
from llama_index.embeddings.openai import OpenAIEmbedding
from llama_index.core.schema import Document
# 1. Initialize the Embedding Model used for detecting topic shifts
embed_model = OpenAIEmbedding(model="text-embedding-3-small")

# 2. Configure the Semantic Splitter
semantic_parser = SemanticSplitterNodeParser(
    buffer_size=1,                         # Number of surrounding sentences to evaluate together
    breakpoint_percentile_threshold=90,     # Split threshold percentile (higher = fewer, larger chunks)
    embed_model=embed_model
)

# 3. Sample document with distinct thematic shifts
raw_text = """
Quantum computing leverages qubits that can exist in superposition states, unlike classical bits.
Superposition allows algorithms to process vast potential outcomes simultaneously.
Entanglement further connects qubit states instantaneously across physical space.
On a totally different topic, baking sourdough bread requires maintaining a wild yeast starter.
You feed the starter equal parts flour and water every 24 hours to encourage fermentation.
Proper gluten development requires folding the dough during the bulk fermentation stage.
"""
doc = Document(text=raw_text)

# 4. Generate Nodes (Chunks)
nodes = semantic_parser.get_nodes_from_documents([doc])
print(f"Total Semantically Coherent Chunks: {len(nodes)}\n")
for i, node in enumerate(nodes):
    print(f"--- Chunk {i+1} ---")
    print(node.get_content().strip())
    print("\n")

Implementation Summary

When working through the Implementation Summary stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Implementation Summary stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Conclusion: Chunking is an Engineering Problem, Not an AI Problem

The Conclusion Chunking is an stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

Operational checklist

For the Operational checklist stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.