Home / Articles / Practical notes: Parent–Child Chunking and Hierarchical Indexing in RAG: A

This article is published in English.

Practical notes: Parent–Child Chunking and Hierarchical Indexing in RAG: A

Operable walkthrough of Practical notes: Parent–Child Chunking and Hierarchical Indexing in RAG: A: contracts, checks, and drop-in code slots for teams shipping this pattern.

2451 words

This walkthrough rebuilds the path from raw materials to a working system for: Parent–Child Chunking and Hierarchical Indexing in RAG: A Practical Guide to Better Retrieval. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Why chunking is actually the hardest part of RAG

When working through the Why chunking is actually stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

A typical failure example

When working through the A typical failure example stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

from sentence_transformers import SentenceTransformer
import numpy as np

model = SentenceTransformer("all-MiniLM-L6-v2")

def cosine_similarity(a, b):
    return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))

query = "What is a risk of caching?"
chunk_a = "Caching improves system performance significantly."
chunk_b = "However, improper cache invalidation can lead to stale data issues."

query_embedding = model.encode(query)
score_a = cosine_similarity(query_embedding, model.encode(chunk_a))
score_b = cosine_similarity(query_embedding, model.encode(chunk_b))

print(f"Chunk A similarity: {score_a:.4f}  - '{chunk_a}'")
print(f"Chunk B similarity: {score_b:.4f}  - '{chunk_b}'")
Chunk A similarity: 0.5891  — 'Caching improves system performance significantly.'
Chunk B similarity: 0.5103  — 'However, improper cache invalidation can lead to stale data issues.'

Parent–child chunking

When working through the Parent child chunking stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Parent child chunking stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

How it works in practice

The How it works in stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

def build_parent_child_index(document_section, parent_text, child_splitter, embedding_model):
    child_chunks = child_splitter(parent_text)
    records = []
    for child_text in child_chunks:
        records.append({
            "text": child_text,
            "embedding": embedding_model.encode(child_text),
            "parent_text": parent_text  # full parent stored directly on every child
        })
    return records


parent_text = """Indexing improves query speed by reducing scan time across large tables.
Caching reduces repeated computation but may introduce staleness if not invalidated properly.
Query optimization involves rewriting SQL queries to use more efficient execution plans."""

child_texts = [
    "Indexing improves query speed by reducing scan time across large tables.",
    "Caching reduces repeated computation but may introduce staleness if not invalidated properly.",
    "Query optimization involves rewriting SQL queries to use more efficient execution plans.",
]

records = [
    {"text": t, "parent_text": parent_text} for t in child_texts
]

for r in records:
    print(f"Child: {r['text'][:50]}...")
    print(f"  → linked to parent ({len(r['parent_text'])} chars)\n")

The key idea: retrieve small, expand big

The The key idea retrieve stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

def parent_child_retrieve(query, child_records, embedding_model, top_k=3):
    query_embedding = embedding_model.encode(query)
    scored = [
        (np.dot(query_embedding, r["embedding"]) /
         (np.linalg.norm(query_embedding) * np.linalg.norm(r["embedding"])), r)
        for r in child_records
    ]
    scored.sort(key=lambda x: x[0], reverse=True)
    top_children = scored[:top_k]

    # Expand each matched child to its parent - but deduplicate first
    seen_parents = set()
    expanded_context = []

    for score, record in top_children:
        parent = record["parent_text"]
        if parent not in seen_parents:
            expanded_context.append(parent)
            seen_parents.add(parent)

     return expanded_context

for r in records:
    r["embedding"] = model.encode(r["text"])

context = parent_child_retrieve("What is a risk of caching?", records, model, top_k=2)
print("Context sent to the LLM:\n")

for c in context:
    print(c)

Why this works so well

The Why this works so stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Why this works so stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Hierarchical indexing: taking it one step further

For the Hierarchical indexing taking it stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

Document: "Cloud Architecture Guide"
  Section: Networking
    Subsection: Load Balancing
      Chunk: Round-robin method
      Chunk: Least connections method
    Subsection: CDN usage
  Section: Security
    Subsection: IAM policies
    Subsection: Encryption

Why hierarchy matters

For the Why hierarchy matters stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

Hierarchical retrieval in action

For the Hierarchical retrieval in action stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Hierarchical retrieval in action stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Parent–child vs hierarchical indexing

When working through the Parent child vs hierarchical stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Implementation pattern (real-world RAG pipeline)

When working through the Implementation pattern real-world RAG stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Common mistakes in real systems

When working through the Common mistakes in real stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Common mistakes in real stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Where this makes the biggest difference

The Where this makes the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

The bigger picture

The The bigger picture stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

Final thought

The Final thought stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Final thought stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Operational checklist

The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.

Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

Score single-turn answers and multi-turn trajectories separately. Aggregate chat scores bury tool-loop failures.

Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.

Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for de861b3bd2cb: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.

For the hardening note 0 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Hardening detail 0/954: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

When working through the hardening note 1 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Hardening detail 1/954: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 2 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Hardening detail 2/954: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.