This article is published in English.
When a RAG bot quotes the wrong insurance deductible
A fixed-size chunker split a dollar amount across lines; rebuild parsing and climb four chunking tiers—from recursive baselines to parent-document retrieval—before trusting answers.
Everything below is a composite narrative for teaching. Company names, people, and dollar amounts are invented; the failure modes are the ones production RAG systems actually hit in regulated insurance and similar high-stakes domains.
Late on a Tuesday, a claims lead pinged engineering: the assistant had told a customer their flood deductible was $500 when the policy said $5,000. The retrieval-augmented bot—“Atlas”—had been live three weeks with a vector store, a solid embedding model, a capable LLM, and a polished UI. What it lacked was a defensible chunking strategy.
The retrieved fragment looked like this:
...applies to all covered perils. Deductible: $5
00 for flood events in Zone A, $500 for wind events, and $1,0
PDF extraction had broken a number across lines. A fixed-size chunker then sliced every 500 characters on top of that break. The embedder indexed a shard that said “$5” and “$500.” The model answered confidently—and wrongly.
What follows is how the ingestion pipeline was rebuilt: one parsing stage that sets the ceiling, then four tiers of chunking that climb toward that ceiling at rising cost.
One framing keeps the economics honest: almost all chunking cost is one-time indexing compute. It runs offline, before any user question. The only chunking-adjacent cost that shows up in query p95 is the embedding model itself. “Expensive” for higher tiers means expensive once per document, not per request.
Stage 0: You cannot chunk structure you destroyed at parse time
The first instinct is a smarter chunker. The better first question is: show the text being chunked, not the PDF. Early Atlas text was a wall of characters. Headings had collapsed into body copy. A deductible table had flattened into a run-on where values from different rows sat beside each other. Page chrome (“Policy Form HO-3 — Page 14 of 62”) repeated every few thousand characters.
No chunker recovers boundaries the parser already destroyed. If headings are gone, a markdown-header splitter has nothing to split on. If tables dissolved, nothing keeps rows intact.
Match each input type to a parser that emits structure the chunker can use:
Native PDFs and DOCX (policy forms, endorsements, guidelines). Prefer layout-aware parsers—LlamaParse, Unstructured.io, Docling, Azure Document Intelligence—over naive text dumps. Require markdown with element types (title, table, list), not a bare string.
Scanned PDFs, images, and faxes. Run OCR (Tesseract is weakest; PaddleOCR, Textract, Document AI, Azure are stronger). Demand text plus layout blocks plus per-block confidence. Confidence later flags “this number is uncertain—do not answer deductibles from it.”
HTML and internal wikis. Strip chrome and emit clean markdown with headings preserved.
Decks and spreadsheets. Keep slide or sheet boundaries so a chunk never mixes unrelated slides.
Audio and video (claims calls, webinars). Whisper, Deepgram, or AssemblyAI should emit a transcript that labels who spoke and when. If speakers are not separated, opposite statements from an adjuster and a customer can fuse into a single misleading paragraph.
After layout-aware re-parse, the deductible section looked like:
## Section 4 — Deductibles
### 4.2 Peril-specific deductibles
| Peril | Zone A | Zone B |
|----------------|---------|---------|
| Flood | $5,000 | $2,500 |
| Wind / hail | $500 | $500 |
| Named storm | $1,000 | $1,000 |
The table returned. The heading tree returned. “$5,000” was one token again. Chunking code had not changed yet, and retrieval was already safer.
Tier 1: Syntactic chunking — cheap, fast, and where you should start
Fixed-size: the prototype that shipped by accident
Split every N characters or tokens (CharacterTextSplitter, tiktoken). Index cost is near zero. Fine for a spike; fatal in production. It cuts mid-sentence, mid-table, mid-number—the exact shape of the deductible incident. Rule after that night: fixed-size never leaves a notebook.
Teams keep rediscovering the same pattern: a demo corpus of short blog posts survives character splitting, then the first multi-column PDF or OCR’d endorsement arrives and answer quality collapses overnight. Treat fixed windows as scaffolding for local experiments, with an explicit checklist item to replace them before any external user traffic.
Overlap: a dial, not a strategy
chunk_overlap repeats the tail of chunk N at the head of N+1. Ten to fifteen percent is common (for example 50 on a 500-token chunk). Overlap is cheap insurance for tiers 1–2; it does not fix a bad boundary—it duplicates it. Expect ~10% more vectors and duplicate hits in top-k when the match sits in the overlap.
When duplicates dominate the neighbor list, rerankers and the LLM waste context window on the same sentence twice. Deduplicating by parent id or by near-identical text hashes after retrieval is a cheap mitigation if overlap must stay high for other reasons.
Sentence/paragraph: respects grammar, ignores the document
Pack whole sentences into a budget (NLTK, spaCy, LlamaIndex SentenceSplitter). Never cuts mid-sentence—so it would have blocked the line-broken number—but it does not know headings, tables, or exclusion lists. Excellent on call transcripts; mediocre on structured policy forms.
Recursive character: the default you ship first
RecursiveCharacterTextSplitter tries separators in priority order—blank line, newline, sentence, word—falling back only when a piece is still too large, then merging small pieces toward chunk_size. A common baseline is 500 tokens with 50 overlap:
from langchain_text_splitters import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(
chunk_size=500,
chunk_overlap=50,
separators=["\n\n", "\n", ". ", " ", ""],
)
chunks = splitter.split_text(policy_markdown)
On the re-parsed deductible section it kept all of 4.2—including the table—because blank lines made a natural unit under 500 tokens. The bug disappeared. Ship recursive-500/50 as the baseline to measure against, not as the finished design.
Write that baseline score on a whiteboard and refuse promotions of “smarter” chunkers that cannot beat it on the same golden questions. Many expensive ideas look brilliant until they are compared to a boring recursive splitter on clean markdown.
Tier 2: Structure-aware chunking — cut where the document already draws lines
With a golden set of ~80 real questions, recursive-500/50 hit ~71%. Failures clustered on long sections spanning chunks and on answers where the model could not tell which policy or section the hit came from.
Markdown / HTML header splitting
MarkdownHeaderTextSplitter / HTMLHeaderTextSplitter cut on heading hierarchy and carry the header path into metadata:
from langchain_text_splitters import MarkdownHeaderTextSplitter
header_splitter = MarkdownHeaderTextSplitter(
headers_to_split_on=[("#", "h1"), ("##", "h2"), ("###", "h3")]
)
sections = header_splitter.split_text(policy_markdown)
# then apply the recursive splitter inside each section
Each fragment carved from subsection 4.2 should carry a breadcrumb such as form HO-3 → deductibles → peril-specific table. Concatenate that breadcrumb onto the text before embedding so the vector encodes location as well as numbers. “Which section?” failures drop; citations can say “per Section 4.2.” Cost stays near zero—but only if Stage 0 emitted real markdown. Flattened text silently becomes one giant chunk. Best on wikis, technical docs, and headed policies.
Metadata fields should travel with the vector: source document id, section path, effective date, and jurisdiction if policies vary by state. Those fields power filters (“only HO-3 in force after 2024-01-01”) that pure similarity search cannot express.
Layout-aware / by_title
Unstructured chunk_by_title (and peers in LlamaParse/Docling) respect parser element boundaries: tables stay whole, titles start chunks, lists keep their intro sentence. Right tool for contracts and complex PDFs. Parse APIs cost money once at index time. A cheap upstream parser nullifies the tier—do not buy a fancy steering wheel for a car with no engine.
Code-aware (AST)
Irrelevant for policy corpora, essential for Terraform/Python RAG. Line splits separate signatures from bodies. tree-sitter or language-aware recursive splitters cut on function/class nodes. Default when the corpus is repos or IaC.
Mixing prose policies and source code in one index without separating splitters usually produces the worst of both worlds: functions truncated mid-body and policy tables shredded by brace-oriented heuristics.
Tier 3: Model-based chunking — pay for judgment only where structure is gone
Around 84% on the golden set, a new failure class appeared: unstructured call transcripts with no headings. Recursive 500-token blocks straddled topic changes—roof claim then mailing address—so embeddings averaged two topics and matched neither well.
Semantic chunking
Embed sentence by sentence; cut where consecutive cosine similarity drops past a threshold (SemanticChunker, any decent embedder). One embedding pass at index time. Transformative on speech; thresholds are corpus-specific. The phone-call threshold shredded dense policy prose into tiny pieces—so use semantic chunking only on unstructured speech and keep policies on tier 2.
Routing by document class—speech versus form versus wiki—beats searching for one universal chunker. The orchestration cost is a classifier or simple file-type switch at ingest; the payoff is fewer muddy embeddings.
Proposition / atomic-fact chunking
A small LLM rewrites passages into standalone statements (“The flood deductible for Zone A under HO-3 is $5,000”). Precision rises because each unit is one fact with context inlined. Cost is one LLM call per passage—sane for small, high-precision corpora such as a 40-page FAQ. Trap: propositions drift from source wording. When customers dispute answers, verbatim citation beats paraphrase. Keep propositions as a retrieval aid and return the original passage for citation.
Compliance and claims organizations will eventually ask for the page image or PDF highlight behind an answer. Architect citation ids early so proposition indexes remain an accelerator rather than the system of record.
Agentic chunking
Hand a whole document to a flagship model to choose boundaries. It works and scales poorly: cost grows linearly with corpus size while value does not. Fine for a few hundred crown-jewel docs; not for millions.
Tier 4: Retrieval-time chunking — search small units, return larger ones
Tiers 1–3 assume the indexed unit is the returned unit. That forces a false trade-off: small chunks embed cleanly but starve the LLM of context; large chunks give context but blur embeddings. Tuning chunk_size forever never finds a universal sweet spot.
Tier 4 splits roles: a search unit sized for the embedder and a delivery unit sized for the LLM. Litmus test: if you need a second store (docstore, parent map, tree), you are in tier 4.
Parent-document / small-to-big
Index small children (150–200 tokens). On match, return the larger parent (whole Section 4.2 or a ~1,500-token subsection):
from langchain.retrievers import ParentDocumentRetriever
from langchain.storage import InMemoryStore
parent_splitter = RecursiveCharacterTextSplitter(chunk_size=1500)
child_splitter = RecursiveCharacterTextSplitter(chunk_size=200)
retriever = ParentDocumentRetriever(
vectorstore=vectorstore,
docstore=InMemoryStore(), # the "second store"
child_splitter=child_splitter,
parent_splitter=parent_splitter,
)
retriever.add_documents(policy_docs)
LlamaIndex AutoMergingRetriever offers a similar hierarchy. Extra index cost is near zero—you embed the same text in smaller pieces. On the golden set this jump alone moved Atlas from ~84% to ~91%: “what’s my flood deductible” hit the precise table row while the LLM still saw the surrounding notes (including how Zone A is defined). Operational cost: a docstore keyed by parent ID that stays in sync on updates.
Deletion and versioning matter here: when Section 4.2 is revised, children and parents must refresh together or the system will retrieve a precise child that opens a stale parent. Treat ingest as a transactional publish of parent, children, and embeddings—not three independent jobs.
Contextual retrieval
Before embedding, ask a mini LLM for one or two situating sentences and prepend them; pair with BM25. Fragments like “Zone A: $5,000 / Zone B: $2,500” become retrievable. Cost is one LLM call per chunk—brutal without prompt caching; with caching the document prefix stays warm and per-chunk calls stay cheap.
Late chunking
Embed the whole document with a long-context embedder, then mean-pool token vectors inside each chunk span. Chunk vectors carry document context without per-chunk LLM calls. Competitive with contextual retrieval, but locks you to long-context embedding models.
Hierarchical / RAPTOR
Cluster chunks, summarize, reclustersummaries into a tree; index every level. Broad questions hit summaries; specifics hit leaves. Highest tier-4 cost and painful when documents change—rebuilds are expensive. Quarterly policy updates made RAPTOR a pass for Atlas.
Static knowledge bases—internal encyclopedias that change yearly—can tolerate tree rebuilds. Living insurance forms cannot. Choose hierarchical methods only when change rate and rebuild budget are written down.
What the rebuilt playbook says
Six weeks after the Slack fire drill, Atlas sat near 93% on a golden set grown to ~300 questions. The short design-review version:
Start with recursive character splitting plus structure-aware cuts at roughly 500 tokens and 50 overlap—it costs almost nothing and gives a measurable baseline. Upgrade first to parent-document retrieval so matching and delivery sizes can differ. Only then spend on contextual retrieval if the evaluation set still blames recall.
Nothing above rescues a bad parse. Repair extraction before tuning splitters.
The one-page version
Stage 0 — Parse. Match parsers to inputs: structured markdown for native docs; text+layout+OCR confidence for scans; clean headed markdown for HTML; slide/sheet bounds for decks; diarized transcripts for audio. This sets the ceiling.
Tier 1 — Syntactic. Fixed-size for prototypes only. Overlap is a dial (10–15%) that costs duplicate top-k hits. Sentence splits respect grammar, ignore document logic. Recursive 500/50 is the default baseline, not the finish line.
Tier 2 — Structure-aware. Header splits carry section paths; only as good as the parser. Layout by_title keeps tables whole; cheap parsers nullify it. AST splits for code repos.
Tier 3 — Model-based. Semantic cuts on similarity drops; tune per corpus. Propositions boost precision on small corpora but drift from wording. Agentic cuts rarely justify at scale.
Tier 4 — Retrieval-time. Match on small units; deliver large ones. Parent-document is the highest-ROI upgrade. Contextual retrieval needs caching. Late chunking is cheaper but embedder-locked. RAPTOR stales on updates. If it needs a second store, it is tier 4.
The deductible incident was never really about choosing LangChain versus LlamaIndex. It was about respecting document structure before math, measuring chunkers on real customer questions, and refusing to hand the model fragments that could not possibly contain a complete numeric fact.
Build the golden set from tickets that already embarrassed the product—wrong deductibles, missing exclusions, confused endorsements—not from synthetic trivia. Score answer correctness and citation faithfulness separately so a fluent wrong number never looks like a win. When a new splitter lands, require it to beat the recursive baseline on both metrics before it touches production traffic.
Finally, keep a kill switch: if OCR confidence on a numeric span is low, or if parent and child versions disagree, the assistant should refuse the deductible question and escalate rather than invent confidence. Users forgive “cannot verify that from the filed form” far faster than a casually wrong five-hundred-dollar answer. That refusal path belongs in the product requirements, not only in a postmortem slide deck shared after the damage is done.