This article is published in English.
Practical notes: The Indexing Layer: How docling-pipelines Prepares Your
Operable walkthrough of Practical notes: The Indexing Layer: How docling-pipelines Prepares Your: contracts, checks, and drop-in code slots for teams shipping this pattern.
This walkthrough rebuilds the path from raw materials to a working system for: The Indexing Layer: How docling-pipelines Prepares Your Documents for RAG. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
What is RAG, and Where Does a Vector Database Fit?
When working through the What is RAG and stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Why Can’t a Regular Database Do This?
When working through the Why Can t a stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
VectorDB Support in docling-pipelines
When working through the VectorDB Support in docling-pipelines stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the VectorDB Support in docling-pipelines stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
How the VectorDB Operator Works
The How the VectorDB Operator stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Hexagonal Architecture: Plugging in Any VectorDB
The Hexagonal Architecture Plugging in stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
VectorDBOperator
│
└── VectorStorePort (interface / abstract contract)
├── OpenSearchAdapter (ships with docling-pipelines)
├── MilvusAdapter (ships with docling-pipelines)
└── YourCustomAdapter (implement VectorStorePort → plug in)
Chunked vs. Non-Chunked Documents
The Chunked vs Non-Chunked Documents stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Chunked vs Non-Chunked Documents stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Stale Chunk Cleanup
For the Stale Chunk Cleanup stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Auto-Detected Vector Dimensions
For the Auto-Detected Vector Dimensions stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Configuring the Full RAG Pipeline
For the Configuring the Full RAG stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Configuring the Full RAG stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
{
"flow_name": "rag-indexing-pipeline",
"global_config": {
"doc_column": "content",
"storage": "in-memory",
"execute_type": "local",
"enable_micro_batching": true,
"micro_batch_size": 10
},
"flow": [
{
"name": "ingest",
"type": "ingest_source",
"config": {
"provider": "filesystem",
"connection_params": { "paths": ["./documents"], "recursive": true },
"include_filter": "pdf,docx,txt"
}
},
{
"name": "extract",
"type": "extract_operator",
"depends_on": ["ingest"],
"config": {
"text_extraction": { "provider": "docling_library", "doc_column": "content" }
}
},
{
"name": "chunk",
"type": "chunker",
"depends_on": ["extract"],
"config": {
"doc_column": "content",
"chunk_size": 512,
"chunk_overlap": 50
}
},
{
"name": "embed",
"type": "embeddings",
"depends_on": ["chunk"],
"config": {
"provider": "litellm",
"embeddings_column": "embeddings",
"provider_config": {
"model_id": "openai/nomic-embed-text",
"api_base": "http://localhost:11434/v1",
"api_key": "${OLLAMA_API_KEY}"
}
}
},
{
"name": "store",
"type": "vectordb",
"depends_on": ["embed"],
"config": {
"provider": "opensearch",
"doc_id_column": "doc_id_hash",
"create_index": true,
"provider_config": {
"index_name": "my_rag_index",
"host": "localhost",
"port": 9200,
"username": "${OPENSEARCH_USERNAME}",
"password": "${OPENSEARCH_PASSWORD}",
"use_ssl": false,
"engine": "faiss",
"algorithm": "hnsw",
"space_type": "l2"
}
}
}
]
}
docling-pipelines --flow-file rag-pipeline.json
Milvus with Dual Embeddings
When working through the Milvus with Dual Embeddings stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Incremental Indexing
When working through the Incremental Indexing stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
The Bigger Picture
When working through the The Bigger Picture stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the The Bigger Picture stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Operational checklist
The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.
Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 6ace628912a6: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.