This article is published in English.
Practical notes: Metadata Enrichment in RAG: The Secret Ingredient for Better
Operable walkthrough of Practical notes: Metadata Enrichment in RAG: The Secret Ingredient for Better: contracts, checks, and drop-in code slots for teams shipping this pattern.
Use this as an operator-facing rebuild of the ideas in “Metadata Enrichment in RAG: The Secret Ingredient for Better Retrieval”: clear stages, ordered code slots, and recovery notes that survive a handoff.
Introduction
The Introduction stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
The Problem with Content-Only Retrieval
The The Problem with Content-Only stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
What Exactly Is Metadata?
The What Exactly Is Metadata stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Employees are entitled to 20 weeks of paid maternity leave.
{
"source": "EU_HR_Policy.pdf",
"region": "Europe",
"department": "Human Resources",
"last_updated": "2025-03-15"
}
How Metadata Improves Retrieval
The How Metadata Improves Retrieval stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
filter = {
"region": "Europe"
}
from sentence_transformers import SentenceTransformer
import numpy as np
embedding_model = SentenceTransformer("all-MiniLM-L6-v2")
chunks = [
{"text": "Employees are entitled to 20 weeks of paid maternity leave.", "region": "United States"},
{"text": "Employees are entitled to 16 weeks of paid maternity leave.", "region": "Europe"},
{"text": "Employees are entitled to 26 weeks of paid maternity leave.", "region": "Asia-Pacific"},
]
for c in chunks:
c["embedding"] = embedding_model.encode(c["text"])
query = "What is the maternity leave policy for employees in the European office?"
query_embedding = embedding_model.encode(query)
def search(chunks, query_embedding, region_filter=None):
candidates = chunks if region_filter is None else [c for c in chunks if c["region"] == region_filter]
scored = [(np.dot(query_embedding, c["embedding"]), c) for c in candidates]
scored.sort(key=lambda x: x[0], reverse=True)
return scored[0][1]
print("Without metadata filter:")
result = search(chunks, query_embedding)
print(f" Returned: '{result['text']}' (region: {result['region']})")
print("\nWith metadata filter (region='Europe'):")
result = search(chunks, query_embedding, region_filter="Europe")
print(f" Returned: '{result['text']}' (region: {result['region']})")
Types of Metadata That Matter in RAG
The Types of Metadata That stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Types of Metadata That stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
1. Source Metadata
For the 1 Source Metadata stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
def format_citation(metadata):
return f"Source: {metadata['source']}, Page {metadata['page']}"
chunk_metadata = {"source": "Employee_Handbook_2025.pdf", "page": 42}
print(format_citation(chunk_metadata))
Source: Employee_Handbook_2025.pdf, Page 42
2. Organizational Metadata
For the 2 Organizational Metadata stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
filter = {"department": "Finance"}
3. Temporal Metadata
For the 3 Temporal Metadata stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the 3 Temporal Metadata stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
chunks = [
{"text": "Employees receive 15 vacation days per year.", "last_updated": "2023-01-10"},
{"text": "Employees receive 20 vacation days per year.", "last_updated": "2025-03-15"},
]
most_recent = max(chunks, key=lambda c: c["last_updated"])
print(most_recent["text"])
Employees receive 20 vacation days per year.
4. Security Metadata
When working through the 4 Security Metadata stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
chunk_metadata = {
"access_level": "confidential",
"allowed_roles": ["HR_Manager", "HR_Admin"]
}
def can_access(user_role, chunk_metadata):
return user_role in chunk_metadata["allowed_roles"]
print(can_access("HR_Manager", chunk_metadata)) # True
print(can_access("Engineering", chunk_metadata)) # False
True
False
Metadata Enrichment During Ingestion
When working through the Metadata Enrichment During Ingestion stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
from langchain_core.documents import Document
doc = Document(
page_content="Employees receive 20 weeks of maternity leave.",
metadata={
"department": "HR",
"region": "Europe",
"source": "EU_HR_Policy.pdf"
}
)
print(doc.page_content)
print(doc.metadata)
Employees receive 20 weeks of maternity leave.
{'department': 'HR', 'region': 'Europe', 'source': 'EU_HR_Policy.pdf'}
def chunk_with_metadata(document, chunk_size=60):
text = document.page_content
chunks = [text[i:i+chunk_size] for i in range(0, len(text), chunk_size)]
return [
Document(page_content=chunk_text, metadata=document.metadata)
for chunk_text in chunks
]
source_doc = Document(
page_content="Employees receive 20 weeks of maternity leave. Eligibility begins after six months of employment.",
metadata={"department": "HR", "region": "Europe", "source": "EU_HR_Policy.pdf"}
)
chunked_docs = chunk_with_metadata(source_doc)
for i, chunk in enumerate(chunked_docs, 1):
print(f"Chunk {i}: {chunk.page_content!r}")
print(f" Metadata: {chunk.metadata}\n")
Chunk 1: 'Employees receive 20 weeks of maternity leave. Eligibility b'
Metadata: {'department': 'HR', 'region': 'Europe', 'source': 'EU_HR_Policy.pdf'}
Chunk 2: 'egins after six months of employment.'
Metadata: {'department': 'HR', 'region': 'Europe', 'source': 'EU_HR_Policy.pdf'}
AI-Generated Metadata
When working through the AI-Generated Metadata stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the AI-Generated Metadata stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
def generate_metadata(text, llm):
prompt = f"""Analyze the following document and return metadata as JSON with these fields:
topic, category, and keywords (a list of 3-5 relevant terms).
Document:
{text}
Respond with only the JSON, no other text."""
response = llm.generate(prompt)
return response
sample_text = "Employees are entitled to 20 weeks of paid maternity leave, with eligibility beginning after six months of continuous employment. Additional unpaid leave may be requested with manager approval."
# In practice, llm.generate() would call an actual model (Claude, GPT, etc.)
# Below is the kind of output this prompt is designed to produce:
example_output = {
"topic": "Employee Benefits",
"category": "HR Policy",
"keywords": ["maternity leave", "eligibility", "employee benefits"]
}
print(example_output)
{'topic': 'Employee Benefits', 'category': 'HR Policy', 'keywords': ['maternity leave', 'eligibility', 'employee benefits']}
Metadata and Hybrid Retrieval
The Metadata and Hybrid Retrieval stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
The Hidden Power of Metadata in Enterprise RAG
The The Hidden Power of stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Conclusion
The Conclusion stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Conclusion stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Operational checklist
When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Freeze a golden set before changing prompts or models. Moving both the system and the yardstick hides regressions.
Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 0ef11f703754: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.
For the hardening note 0 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 0/765: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 1 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 1/765: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 2 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 2/765: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 3 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 3/765: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 4 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 4/765: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 5 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 5/765: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 6 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 6/765: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 7 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 7/765: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.