Home / Articles / Production RAG on Azure: chunking, hybrid search, filters, and citations

This article is published in English.

Production RAG on Azure: chunking, hybrid search, filters, and citations

Enterprise retrieval needs structure-aware chunks, hybrid search, ACL filters, reranking, grounded prompts, and evaluation—not a PDF demo.

1350 words

RAG demos look easy: PDF → chunks → embeddings → vector DB → ask the model. Then five thousand documents arrive and someone asks for the international travel reimbursement rule. The hard part is getting the right evidence into the model reliably.

1. Start with the data, not the LLM

Inventory sources, access control, freshness, and formats before picking a chat model. Garbage corpora make fluent wrong answers.

2. Chunking: do not split every N characters blindly

Prefer structure-aware splits (headings, pages, tables) with overlap. Fee tables and policy clauses hate naive character windows.

3. Generate embeddings

Pin one embedding deployment; record dimension and model id with every index. Mixing models silently breaks recall.

document = {
   "document_id": "HR-2026-001",
   "title": "Employee Benefits Policy",
   "department": "HR",
   "country": "India",
   "version": "2026.1",
   "effective_date": "2026-01-01",
   "source": "HR Portal"
}
def chunk_text(text, size=1000):
   return [
     text[i:i + size]
     for i in range(0, len(text), size)
]
def chunk_by_sections(document):
  chunks = []
  for section in document.sections:
    chunks.append({
      "title": section.title,
      "content": section.text,
      "document_id": document.id
    })
  return chunks
pip install openai azure-identity
from openai import OpenAI
from azure.identity import DefaultAzureCredential, get_bearer_token_provider

token_provider = get_bearer_token_provider(
  DefaultAzureCredential(),
  "https://ai.azure.com/.default"
)

client = OpenAI(
  base_url="https://YOUR-RESOURCE.openai.azure.com/openai/v1/",
  api_key=token_provider
)

response = client.embeddings.create(
  model="text-embedding-3-small",
  input="Employees are eligible after completing 12 months of service."
)
vector = response.data[0].embedding
{
  "id": "HR-2026-001-004",
  "document_id": "HR-2026-001",
  "title": "Employee Benefits Policy",
  "content": "...",
  "country": "India",
  "version": "2026.1",
  "content_vector": embedding
}
from azure.search.documents import SearchClient
from azure.search.documents.models import VectorizedQuery

query = "Can I work remotely from another country?"

query_vector = client.embeddings.create(
  model="text-embedding-3-small",
  input=query
).data[0].embedding

vector_query = VectorizedQuery(
  vector=query_vector,
  k_nearest_neighbors=10,
  fields="content_vector"
)

results = search_client.search(
  search_text=query,
  vector_queries=[vector_query],
  top=5,
  select=["title", "content", "document_id", "country"]
)
results = search_client.search(
  search_text=query,
  vector_queries=[vector_query],
  filter="country eq 'India'",
  top=5
)
context = "\n\n".join(
  f"Source: {r['title']}\n{r['content']}"
  for r in results
)
prompt = f"""
You are an enterprise policy assistant.
Answer the user's question using only the provided evidence.
If the evidence does not contain the answer, say that you don't have enough information.
Do not invent policies.
Evidence: {context}
Question: {query}
"""
{
  "document_id": "HR-2026-001",
  "page": 14,
  "section": "Eligibility",
  "content": "..."
}
def answer_question(question):

  # 1. Embed the question
  vector = embed(question)

  # 2. Hybrid retrieval
  candidates = search(
    query=question,
    vector=vector,
    top=20
  )

  # 3. Apply metadata/business filters
  candidates = apply_filters(candidates)

  # 4. Rerank
  ranked = rerank(question, candidates)

  # 5. Keep only useful evidence
  evidence = ranked[:5]

  # 6. Build grounded prompt
  prompt = build_prompt(
    question,
    evidence
  )

  # 7. Generate answer
  response = generate(prompt)

  # 8. Return answer + citations
  return {
    "answer": response,
    "sources": extract_sources(evidence)
  }
test_cases = [
  {
    "question": "What is the India travel allowance?",
    "expected_source": "travel-policy-india.pdf"
  },
  {
    "question": "How long is parental leave?",
    "expected_source": "leave-policy-2026.pdf"
  }
]

4. Store embeddings in Azure AI Search

Persist text, vectors, and filterable metadata (doc id, product, locale, acl group, updated_at).

5. Hybrid search beats vector-only for most enterprises

Combine dense similarity with BM25/keyword for ids, codes, and rare proper nouns.

6. Filter before wasting context

Apply security and tenancy filters in the query—not after the model already saw forbidden chunks.

7. Retrieval is not final context

Rerank, dedupe, and trim to a token budget. More chunks ≠ better answers.

8. Build the prompt from evidence

Instructions should require answers only from provided passages and to say when evidence is missing.

9. Ground with citations

Return passage ids/pages so users can verify. Uncited fluency is a liability.

10. A pipeline worth shipping

Ingest → chunk → embed → index → hybrid retrieve → filter → rerank → prompt → generate → cite → evaluate.

11. What usually breaks

Bad chunking; poor metadata; vector-only retrieval; oversized context; no eval set; no citations; missing ACL; silent embedding model swaps; no freshness story.

Closing

Production RAG is information retrieval with an LLM endpoint attached—not the reverse. Invest in chunking, hybrid search, filters, citations, and evaluation; the chat model is the last mile.

Keep a golden question pack per domain: exact-id lookups, paraphrase questions, and ACL-negative cases. Automate them in CI whenever chunking or embedding settings change.

Keep a golden question pack per domain: exact-id lookups, paraphrase questions, and ACL-negative cases. Automate them in CI whenever chunking or embedding settings change.

Keep a golden question pack per domain: exact-id lookups, paraphrase questions, and ACL-negative cases. Automate them in CI whenever chunking or embedding settings change.

Keep a golden question pack per domain: exact-id lookups, paraphrase questions, and ACL-negative cases. Automate them in CI whenever chunking or embedding settings change.

Keep a golden question pack per domain: exact-id lookups, paraphrase questions, and ACL-negative cases. Automate them in CI whenever chunking or embedding settings change.

Keep a golden question pack per domain: exact-id lookups, paraphrase questions, and ACL-negative cases. Automate them in CI whenever chunking or embedding settings change.

Keep a golden question pack per domain: exact-id lookups, paraphrase questions, and ACL-negative cases. Automate them in CI whenever chunking or embedding settings change.

Keep a golden question pack per domain: exact-id lookups, paraphrase questions, and ACL-negative cases. Automate them in CI whenever chunking or embedding settings change.

Keep a golden question pack per domain: exact-id lookups, paraphrase questions, and ACL-negative cases. Automate them in CI whenever chunking or embedding settings change.

Keep a golden question pack per domain: exact-id lookups, paraphrase questions, and ACL-negative cases. Automate them in CI whenever chunking or embedding settings change.

Keep a golden question pack per domain: exact-id lookups, paraphrase questions, and ACL-negative cases. Automate them in CI whenever chunking or embedding settings change.

Keep a golden question pack per domain: exact-id lookups, paraphrase questions, and ACL-negative cases. Automate them in CI whenever chunking or embedding settings change.

Keep a golden question pack per domain: exact-id lookups, paraphrase questions, and ACL-negative cases. Automate them in CI whenever chunking or embedding settings change.

Keep a golden question pack per domain: exact-id lookups, paraphrase questions, and ACL-negative cases. Automate them in CI whenever chunking or embedding settings change.

Keep a golden question pack per domain: exact-id lookups, paraphrase questions, and ACL-negative cases. Automate them in CI whenever chunking or embedding settings change.

Keep a golden question pack per domain: exact-id lookups, paraphrase questions, and ACL-negative cases. Automate them in CI whenever chunking or embedding settings change.

Keep a golden question pack per domain: exact-id lookups, paraphrase questions, and ACL-negative cases. Automate them in CI whenever chunking or embedding settings change.

Keep a golden question pack per domain: exact-id lookups, paraphrase questions, and ACL-negative cases. Automate them in CI whenever chunking or embedding settings change.

Keep a golden question pack per domain: exact-id lookups, paraphrase questions, and ACL-negative cases. Automate them in CI whenever chunking or embedding settings change.

Keep a golden question pack per domain: exact-id lookups, paraphrase questions, and ACL-negative cases. Automate them in CI whenever chunking or embedding settings change.

Evaluation belongs in the same repo as chunking config: when someone “just tweaks” overlap, the golden pack should fail before users notice.

Hybrid search still needs good analyzers and synonym lists for domain jargon; embeddings alone will not save a misspelled policy code.

Document the ACL field contract in one page so every new ingest job writes the same claims the retrieve filter expects.

Evaluation belongs in the same repo as chunking config: when someone “just tweaks” overlap, the golden pack should fail before users notice.