Home / Articles / Practical notes: Designing RAG for 100 Million Documents

This article is published in English.

Practical notes: Designing RAG for 100 Million Documents

Operable walkthrough of Practical notes: Designing RAG for 100 Million Documents: contracts, checks, and drop-in code slots for teams shipping this pattern.

5618 words

Use this as an operator-facing rebuild of the ideas in “Designing RAG for 100 Million Documents”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Here is the requirements:

For the Here is the requirements stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

Start With the Math

For the Start With the Math stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

Documents              = 100,000,000
Average chunks/document = 30
Embedding dimensions    = 768
Storage/dimension       = 4 bytes (float32)
100,000,000 × 30
= 3,000,000,000 vectors
3B × 768 × 4 bytes
≈ 9.2 TB
ANN index structures
chunk text
document metadata
ACL metadata
IDs
lexical indexes
replicas
snapshots
WALs
temporary indexes
operational headroom
100M documents × 50 chunks
= 5 billion vectors

The Architecture you Would Build

For the The Architecture you Would stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the The Architecture you Would stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

The Source of Truth Is Not Your Vector Database

When working through the The Source of Truth stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Ingestion Must Be Asynchronous

When working through the Ingestion Must Be Asynchronous stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

POST /documents
       │
       ▼
Store document
       │
       ▼
Create ingestion event
       │
       ▼
Return 202 Accepted

Every Ingestion Step Should Be Idempotent

When working through the Every Ingestion Step Should stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Every Ingestion Step Should stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

document_id
document_version
chunk_index
embedding_model_version
hash(
    tenant_id +
    document_id +
    document_version +
    chunk_index
)
upsert(chunk_id, ...)

Chunking Becomes an Infrastructure Decision

The Chunking Becomes an Infrastructure stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

embedding compute
vector storage
indexing work
retrieval fan-out
reindexing time
replication traffic
{
  "chunk_id": "c_982...",
  "document_id": "doc_51...",
  "document_version": 8,
  "tenant_id": "tenant_17",
  "page": 43,
  "section": "Risk Factors",
  "language": "en",
  "created_at": "...",
  "acl_groups": ["finance", "executives"],
  "embedding_version": "embed_v4"
}

Do Not Search the Entire Corpus Unless You Actually Have To

The Do Not Search the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

organization = current tenant
region       = Europe
topic        = refund policy
time         = current quarter
permissions  = documents user may access

Sharding Is How the Vector Layer Becomes Distributed

The Sharding Is How the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Sharding Is How the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

tenant_id
region
document namespace
language
time partition
Query
 ↓
128 shards
tenant_id = 429
 ↓
Shard group 18
 ↓
4 shards

Replication Solves a Different Problem

For the Replication Solves a Different stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

Shard A → Node 1
Shard B → Node 2
Shard C → Node 3
Shard A → Node 1 + Node 4
Shard B → Node 2 + Node 5
Shard C → Node 3 + Node 6

Vector Search Alone Is Not Enough

For the Vector Search Alone Is stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

RRF(d) = Σ 1 / (k + rank_i(d))

Be Careful With Sequential Prefiltering

For the Be Careful With Sequential stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Be Careful With Sequential stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

3B documents
   ↓
BM25 returns 10,000
   ↓
vector search only those
tenant
permissions
document status

Authorization Must Happen Before Evidence Reaches the LLM

When working through the Authorization Must Happen Before stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

Engineering documents
HR documents
Board documents
Legal documents
Executive compensation
Customer contracts
Vector Search
     ↓
Retrieve confidential chunk
     ↓
Send chunk to LLM
     ↓
Check permission
User
  ↓
Identity
  ↓
Groups / roles / ACL
  ↓
Authorized search space
  ↓
Retrieval

Retrieval Should Produce Candidates, Not Context

When working through the Retrieval Should Produce Candidates stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Vector candidates: 200
Lexical candidates: 200
          │
          ▼
        Fusion
          │
          ▼
    ~250 unique chunks
          │
          ▼
       Reranker
          │
          ▼
       top 20–40

Then Remove Redundancy

When working through the Then Remove Redundancy stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Then Remove Redundancy stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Chunk 1 → page 14
Chunk 2 → page 14
Chunk 3 → page 15
Chunk 4 → page 14
Chunk 5 → page 15
deduplication
document diversity
section diversity
near-duplicate detection
adjacent-chunk merging
token budgeting
chunk 46
chunk 47
chunk 48

Context Construction Is Its Own Layer

The Context Construction Is Its stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

context = "\n\n".join(retrieved_chunks)
Retrieved Evidence
        ↓
Deduplicate
        ↓
Merge related chunks
        ↓
Enforce token budget
        ↓
Preserve source IDs
        ↓
Order evidence
        ↓
Generate context
SOURCE S1
Document: Employee Handbook
Page: 42
Version: 18
Text: ...
SOURCE S2
Document: European Leave Addendum
Page: 7
Version: 4
Text: ...
QUESTION
...

Query Understanding Should Be Cheap

The Query Understanding Should Be stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

metric = revenue
region = Europe
year = 2025
intent = financial comparison
How much did revenue grow in Asia in 2025?
small model
    ↓
classification
small model
    ↓
query rewriting
embedding model
    ↓
retrieval
reranker
    ↓
relevance
large model
    ↓
final reasoning
largest model everywhere

Query Decomposition Helps With Multi-Hop Questions

The Query Decomposition Helps With stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Query Decomposition Helps With stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Question 1:
Which European product had the largest decline?
Question 2:
What explanation did management provide for that product?
Complex Query
                         │
                         ▼
                 Query Decomposer
                    /         \
                   /           \
                  ▼             ▼
              Subquery A    Subquery B
                  │             │
                  ▼             ▼
              Retrieval     Retrieval
                   \           /
                    \         /
                     ▼       ▼
                  Evidence Join
                       │
                       ▼
                      LLM

Freshness Changes the Indexing Strategy

For the Freshness Changes the Indexing stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

Document updated
       │
       ▼
new document version
       │
       ▼
parse changed document
       │
       ▼
generate new chunks
       │
       ▼
embed
       │
       ▼
write new index entries
       │
       ▼
mark previous version inactive
document_id = D17
version 41 → status=inactive
version 42 → status=active
status = active

Model Versioning Matters Too

For the Model Versioning Matters Too stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

embedding_model
embedding_version
embedding_dimensions
chunking_version
parser_version
Documents
                    │
          ┌─────────┴─────────┐
          ▼                   ▼
   embedding model V3   embedding model V4
          │                   │
          ▼                   ▼
      Index V3            Index V4
                              │
                              ▼
                         shadow traffic
                              │
                              ▼
                           evaluate
                              │
                              ▼
                         switch alias

Caching Should Exist at Several Levels

For the Caching Should Exist at stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Caching Should Exist at stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

embedding
→ distributed retrieval
→ reranking
→ large LLM
User Query
    │
    ▼
Exact Response Cache
    │ miss
    ▼
Semantic Cache
    │ miss
    ▼
Retrieval Cache
    │ miss
    ▼
Search
hash(normalized_query)
      ↓
query embedding
Policy version 17
Policy version 18

Failure Handling Matters More Than Happy-Path Latency

When working through the Failure Handling Matters More stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

embedding provider unavailable
one vector shard unavailable
search cluster degraded
reranker timeout
LLM rate limited
LLM provider outage
queue backlog growing
parser crashing on malformed PDF
OCR worker running out of memory
Vector search fails
        ↓
Can lexical search answer?
        ↓
yes
        ↓
return degraded retrieval mode
Primary LLM unavailable
        ↓
Fallback model
        ↓
Lower quality but service continues
Document parser fails 5 times
        ↓
Dead-letter queue
        ↓
record failure reason
        ↓
operator / automated remediation

Backpressure Is Essential

When working through the Backpressure Is Essential stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

20M chunks/hour
25M chunks/hour
200M chunks/hour
embedding service overloaded
        ↓
timeouts
        ↓
retries
        ↓
more load
        ↓
more failures
        ↓
retry storm
incoming workload
      ↓
durable queue
      ↓
workers consume at safe rate
      ↓
backlog temporarily grows

Observability Has to Measure Retrieval Quality

When working through the Observability Has to Measure stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Observability Has to Measure stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

HTTP 200 rate
CPU
memory
disk
p95 latency
error rate
requests/second
request_id: rq_912
query rewrite:
"parental policy?" → "parental leave policy"
vector retrieval:
200 candidates
145 ms
lexical retrieval:
200 candidates
91 ms
fusion:
278 unique candidates
reranking:
278 → 20
212 ms
context:
13 chunks
8,420 tokens
LLM:
input 9,104 tokens
output 681 tokens
1.8 sec
citations:
3 documents
total:
2.4 sec

Evaluate the Retrieval Pipeline Separately From the LLM

The Evaluate the Retrieval Pipeline stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

Retrieval failure

The Retrieval failure stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

Correct document exists
       ↓
retriever misses it
       ↓
LLM never sees it
       ↓
wrong answer

Generation failure

The Generation failure stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Generation failure stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

correct evidence
       ↓
context
       ↓
LLM
       ↓
incorrect interpretation

Latency Should Have a Budget

For the Latency Should Have a stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

API + auth                    50 ms
query understanding          100 ms
retrieval                    300 ms
fusion                        20 ms
reranking                    250 ms
context construction          30 ms
LLM time-to-first-token     1,200 ms
network / safety margin      550 ms
-----------------------------------
target                     ~2,500 ms

Cost Needs the Same Treatment

For the Cost Needs the Same stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

object storage
document parsing
OCR
embedding generation
vector storage
vector replicas
lexical search
metadata database
network transfer
reranking
LLM input tokens
LLM output tokens
observability
backups
cost / 1,000 indexed documents
cost / million chunks
cost / query
cost / successful answer
cost / tenant
Tenant A
5% of revenue
42% of LLM spend

Don’t Put Everything in Expensive Storage

For the Don t Put Everything stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Don t Put Everything stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

expensive / fast
                      ▲
                      │
               vector indexes
               search indexes
               Redis caches
                      │
               metadata DB
                      │
               object storage
                      │
                      ▼
                 cheap / large

Hot and Cold Data May Need Different Treatment

When working through the Hot and Cold Data stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Last 30 days        → 70% of queries
Last 12 months      → 25%
Older archive       → 5%
Query
  │
  ├────→ Hot index
  │
  └────→ Archive index when necessary

Multi-Tenancy Changes Everything

When working through the Multi-Tenancy Changes Everything stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

10,000 organizations
10,000 completely independent vector clusters
Small tenants
        ↓
shared shard groups
Medium tenants
        ↓
partitioned shared infrastructure
Very large tenants
        ↓
dedicated shard groups

The LLM Should Be Near the End of the Architecture

When working through the The LLM Should Be stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the The LLM Should Be stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

User
 ↓
API Gateway
 ↓
Authentication
 ↓
Authorization
 ↓
Rate Limiter
 ↓
Query Understanding
 ↓
Shard Routing
 ↓
Hybrid Candidate Retrieval
 ↓
Rank Fusion
 ↓
Reranker
 ↓
Deduplication
 ↓
Context Builder
 ↓
LLM
 ↓
Citation Validation
 ↓
Response

The Full Production Architecture

The The Full Production Architecture stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

CLIENTS
                                 │
                                 ▼
                         ┌──────────────┐
                         │ API Gateway  │
                         └──────┬───────┘
                                │
                         ┌──────▼───────┐
                         │ Auth / ACL   │
                         │ Rate Limits  │
                         └──────┬───────┘
                                │
                                ▼
                     ┌────────────────────┐
                     │ Query Orchestrator │
                     └─────────┬──────────┘
                               │
                 ┌─────────────┼──────────────┐
                 │             │              │
                 ▼             ▼              ▼
              Cache      Query Rewrite    Metadata
                                           Filters
                 │             │              │
                 └─────────────┼──────────────┘
                               ▼
                        Shard Router
                               │
                  ┌────────────┴────────────┐
                  ▼                         ▼
           Lexical Search              Vector Search
                  │                         │
                  └───────────┬─────────────┘
                              ▼
                         Rank Fusion
                              │
                              ▼
                           Reranker
                              │
                              ▼
                         Deduplicate
                              │
                              ▼
                       Context Builder
                              │
                              ▼
                             LLM
                              │
                              ▼
                     Citation Validation
                              │
                              ▼
                          RESPONSE
================================================================
INGESTION SYSTEM
Documents
                            │
                            ▼
                       Object Store
                            │
                            ▼
                       Event Queue
                            │
               ┌────────────┼────────────┐
               ▼            ▼            ▼
            Parser        Parser       Parser
               │            │            │
               └────────────┼────────────┘
                            ▼
                      Normalization
                            │
                            ▼
                         Chunker
                            │
             ┌──────────────┴──────────────┐
             ▼                             ▼
        Embedding Queue               Text Index Queue
             │                             │
       ┌─────┼─────┐                       ▼
       ▼     ▼     ▼                 Lexical Index
      GPU   GPU   GPU
       │     │     │
       └─────┼─────┘
             ▼
       Vector Index
================================================================
SUPPORTING SYSTEMS
Metadata DB         → document state, versions, ownership
Object Storage      → original source documents
Vector Cluster      → semantic retrieval
Search Cluster      → lexical retrieval
Redis               → caches and rate limiting
Message Queue       → asynchronous ingestion
Observability       → traces, metrics, evaluation
Secrets / IAM       → service authorization
Evaluation System   → retrieval + answer-quality tests

What you Would Not Do

The What you Would Not stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

One giant vector collection with no routing strategy
Synchronous PDF ingestion
Only vector retrieval, no lexical search
ACL filtering after retrieval
No document versioning
No embedding model version
No reranking
Sending top-100 chunks directly to the LLM
Using the LLM for every trivial classification task
No dead-letter queue
No retrieval-level evaluation
No tenant-level cost attribution
Treating the vector DB as permanent document storage
Re-embedding the entire corpus for every small content change
Benchmarking only average latency instead of tail latency

The Core Design Principle

The The Core Design Principle stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The The Core Design Principle stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

question
→ vector search
→ LLM
billions of possible evidence units
              ↓
        routing constraints
              ↓
      relevant partitions
              ↓
      cheap candidate search
              ↓
        hundreds of items
              ↓
      expensive reranking
              ↓
        tens of passages
              ↓
      context optimization
              ↓
             LLM
              ↓
      grounded answer
Cheap operation      → huge search space
Moderate operation   → smaller candidate set
Expensive operation  → tiny candidate set
Most expensive model → final context only

Final Takeaway

For the Final Takeaway stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

Operational checklist

When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.

Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.

Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for 1f03733b8d4b: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.