Home / Articles / When citations look perfect and the headcount is still wrong

This article is published in English.

When citations look perfect and the headcount is still wrong

A salary RAG cited the right PDF and still undercounted twelve employees as one—hybrid search, table-aware compression, and pipeline traces fixed it.

3139 words

A payroll question should be boring. Ask how many people appear in the April 2025 salary register and you expect a count, a citation, and nothing clever. The stack returned a polished sentence: one employee, source PDF named, page stamped. The register listed twelve.

That failure mode is what makes retrieval-augmented generation dangerous in practice. The process did not throw. Logs stayed quiet. The answer looked like every correct answer from the same week. Soft failure with a footnote is worse than a hard crash, because nobody escalates a polite paragraph.

This write-up reconstructs the pipeline that produced the lie, the diagnostics that exposed it, and the concrete changes that brought the count back to twelve. The stack is deliberately ordinary: pdfplumber for layout-aware PDF text, a LangChain text splitter only for chunking, MiniLM embeddings, Qdrant for vectors, a hybrid dense-plus-BM25 retrieval mix, a cross-encoder reranker, and a compressor that was supposed to shrink context before the generator ran. None of those choices were exotic. The bugs lived in how they interacted on tables and identifiers.

Ingestion once, answering every time

Documents land once. Questions arrive forever. Keeping those paths separate matters when you debug, because a broken answer can come from stale vectors, bad filters, or a generator that never saw the right rows.

Ingestion walks each PDF, extracts text with layout awareness, splits into overlapping chunks, embeds, and upserts into Qdrant with metadata that later filters can use:

PDF → extract text per page → cut into chunks
    → convert each chunk to 384 numbers → store in a vector database

Answering is a different graph. Embed the question, retrieve candidates, optionally blend keyword hits, rerank, compress, then ask the model with citations attached:

question → search → rerank → compress → ask the LLM → cited answer
              ↓         ↓          ↓
          5 chunks  3 chunks   ~600 chars

On paper that flow is textbook. In production the textbook assumptions fail on employee registers, invoice tables, and any document where the answer is a structure rather than a slogan.

Why each piece earned its seat

pdfplumber is not the fastest PDF library. It is one of the few that preserves enough layout signal to keep salary rows from collapsing into a single paragraph soup. Faster extractors that flatten tables into run-on sentences make later “count the employees” questions nearly impossible, because the model never sees row boundaries.

Chunking used LangChain’s recursive character splitter and nothing else from the framework for this project. The goal was predictable window sizes with overlap, not an opinionated RAG chain. Embeddings settled on all-MiniLM-L6-v2 after a short bake-off: smaller models missed identifier tokens; larger models added latency without fixing the failure that mattered. Qdrant won for local-to-cloud continuity and payload filters that matched how the app already thought about document types.

Hybrid search mixed semantic similarity with BM25-style keyword scoring at roughly 0.65 / 0.35. Pure semantic search was excellent at “what is parental leave policy” and terrible at STL/2025-26/003. Reranking with a cross-encoder cleaned the shortlist. Compression was meant to keep only high-signal sentences before the LLM call. That last stage is where the twelve-became-one story really started.

The confident wrong answer

The salary question retrieved a plausible page, cited it, and undercounted. Looking at raw retrieval scores made the first crack visible: semantic neighbors looked “close enough” while the exact register rows competed poorly against narrative HR prose elsewhere in the corpus.

0.781  Invoice_INV2025001_Arvind.pdf     <- returned
0.779  Invoice_INV2025002_Trident.pdf
0.776  Invoice_INV2025003_Welspun.pdf    <- actually wanted

Meaning similarity cannot be asked to do exact token work. Adding BM25 and blending scores changed the ranking so the identifier and table-bearing chunks rose:

0.854  Invoice_INV2025003_Welspun.pdf    <- correct
0.549  Invoice_INV2025001_Arvind.pdf
0.545  Invoice_INV2025002_Trident.pdf

That alone did not restore the count. It only put the right document in play. The next traps lived inside the document.

Employee IDs as a survival experiment

Rather than argue about vibes, the employee IDs ST001 through ST012 were traced through each stage: extraction, chunking, retrieval, rerank, compression, and final prompt assembly.

retrieved chunk    ST001 ST002 ST003 ST004 ST005 ... ST012   (12 present)
after rerank       ST001 ST002 ST003 ST004 ST005 ... ST012   (12 present)
after compression  ST001 ST002                                (2 present)
prompt sent        ST001 ST002                                (2 present)

Several IDs vanished between compression and the model. The compressor kept a fixed budget of “top sentences.” A table row counts as a sentence. A dense salary table burns the budget on early rows while later rows—and the prose that explains them—never enter the prompt. The model then honestly reports whatever subset it saw.

# before — keep the top N
best = scored[:40]
# after — keep the best until the character budget runs out
budget = MAX_CONTEXT_CHARS // TOP_K_AFTER_RERANK
for sentence, score in ranked:
    if best and len(sentence) + 2 > budget:
        continue
    best.append(sentence)
    budget -= len(sentence) + 2

Once compression shuffled and truncated rows, the salary table arrived as a deck of cards dealt out of order. Every number might exist somewhere in the context window, yet the structure required to count distinct employees was gone. Counting people in a table someone has cut into ranked snippets is a different task from reading the table.

Reranker scores that looked fine and still hurt

Cross-encoder scores around the critical rows looked “reasonable.” Reasonable is not the same as preserving a complete table.

0.446  keep   ST001 Rajesh Kumar M CFO 75,000 ...
0.323  keep   ST002 Priya Sundaram Finance Mgr 45,000 ...
0.420  keep   ST003 Murugan K Sr Accountant 28,000 ...
0.259  DROP   ST005 Selvam R Production Mgr 38,000 ...   ← below 0.30
0.422  keep   ST006 Karthik P Supervisor 18,000 ...
# pick the best by score…
chosen = set(id(p) for p in best)
# …then put them back in document order before joining
best = [p for p in scored if id(p) in chosen]

The repair was not “raise the threshold everywhere.” It was to detect table-like lines and exempt them from aggressive compression. A practical heuristic: if a line contains four or more numbers and sits among at least three similar lines in the same chunk, treat it as a table row and keep it regardless of the sentence score.

ROW_MIN_NUMBERS = 4
TABULAR_MIN_ROWS = 3
def is_row(sentence):
    return len(NUMBER_RE.findall(sentence)) >= ROW_MIN_NUMBERSif looks_tabular(sentences):
    survivors = [p for p in scored if is_row(p[0]) or p[1] >= threshold]
else:
    survivors = [p for p in scored if p[1] >= threshold]

With table rows exempted, all twelve IDs survived into the prompt on the next run. The generator finally had a structure it could count.

Small fixes that removed other landmines

While the table path was open, a few adjacent issues got the same treatment: empty environment variables that silently replaced intended defaults, document-type filters that behaved differently between local Qdrant and Qdrant Cloud, and answer payloads that returned only prose when operators needed the pipeline trace.

# the context cap must be at least CHUNK_SIZE × TOP_K_AFTER_RERANK
MAX_CONTEXT_CHARS = 5000   # was 2000 while 1200 × 3 = 3600

Every answer now returns the pipeline, not only the final sentence—retrieved IDs, scores, compression decisions, and filters—so the next wrong count can be dissected without guessing:

{
  "answer": "12 employees are listed...",
  "citations": [{"doc": "Salary_Register_April_2025.pdf", "page": 1}],
  "pipeline": {
    "retrieved": 5,
    "after_rerank": 3,
    "chars_before_compression": 3847,
    "chars_after_compression": 2103
  }
}

Filtering by document type worked against a local Qdrant instance and failed the same way in Qdrant Cloud until the payload indexes and filter shapes matched what the cloud deployment actually indexed:

500 — Index required but not found for "doc_type"

Related footgun in Python: os.getenv("KEY", "default") returns an empty string when the variable exists but is blank. That is not the default. Prefer or when empty must fall through:

QDRANT_URL = os.getenv("QDRANT_URL") or "http://localhost:6333"

Where the stack landed

After hybrid retrieval, table-aware compression, and explicit pipeline telemetry, the April register question returned twelve with citations that pointed at the rows that justified the count.

Retrieval:    hybrid, 0.65 semantic + 0.35 BM25
Reranking:    cross-encoder, 5 in → 3 out
Compression:  sentence-level, table-aware, ~35% reduction
Grounding:    prompt + threshold + temperature 0.0
Evaluation:   13/15 on the golden set, refusals tested
Latency:      ~4s end to end
Cost:         ~$0.003 per fifteen-question run

None of this required a new foundation model. It required treating RAG as a data pipeline with measurable survival rates for the tokens that matter. Tutorials sell “embed, retrieve, generate.” Production sells “prove the rows reached the prompt.”

Operational habits that keep the count honest

Keep a golden question set that includes at least one pure count over a table, one exact identifier lookup, and one semantic policy question. If only the semantic questions pass, the demo is lying about readiness.

Log chunk IDs through compression. If an ID present after retrieval is absent after compression, that is a bug even when the final answer happens to be right today.

Prefer explicit hybrids for corpora that mix prose and registers. Dense retrieval alone will keep winning on essays and losing on codes.

Treat citations as necessary but not sufficient. A page number next to a wrong headcount is still a wrong headcount.

When moving from local Qdrant to cloud, re-verify payload indexes and filter behavior with the same fixtures. API compatibility does not imply identical indexing defaults.

Budget context for structure, not only for “important sentences.” Tables are structure. Compressing them like blog paragraphs is how twelve becomes one.

Why soft failures deserve harder gates

Teams often gate RAG launches on “answers look good in a spreadsheet of happy-path prompts.” That bar misses the failure that matters in payroll, inventory, and compliance: fluent undercounts. Add automated checks that assert extracted entity sets survive into the prompt for known fixtures. Fail the build when ST001–ST012 do not all appear after compression on the salary fixture.

The same gate can assert hybrid blend weights are actually enabled in the running config, not only in a README. Drift between notebook experiments and the service config is a common way BM25 “disappears” after a refactor.

Finally, teach operators to distrust pretty citations. The product UI should surface retrieval and compression traces next to the answer for internal users until the golden set stays green for a release cycle. External users can keep the clean paragraph; internal users need the autopsy kit.

Closing

The register never said there was one employee. The pipeline did, after it quietly discarded the rows that would have corrected it. Fixing that meant hybrid search for identifiers, table-aware compression for structure, and answers that carry their own evidence trail. Once those pieces landed, twelve stayed twelve—and the next confident citation had something real to stand on.

That is the bar worth keeping: not whether the model can sound sure, but whether the system can prove the facts that justify the certainty.

A deeper look at hybrid scoring in practice

Dense retrieval encodes the question and each chunk into the same vector space and ranks by cosine or dot product. That works when the user’s language matches the document’s language: parental leave, severance, remote work policy. It fails when the user pastes an opaque token that barely appears in training data and barely co-occurs with nearby words in the embedding space. Employee codes, invoice numbers, and contract IDs are exactly that class of token.

BM25 and related lexical scorers invert the problem. They care that the token appears, how rare it is in the corpus, and how often it appears in the candidate chunk. They do not care that “STL/2025-26/003” is semantically near nothing. Blending the two scores is not philosophical; it is acknowledging that payroll corpora contain both essay-like policies and registry-like tables.

The 0.65 / 0.35 mix used here is a starting point, not a law of nature. Corpora heavy on codes may need more lexical weight. Corpora heavy on narrative may need less. What matters is measuring both classes of question in the golden set and refusing to ship a blend that only passes narrative prompts.

When implementing the blend, normalize scores before mixing. Raw dense similarities and raw BM25 scores live on different scales. Teams that add unnormalized numbers often discover that one channel dominates by accident. Min-max or rank fusion (RRF) are both acceptable if they are tested against fixtures that include exact IDs.

Compression as a lossy codec for tables

Context compression exists because models have finite windows and because irrelevant sentences dilute attention. The default story is: score sentences for relevance, keep the top k, discard the rest. That story assumes sentences are interchangeable units of meaning. Table rows are not interchangeable units of meaning. Row 7 without rows 1–6 is not a slightly worse table; it is a broken table.

The exemption heuristic—many numbers on a line, several such lines nearby—is intentionally crude. It will keep some non-table lines that look numeric, and that is acceptable. False positives cost tokens. False negatives cost correctness. For payroll and inventory, correctness wins.

More elaborate detectors can use PDF structure from pdfplumber: cell bounding boxes, aligned columns, repeated x-coordinates. Those signals are excellent when available. The sentence heuristic remains useful as a fallback when the text has already been flattened into markdown or plain text before compression runs.

Another failure mode is shuffling. Even if all rows survive, reordering them by relevance score can destroy totals and counts. Prefer stable document order for exempted table rows. Rank prose freely; keep tables in reading order.

Citations that defend the wrong answer

Citation UX often highlights the PDF and page that contributed the most retrieved chunk. When compression later drops half the table, the citation still points at the right file. Users read that as confirmation. Product design should either cite the specific rows used in the final prompt or show a retrieval trace. Otherwise the interface becomes an accomplice.

For regulated domains, store the exact prompt context hash with the answer. If an auditor asks why the system said “one employee,” you can replay the truncated table and show the bug rather than debating model temperament.

Local versus cloud vector databases

The document-type filter that worked locally and failed in Qdrant Cloud is a reminder that “same API” is not “same index configuration.” Payload indexes, keyword filters, and null handling differ across deployment modes. Fixture tests should run against the same class of deployment you productionize. A green local suite with a red cloud filter is how silent empty retrievals ship.

Empty environment variables deserve the same paranoia. Configuration systems that export blank strings for unset secrets will bypass defaults in languages where empty is truthy enough to win. Normalize config at process start: treat blank as missing, then apply defaults, then fail closed if required keys are still absent.

Measuring survival, not vibes

The employee-ID survival experiment is the template. Pick entities that must appear for a correct answer. Assert they exist after extraction, after chunking, after retrieval, after rerank, and after compression. Plot drop-offs. The first cliff is usually the bug.

Extend the idea to numeric aggregates. If the question is a sum, assert that every addend row reached the prompt. If the question is a distinct count, assert the distinct key set is complete. These checks are cheap compared with production incidents.

What not to blame first

It is tempting to blame the LLM when the count is wrong. Sometimes the model truly cannot count. More often the model never received a countable structure. Swap models only after the survival plot is green. Otherwise you will “fix” a counting bug by moving to a larger model that hallucinates the number you wanted—until the next register arrives.

Similarly, resist ripping out the whole framework because one compressor misbehaved. Isolate the stage, add the exemption, add the test, and move on. Broad rewrites feel productive and often reintroduce the same class of bug under new names.

A minimal hardening checklist

  1. Golden questions: table count, exact ID, semantic policy.
  2. Hybrid retrieval with normalized blend or RRF.
  3. Table-aware compression with stable row order.
  4. Pipeline traces on every internal answer.
  5. Config normalization that treats blank env vars as missing.
  6. Cloud index fixtures that mirror production filters.
  7. Build gates on entity survival for known documents.
  8. Citations tied to prompt contents, not only to retrieved files.

Run that checklist before calling a payroll RAG “done.” The April register will not be the last document that looks simple and fails politely.

Aftermath for the team

Once the twelve-employee answer stabilized, the same telemetry caught two quieter bugs: a stale collection alias after a re-ingest, and a reranker timeout that fell back to unre-ranked hybrid hits without marking the answer as degraded. Both would have been invisible if the API returned only a string. Returning structure—scores, timings, fallbacks—turned the assistant into something operators could trust enough to debug at 2 a.m.

That is the real product lesson. RAG systems are not chat skins over PDFs. They are data pipelines that speak in paragraphs. Treat them like pipelines: measure drops, preserve structure, and never let a citation substitute for proof.

One more pass on the original wrong answer

Replaying the bad response with traces attached makes the story almost dull. Retrieval nearly had the right PDF. Hybrid scoring finished that job. Compression then spent its sentence budget on the first numeric rows and dropped the rest. The model counted what remained and cited the file. Every stage did something defensible in isolation. Together they manufactured confidence.

That is why stage-local metrics mislead. Retrieval@k can look healthy while compression kills the answer. End-to-end entity survival is the metric that matches user harm. Adopt it early, especially when documents are tables wearing PDF clothing.

If you take nothing else from the twelve-employee incident, take this: polite RAG failures are pipeline bugs until proven otherwise. Instrument the path, protect structure, and make the system show its work before you trust its voice. Soft failures deserve hard gates, repeated fixtures, and operators who can see every drop before users ever trust a cited headcount again.

Entity survival through compression remains the simplest honest gate for table-heavy corpora.

When the next register arrives, rerun the same ID survival plot before trusting any newly confident citation again.