This article is published in English.
Why RAG answers look complete when Azure DevOps search proves they are not
Closing completeness gaps in a graph-backed Azure DevOps RAG system: deterministic tag checks, graph walks, context ceilings, and read-only query tools.
A system that loads Azure DevOps history into a graph database can answer questions about projects, ownership, and shipped work with readable prose and clickable citations. For a while that looked sufficient—until answers were checked against a blunt baseline: a raw Azure DevOps query. The gap between a polished reply and an exhaustive list is where “complete” retrieval quietly fails. Closing that gap required deterministic steps, measured tests, and several fixes that created new problems of their own.
The question that exposed the gap
Early questions produced clean, sourced answers with little invention. A broader prompt—“What AI work is being done in this org?”—still looked fine: a few paragraphs, a handful of citations, nothing obviously wrong. The same question run through az boards query, a brute-force search with no ranking intelligence, returned dozens of items the polished answer never mentioned. Clusters about evaluating coding assistants, comparing model costs, or debugging a specific tool were absent because they shared little wording with the question and never ranked near the top of semantic search.
That failure mode is easy to miss: a well-ranked answer is not the same as a complete answer, and the system offers no signal telling you which one you received.
A prompt that only sometimes helped
The first impulse was to instruct the model to be more careful—on broad questions, inspect category tags before answering instead of trusting search alone. It helped intermittently. The identical wording produced different behaviour across runs: sometimes the model checked tags, sometimes it skipped them. Nothing in the question changed—only whether the instruction was followed that turn.
An instruction to a language model is a nudge, not a guarantee. When completeness matters, thoroughness cannot depend on the model deciding, in the moment, whether to be thorough.
Making the path deterministic
The next change stopped asking. Code always checks tags on every broad question, regardless of whether the model believes it is necessary. That alone moved a hard ground-truth cluster from roughly half found to about seven of sixteen items—better, still incomplete. Further steps came from concrete gaps found in testing, not speculation:
A step that mines tag-search hits for repeated names and phrases appearing two or three times, then searches those names directly—so a product name surfaces even when nobody queried it explicitly, once the small cluster that contains it is already in view.
A step that walks graph relationships, not only wording, because sibling work items can share a parent while sharing zero title words. An item titled like a seeded test repository with representative coding tasks may say nothing about AI in its own text and only connect through hierarchy.
A step for repositories, which do not carry the same category tags as work items and would otherwise stay invisible to tag-driven paths.
A step for people, after questions like how often two collaborators worked together produced confidently wrong answers.
Each closed a measured gap. Eventually the sixteen-item hardest case resolved completely—not mostly, entirely.
More retrieved context made answers worse
Counterintuitively, once retrieval improved enough that the model’s pool grew from a few hundred items to over a thousand, answers became shorter and less complete. The model did not crash; it quietly wrote less despite having more true material available. Past a volume threshold, answer quality does not scale with context size—it degrades. Retrieval metrics improved while final-answer metrics declined in the same run. The remedy was not “add more,” but finding the ceiling and staying under it.
A transparency fix that flooded narrow questions
When the model summarized away real material instead of naming it, a new answer section listed items found but not named, so nothing vanished silently. The first version dumped more than a thousand loosely related items into replies for narrow questions, because the list generator could not tell precise matches from generic tag noise. A broad category pulled everything remotely related and treated it as equally worth surfacing.
The durable fix was not only a numeric cap. It was separating two kinds of “found but unnamed”: a small set of precise matches that should always appear, and a large pool of loosely tagged noise that needs an honest cap plus a note that more exists. That distinction mattered more than the cap itself.
Giving the model tools—with guardrails
A useful question shaped the rest of the work: why can a human answer what the system cannot when both see the same data? Humans can write a new query when fixed tools misfit, and they double-check answers that feel off. The model had neither ability. It received a tool to write its own read-only database query—enforced read-only by the database, time-limited, and capped in result size.
Two more rounds were required. Asked how often two named people collaborated, the model first assumed a shared surname from co-mention, got an empty result, and guessed from unrelated comments instead of questioning the empty set. The rule became: resolve each name to its exact record first, and treat an empty self-query as evidence the query is wrong, not that the count is zero.
On the next attempt it resolved both people, ran the right query, obtained twenty-five shared items—and still did not state the number, because every factual claim needed a citation id and a computed count had none. Following the letter of the citation rule discarded a true answer. An explicit exception was required: a computed number may be stated plainly without a source id.
Neither failure was incapability; both were correct obedience to instructions that did not cover the situation. That distinction matters more than it first appears.
How testing stayed honest
Ground truth was not a vibe check. For the hardest cluster, sixteen known items were listed up front from the exhaustive query, then scored after each retrieval change: how many appeared in the final answer, how many were cited, how many were silently dropped. That scoreboard is what made “seven of sixteen” and later “sixteen of sixteen” meaningful. Without an external exhaustive list, “looks complete” is only aesthetics.
Ordering the pipeline without drowning the model
Deterministic steps still need an order. Tag expansion first widens the candidate set; name mining then deepens it; graph walks add structural neighbours; repository and people steps fill modality gaps. Only after that pool is built does a size ceiling trim what enters the prompt. Reversing that order—asking the model to be thorough before the pool is built—returns to the nudge problem. Completeness work is pipeline design, not better adjectives in the system prompt.
Citations versus computed facts
The citation-every-claim rule protects against invention when the model quotes work items. It becomes harmful when the model performs arithmetic or joins over tool results. Separating “quoted claim” from “computed aggregate” in the instructions restored the twenty-five-collaboration count without weakening citation discipline for narrative claims. Tool-using RAG systems need both rules, stated explicitly.
Where the work stands
A raw database query is exhaustive by construction: matching rows cannot be missed. It also cannot explain, group, or narrate meaning—it returns a list, not an answer. The system now matches that query’s completeness on the hard cases that were found and tested, while still explaining findings in prose with verifiable sources. That result is measured, not assumed.
It is not a solved problem. Every fix above exists because a specific failure appeared when answers were checked against something exhaustive. That method finds gaps; it does not prove none remain. After the changes above, a new question type still surfaced a citation used to support a claim about a person when the cited source said nothing about that person—found, not yet fixed at the time of writing. The pattern continues: something that looks complete is checked, and something new appears.
The honest summary is not “RAG completeness is solved.” It is that every specific gap that could be found and tested was closed, with before-and-after numbers. There is no way to promise there is no next gap, and no system of this kind should imply otherwise. “The model usually gets it right” is not a foundation for reliability—usually is exactly the word that fails on the question that mattered.