Home / Articles / Practical notes: Dartboard RAG: When Top-K Returns Three Versions of the Same

This article is published in English.

Practical notes: Dartboard RAG: When Top-K Returns Three Versions of the Same

Operable walkthrough of Practical notes: Dartboard RAG: When Top-K Returns Three Versions of the Same: contracts, checks, and drop-in code slots for teams shipping this pattern.

2352 words

The following notes reconstruct a practical path around “Dartboard RAG: When Top-K Returns Three Versions of the Same Chunk”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing.

The duplicate-context problem

When working through the The duplicate-context problem stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Top 3:
1. "Greenhouse gases trap heat in the atmosphere, causing warming..."  (sim 0.91)
2. "Atmospheric greenhouse gases are the primary driver of climate change..."  (sim 0.89)
3. "The trapping of heat by greenhouse gases leads to rising temperatures..."  (sim 0.88)

The dartboard analogy

When working through the The dartboard analogy stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Chunks plotted by relevance to query
                    (closer to center = higher cosine sim)

                       ●●●         ← cluster of near-duplicate chunks
                      ●●●          (all about "greenhouse gases")
                      ● bull's-eye = QUERY

                                ●     ← chunk about deforestation
                                       (relevant but different topic)

                    ●           ●     ← chunks about agriculture, ocean carbon

                       ●  ●          ← chunks about historical climate


   STANDARD TOP-3 picks:           DARTBOARD TOP-3 picks:
   3 closest darts                 1 closest, then darts that are also
   = 3 darts in the same           good but spread across the board
     spot near bull's-eye          = better coverage of relevant content

The pipeline

When working through the The pipeline stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Step 1 — Over-fetch with FAISS

When working through the Step 1 Over-fetch with stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

fetch_k = self.k * self.oversampling     # default: 5 × 3 = 15
candidates = vector_store.search(query_embedding, k=fetch_k)

Step 2 — Compute distance matrices

When working through the Step 2 Compute distance stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Step 2 Compute distance stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

# Normalize all vectors so dot product = cosine similarity
query_norm = query_vec / np.linalg.norm(query_vec)
cand_norm = candidate_matrix / np.linalg.norm(candidate_matrix, axis=1, keepdims=True)

# Distance = 1 - cosine_similarity
query_distances = 1.0 - np.dot(query_norm, cand_norm.T)        # (1, N)
document_distances = 1.0 - np.dot(cand_norm, cand_norm.T)      # (N, N)

Step 3 — Convert to log-normal probabilities

The Step 3 Convert to stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

def lognorm(dist, sigma):
    return -np.log(sigma) - 0.5 * np.log(2 * np.pi) - dist**2 / (2 * sigma**2)

Step 4 — The greedy selection loop

The Step 4 The greedy stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

# Step 1: pick most relevant first
most_relevant_idx = np.argmax(query_probs)
selected_indices = [most_relevant_idx]
max_distances = doc_probs[most_relevant_idx].copy()    # diversity tracker

# Step 2-6: iteratively add diverse + relevant chunks
while len(selected_indices) < num_results:
    # For each candidate, compute "diversity from any selected"
    updated_distances = np.maximum(max_distances, doc_probs)

    # Combine relevance + diversity
    combined = (diversity_weight * updated_distances
                + relevance_weight * query_probs[np.newaxis, :])

    # Aggregate per candidate (logsumexp for numerical stability)
    normalized = logsumexp(combined, axis=1)

    # Mask already-selected
    for idx in selected_indices:
        normalized[idx] = -np.inf

    # Pick the best
    best_idx = np.argmax(normalized)
    max_distances = updated_distances[best_idx]
    selected_indices.append(best_idx)

What the math is really doing (intuitively)

The What the math is stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The What the math is stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Relevance →
            Low              ●◯

                   ●         ●           ←  Standard top-k picks these 3
                                              (highest relevance, regardless of diversity)
                    ●        ●
                       ●●●●●●  ●  ●  ●     ← Many similar high-relevance chunks
            High         (cluster)


            Diversity ↓
            from
            selected
                        ↓
                     ↓     ↓   ←  Dartboard picks 1 from cluster,
                                 then far-away ones with high relevance still

The weights — what each one does

For the The weights what each stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

relevance_weight = 1.0     # how much we care about chunks being close to query
diversity_weight = 1.0     # how much we care about chunks being different from each other

A worked example: the duplicate corpus test

For the A worked example the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

1. "Greenhouse gases cause warming..."  (sim 0.91)
2. "Greenhouse gases cause warming..."  (sim 0.91)  ← DUPLICATE
3. "Greenhouse gases cause warming..."  (sim 0.91)  ← DUPLICATE
4. "Greenhouse gases cause warming..."  (sim 0.91)  ← DUPLICATE
5. "Greenhouse gases cause warming..."  (sim 0.91)  ← DUPLICATE

Unique results: 1/5
1. "Greenhouse gases cause warming..."  (highest relevance — wins first pick)
2. "Deforestation reduces the carbon sink..."  (different chunk, still relevant)
3. "Industrial agriculture emits methane..."  (third unique cause)
4. "Land-use changes alter surface albedo..."  (fourth unique cause)
5. "Fossil fuel combustion is the largest CO₂ source..."  (related to #1 but different angle)

Unique results: 5/5

The essence in a few lines

For the The essence in a stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

def dartboard_select(query_emb, candidate_embs, k=5, sigma=0.1):
    # 1. Compute distance matrices
    query_dist = 1 - cosine(query_emb, candidate_embs)        # query→each
    doc_dist   = 1 - cosine(candidate_embs, candidate_embs)   # each→each

    # 2. Convert distances to log-probabilities
    query_probs = lognorm(query_dist, sigma)
    doc_probs   = lognorm(doc_dist, sigma)

    # 3. Pick most relevant first
    selected = [np.argmax(query_probs)]
    max_distances = doc_probs[selected[0]].copy()

    # 4. Iteratively add diverse + relevant
    while len(selected) < k:
        updated = np.maximum(max_distances, doc_probs)
        combined = updated + query_probs[np.newaxis, :]   # equal weights = sum
        scores = logsumexp(combined, axis=1)
        for idx in selected:
            scores[idx] = -np.inf       # don't re-select

        best = np.argmax(scores)
        max_distances = updated[best]
        selected.append(best)

    return selected

For the The essence in a stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Knobs you might turn

When working through the Knobs you might turn stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Where this earns its keep, and where it doesn’t

When working through the Where this earns its stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

The bigger idea worth taking with you

When working through the The bigger idea worth stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the The bigger idea worth stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

A Final Thought

The A Final Thought stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

Operational checklist

For the Operational checklist stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.

Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.

Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for fd4991fea9d9: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.