实用笔记:RAG中的5种重排技术——从快速检索到精准排序
《实用笔记》操作指南:RAG中的5种重排序技术——从快速检索到精准匹配;同时为采用该模式的团队提供模板、检查清单及可直接使用的代码片段。
本指南将逐步构建从原始材料到可运行系统的完整流程,主题为:RAG中的5种重排序技术:从快速检索到精准上下文。重点在于可操作的步骤、明确的检查点以及可直接放入代码库的代码,无需猜测其用途。 在概览阶段,应在修改代码之前明确输入参数、各步骤的负责人以及完成标准。操作人员应能够从已知的检查点重新运行该步骤,而无需推测隐藏状态。 除了功能结果外,还需记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在流程从演示环境转向共享环境时出现意外费用。
无人提及的检索瓶颈
在处理“检索瓶颈无人阶段”时,首先需列出相关约定:所需输入、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 应将配置信息与应用程序代码分开存放。环境文件、密钥存储以及功能开关应集中管理,以便操作人员无需查看整个系统结构即可进行审计。 在调整提示词之前,需先用固定的问题集来衡量召回率。仅仅更换提示词很难解决检索效果不佳的问题。
什么是重排序?
在处理“什么是重排”这一阶段时,首先写下相关约定:所需的输入参数、成功标志以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持一致性。 同时记录正常流程和异常恢复流程。重试机制、人工审核环节以及错误消息处理都是产品本身的组成部分,而非后续需要优化的内容。 在调整提示词之前,先使用固定的问题集来衡量召回率。仅仅更换提示词很难解决检索效果不佳的问题。
检索与重排
在处理检索与重排阶段时,首先明确相关约定:所需输入、成功标志以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持一致性。 优先选择小型、可测试的单元,而非庞大的脚本。当某个步骤失败时,故障应指向单一责任点,而非复杂的流程链。 在调整提示词之前,先使用固定的问题集来衡量召回率。仅仅更换提示词很难解决检索效果不佳的问题。 在处理检索与重排阶段时,首先明确相关约定:所需输入、成功标志以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持一致性。 在功能结果之外,还需记录执行时间以及令牌或查询成本。提前了解成本情况,可避免从演示环境过渡到共享环境时出现意外费用。
| Aspect | Initial Retrieval | Reranking |
| ------------------- | ------------------------ | ----------------------------- |
| Goal | Find candidates fast | Judge true relevance |
| Speed | Milliseconds | Tens to hundreds of milliseconds |
| Input | Query + index | Query + top-k candidates |
| Scoring depth | Shallow (embedding dot product) | Deep (cross-attention, token interaction) |
| Cost | Low (local compute) | Higher (model inference) |
| When to use | Every query | On top-k candidates only |
五种重排序技术
将“五种重排序技术”这一阶段视为可度量的框架来使用效果最佳。在扩大范围之前,先记录一个成功的案例、一个失败案例以及回滚说明。 将配置置于应用程序代码之外。环境文件、密钥存储和功能标志应集中存放于一个位置,以便操作人员无需查看整个系统结构即可进行审计。 将分块策略与检索策略分开。当质量指标发生变化时,修改其中一项不应迫使重新编写另一项。
1. 交叉编码器重排序
将1个交叉编码器重排序阶段视为可度量的模型表面时,其效果最佳。在扩大范围之前,需记录一个成功的处理案例、一个失败案例以及回滚说明。同时记录正常流程与故障恢复流程。重试机制、人工审核环节以及死信处理都是产品本身的组成部分,而非后续需要补充的功能。应将分块策略与检索策略分开,当质量指标发生变化时,修改其中一项不应强制要求重新编写另一项。
from sentence_transformers import CrossEncoder
# Load a cross-encoder reranker
# ms-marco-MiniLM-L-6-v2 is fast and accurate for general use
cross_encoder = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2")
def rerank_with_cross_encoder(query: str, retrieved_docs: list[str], top_k: int = 5):
"""
Rerank retrieved documents using a cross-encoder.
Args:
query: The user question
retrieved_docs: List of document chunks from initial retrieval
top_k: Number of documents to return after reranking
Returns:
List of (document, score) tuples, sorted by relevance
"""
# Create query-document pairs
pairs = [[query, doc] for doc in retrieved_docs]
# Get relevance scores
scores = cross_encoder.predict(pairs)
# Combine docs with scores and sort
scored_docs = list(zip(retrieved_docs, scores))
scored_docs.sort(key=lambda x: x[1], reverse=True)
return scored_docs[:top_k]
# Example usage
query = "What are the side effects of amoxicillin?"
retrieved = [
"Amoxicillin is a penicillin antibiotic used to treat bacterial infections.",
"Common side effects include nausea, vomiting, and diarrhea.",
"The drug was first discovered in 1958 by researchers at Beecham.",
"Patients with penicillin allergies should avoid amoxicillin.",
"Side effects may include rash, itching, and in rare cases, anaphylaxis.",
]
top_docs = rerank_with_cross_encoder(query, retrieved, top_k=3)
for doc, score in top_docs:
print(f"Score: {score:.4f} | {doc}")
2. 互斥排名融合(RRF)
将“双向排名融合”阶段视为可度量的对象时,其效果最佳。在扩大范围之前,需记录一份理想案例、一个故障实例以及回滚说明。 相较于庞大的脚本,应优先选择小型且可测试的单元。当某一步骤出现故障时,故障原因应能明确指向某个具体责任方,而非复杂的流程链。 应将分块策略与检索策略分开。当质量指标发生变化时,修改其中一项不应迫使重新编写另一项。 将“双向排名融合”阶段视为可度量的对象时,其效果最佳。在扩大范围之前,需记录一份理想案例、一个故障实例以及回滚说明。 除了功能结果外,还需记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外费用。
def reciprocal_rank_fusion(rankings: list[list[str]], k: int = 60) -> list[tuple[str, float]]:
"""
Merge multiple document rankings using Reciprocal Rank Fusion.
Args:
rankings: List of rankings, where each ranking is a list of document IDs
ordered from most to least relevant
k: RRF constant (default 60, as recommended in the original paper)
Returns:
List of (document_id, rrf_score) tuples, sorted by fused score
"""
scores = {}
for ranking in rankings:
for rank, doc_id in enumerate(ranking, start=1):
if doc_id not in scores:
scores[doc_id] = 0.0
# RRF formula: 1 / (k + rank)
scores[doc_id] += 1.0 / (k + rank)
# Sort by score descending
return sorted(scores.items(), key=lambda x: x[1], reverse=True)
# Example: merging BM25 and vector search results
bm25_results = ["doc_5", "doc_2", "doc_8", "doc_1", "doc_9"]
vector_results = ["doc_1", "doc_5", "doc_3", "doc_8", "doc_7"]
fused = reciprocal_rank_fusion([bm25_results, vector_results])
print("Fused ranking:")
for doc_id, score in fused:
print(f" {doc_id}: {score:.4f}")
# Notice: doc_5 and doc_1 appear in both retrievers and get boosted to the top
3. Cohere Rerank API
对于3个Cohere Rerank API阶段,在修改代码之前需明确输入参数、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 配置信息应置于应用程序代码之外。环境文件、密钥存储以及功能标志应集中存放于一个位置,以便操作人员无需查看整个系统结构即可进行审核。 需注明实际作为答案依据的段落。如果没有引用,操作人员就无法区分是幻觉内容还是索引缺失导致的错误。
import cohere
from dotenv import load_dotenv
import os
load_dotenv()
# Initialize Cohere client
co = cohere.Client(os.getenv("COHERE_API_KEY"))
def rerank_with_cohere(query: str, documents: list[str], top_k: int = 5):
"""
Rerank documents using Cohere's managed Rerank API.
Args:
query: The user question
documents: List of document chunks from initial retrieval
top_k: Number of documents to return
Returns:
List of (document, relevance_score) tuples
"""
response = co.rerank(
model="rerank-v3.5",
query=query,
documents=documents,
top_n=top_k,
return_documents=True
)
results = []
for result in response.results:
results.append((
result.document.text,
result.relevance_score
))
return results
# Example usage
query = "How do I handle authentication in a FastAPI app?"
docs = [
"FastAPI is a modern web framework for building APIs with Python.",
"To add authentication, use OAuth2PasswordBearer and JWT tokens.",
"Pydantic models in FastAPI provide automatic request validation.",
"The OAuth2PasswordBearer class expects a token URL endpoint.",
"FastAPI was created by Sebastián Ramírez and released in 2018.",
]
ranked = rerank_with_cohere(query, docs, top_k=3)
for doc, score in ranked:
print(f"Score: {score:.4f} | {doc}")
4. ColBERT
在4 ColBERT阶段,修改代码之前需明确输入内容、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 需同时记录正常流程和异常恢复流程。重试机制、人工审核环节以及错误处理都是产品本身的组成部分,而非后续需要补充的功能。 必须引用实际作为答案依据的段落。没有引用的话,操作人员就无法区分幻觉内容与索引缺失问题。
from colbert import Searcher
from colbert.infra import Run, RunConfig
def setup_colbert_searcher(index_path: str, checkpoint: str):
"""
Initialize a ColBERT searcher for late-interaction reranking.
Args:
index_path: Path to the pre-built ColBERT index
checkpoint: Path to the ColBERT model checkpoint
Returns:
Configured Searcher instance
"""
with Run().context(RunConfig(nranks=1, experiment="reranking")):
searcher = Searcher(
index=index_path,
checkpoint=checkpoint
)
return searcher
def rerank_with_colbert(searcher, query: str, doc_ids: list[str], top_k: int = 5):
"""
Rerank documents using ColBERT's late interaction.
Args:
searcher: Initialized ColBERT Searcher
query: The user question
doc_ids: List of document IDs from initial retrieval
top_k: Number of documents to return
Returns:
List of (doc_id, score) tuples
"""
# Search within the candidate set
results = searcher.search(
query,
k=top_k,
filter_fn=lambda pid: pid in doc_ids # Only rerank candidates
)
return list(zip(results[0], results[2])) # doc_ids, scores
# Note: ColBERT requires a pre-built index and model checkpoint.
# For production use, build the index once and load it at startup.
5. 以LLM作为评判者
在“5 LLM作为裁判”阶段,修改代码之前需明确输入内容、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 相较于庞大的脚本,更应采用小型且可测试的单元。当某一步骤失败时,故障原因应能指向单一责任主体,而非复杂的流程链。 若后续步骤为代码调用或工具调用,相比自由形式的文字描述,带有结构化格式且经过模式验证的输出更为合适。 在“5 LLM作为裁判”阶段,修改代码之前需明确输入内容、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 除了功能结果外,还需记录执行时间以及token或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外费用。
You are evaluating documents for a retrieval system.
Query: {query}
Document: {document}
Rate how relevant this document is for answering the query.
Respond with a single integer from 1 to 10, where 10 means perfectly relevant.
Relevance score:
from openai import OpenAI
import os
from dotenv import load_dotenv
load_dotenv()
client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
def score_document_with_llm(query: str, document: str) -> int:
"""
Ask an LLM to score a document's relevance to a query.
Args:
query: The user question
document: A candidate document chunk
Returns:
Integer relevance score from 1-10
"""
prompt = f"""You are evaluating documents for a retrieval system.
Query: {query}
Document: {document}
Rate how relevant this document is for answering the query.
Respond with a single integer from 1 to 10, where 10 means perfectly relevant.
Be strict: only give high scores to documents that directly help answer the query.
Relevance score:"""
response = client.chat.completions.create(
model="gpt-4.1-mini",
messages=[{"role": "user", "content": prompt}],
temperature=0,
max_tokens=5
)
try:
score = int(response.choices[0].message.content.strip())
return max(1, min(10, score)) # Clamp to 1-10
except ValueError:
return 5 # Default on parse failure
def rerank_with_llm_judge(query: str, documents: list[str], top_k: int = 3):
"""
Rerank documents using an LLM as a relevance judge.
Args:
query: The user question
documents: List of candidate document chunks
top_k: Number of documents to return
Returns:
List of (document, score) tuples, sorted by relevance
"""
scored = []
for doc in documents:
score = score_document_with_llm(query, doc)
scored.append((doc, score))
scored.sort(key=lambda x: x[1], reverse=True)
return scored[:top_k]
# Example usage
query = "What are the tax implications of RSU vesting for employees in California?"
docs = [
"RSUs are restricted stock units granted to employees as part of compensation.",
"In California, RSU income is taxed as ordinary income at vesting, not at grant.",
"Employers typically withhold federal and state taxes at vesting time.",
"Stock options and RSUs have different tax treatments under IRS rules.",
"California has one of the highest state income tax rates in the US.",
]
ranked = rerank_with_llm_judge(query, docs, top_k=3)
for doc, score in ranked:
print(f"Score: {score}/10 | {doc}")
应该选择哪一个?
在确定应使用哪种方案时,首先需列出相关契约:所需的输入参数、成功信号以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持一致性。 应将配置信息与应用程序代码分开存放。环境文件、密钥存储以及功能开关应集中管理,这样操作人员无需查看整个系统结构即可进行审计。 在调整提示词之前,需先使用固定的问题集来测试召回率。仅仅更换提示词往往无法解决检索效果不佳的问题。
| Technique | Best For | Latency | Cost |
| --------------------- | ------------------------------------------------- | ------------ | -------------- |
| Cross-Encoder | Maximum quality on top-k candidates | 50-200ms | Local GPU/CPU |
| RRF | Hybrid retrieval without adding model inference | ~0ms | Free |
| Cohere Rerank API | Speed without operational overhead | 100-300ms | Per API call |
| ColBERT | Large-scale, low-latency use cases | 20-100ms | Index + GPU |
| LLM-as-a-Judge | Complex, high-value queries (medical, legal) | 1-5 seconds | Per API call |
总结
在进入“最终思考”阶段时,首先写下相关契约:所需的输入参数、成功标志以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持一致性。
操作检查表
将“操作检查表”阶段视为可衡量的指标,效果会更好。在扩大范围之前,需记录一份最佳处理案例、一个失败案例以及回滚说明。
要把这个阶段视为输入与经过验证的输出之间的契约。为相关文档命名,明确成功标准,杜绝无声的半完成状态。
应将分块策略与检索策略分开。当质量指标发生变化时,修改其中一项不应强制重新编写另一项。
需分别对单轮回复和多轮对话流程进行评分。若仅汇总聊天评分,就会掩盖工具循环故障问题。
编写简短的操作手册:说明如何轮换密钥、如何清空队列以及如何回滚最近的导入操作。
需同时记录正常流程与故障恢复流程。重试机制、人工审核环节以及死信处理都是产品不可或缺的部分,而非后续才需要补充的功能。
在推广该技术栈之前,应冻结版本,为关键流程保存标准记录,并确认好回滚步骤。共享环境需要设置速率限制、进行租户检查,同时明确密钥轮换的负责人。与其展示花哨的一次性演示,不如注重扎实的可靠性。
关于16f80a919c4e的批处理说明:不要将提供者密钥放入代码仓库,为每个会话设置令牌上限,并将转录内容存储在评估测试用例的旁边,以便后续更换模型时仍能保持可比性。