实用笔记:从零构建RAG系统——实战操作,无需支付API费用
《实用笔记:从零构建RAG系统——实战指南,无需支付API费用》的操作流程详解:为采用该架构的团队提供的合同、检查清单以及可直接使用的代码模板。
以下笔记为“从零构建RAG系统——实战操作,无需支付API费用”提供了一条实用路径。重点在于契约、校验机制以及可直接插入的代码占位符,而非激励性表述。 在完成概览阶段时,首先明确契约内容:所需输入、成功信号以及部分失败时的处理方式。这样的清单能确保后续代码修改的规范性。 同时记录正常流程与异常恢复路径。重试机制、人工审核环节以及死信处理都是产品本身的组成部分,而非后续的优化内容。
RAG究竟是什么(60秒)
将 RAG 的实际运作视为一个可测量的对象最为有效。在扩大范围之前,先记录一份优秀的示例文本、一个失败案例以及回滚说明。 优先选择小型且可测试的单元,而非庞大的脚本。当某个步骤出现故障时,故障应指向单一责任主体,而非复杂的流程链。 将分块策略与检索策略分开。当质量指标发生变化时,修改其中一项不应迫使重新编写另一项。
question ──► [embed] ──► [search your docs] ──► top chunks ──┐
▼
[LLM: "answer using this context"] ──► answer
步骤 0 — 设置
将第0步的设置阶段视为可度量的对象最为有效。在扩大范围之前,先记录一份理想的输出样本、一个失败案例以及回滚说明。 把这一阶段视为输入与经过验证的输出之间的契约。为相关成果命名,明确成功标准,绝不允许出现悄无声息的半完成状态。 将分块策略与检索策略分开。当质量指标发生变化时,修改其中一项不应强制要求重新编写另一项。
pip install sentence-transformers transformers torch numpy
第1步 — 模型从未见过的知识库
将第一步A的知识阶段视为可度量的对象来处理,效果最佳。在扩大范围之前,先记录一份理想的操作日志、一个故障案例以及回滚说明。 在功能结果旁同时记录处理时间以及令牌或查询成本。提前了解成本情况,可避免在系统从演示环境过渡到共享环境时出现意外账单。 为每轮对话和每次会话设定令牌预算。智能工具往往会大量消耗上下文资源,设置上限能防止演示环境变成令人意外的费用来源。 将第一步A的知识阶段视为可度量的对象来处理,效果最佳。在扩大范围之前,先记录一份理想的操作日志、一个故障案例以及回滚说明。 需同时记录正常流程和异常恢复流程。重试机制、人工审核环节以及死信处理都是产品功能的一部分,而非后续需要补充的内容。
# rag.py
DOCUMENTS = [
"""Nimbus is a fictional note-taking app launched in 2023. The free plan,
called Nimbus Lite, allows up to 50 notes and 1 GB of storage. There are no
collaboration features on the free plan.""",
"""Nimbus Pro costs 8 dollars per month billed annually, or 10 dollars billed
monthly. Pro removes the note limit, gives 50 GB of storage, and unlocks
real-time collaboration with up to 5 people per note.""", """Nimbus stores all notes encrypted at rest using AES-256. End-to-end
encryption is only available on the Pro plan and must be enabled manually in
Settings > Security. Once enabled it cannot be turned off for that note.""", """The Nimbus mobile app supports offline editing. Changes made offline are
queued and sync automatically the next time the device is online. If two
devices edit the same note offline, Nimbus keeps both versions and flags a
conflict for the user to resolve.""", """Nimbus offers a 30-day refund policy on all paid plans, no questions asked.
Refunds are processed to the original payment method within 5 business days.
Annual plans cancelled after 30 days are not refundable but stay active until
the end of the billing period.""", """Nimbus support is available via email at help@nimbus.example and live chat.
Live chat is only staffed for Pro customers, Monday to Friday, 9am to 6pm UTC.
Free-plan users receive email support with a typical 48-hour response time.""",
]
from transformers import pipeline
gen = pipeline("text2text-generation", model="google/flan-t5-base")
print(gen("How much does Nimbus Pro cost?", max_new_tokens=50)[0]["generated_text"])
第二步 — 分块处理
在第二阶段的切分步骤中,应在修改代码之前明确输入内容、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。相比冗长的脚本,更应采用小型且可测试的单元。当某个步骤失败时,故障原因应指向单一责任模块,而非复杂的流程链。必须引用实际作为答案依据的段落;没有引用的话,操作人员就无法区分是幻觉内容还是索引缺失所致。
def chunk_text(text, chunk_size=60, overlap=15):
"""Split text into overlapping chunks of `chunk_size` words."""
words = text.split()
chunks = []
start = 0
while start < len(words):
end = start + chunk_size
chunks.append(" ".join(words[start:end]))
if end >= len(words):
break
start = end - overlap # step back by `overlap` so context isn't cut
return chunks
# Build our chunk list, remembering which doc each chunk came from
chunks = []
for doc_id, doc in enumerate(DOCUMENTS):
for c in chunk_text(doc):
chunks.append({"doc_id": doc_id, "text": c})print(f"{len(DOCUMENTS)} documents -> {len(chunks)} chunks")
for c in chunks[:3]:
print("-", c["text"][:70], "...")
第三步 — 嵌入模型:将文本转换为向量
在第三步的嵌入转换阶段,修改代码之前需明确输入内容、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,定义成功判定标准,并拒绝默许的半完成状态。 需引用实际作为答案依据的段落。没有引用的话,操作人员就无法区分幻觉内容与索引缺失的情况。
from sentence_transformers import SentenceTransformer
embedder = SentenceTransformer("all-MiniLM-L6-v2")# Embed every chunk. normalize_embeddings=True makes the vectors unit-length,
# which lets us measure similarity with a simple dot product later.
chunk_texts = [c["text"] for c in chunks]
chunk_vectors = embedder.encode(chunk_texts, normalize_embeddings=True)print("vector shape:", chunk_vectors.shape) # (num_chunks, 384)
import numpy as np
pairs = embedder.encode(
["the price of the pro plan", "how much does it cost", "the weather in Paris"],
normalize_embeddings=True,
)
print("price vs cost :", round(float(pairs[0] @ pairs[1]), 3)) # should be HIGH
print("price vs weather:", round(float(pairs[0] @ pairs[2]), 3)) # should be LOW
第四步——检索:找到能回答问题的片段
在修改代码之前,针对第4步的检索查找阶段,需明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测其中的隐藏状态。应在功能结果旁记录执行时间以及令牌或查询成本。提前显示成本信息,可避免在流程从演示环境转向共享环境时出现意外费用。需注明实际用于得出答案的相关内容;若没有引用依据,操作人员就无法区分是幻觉结果还是索引缺失导致的错误。在修改代码之前,针对第4步的检索查找阶段,需明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测其中的隐藏状态。需同时记录正常流程和异常恢复流程。重试机制、人工审核环节以及错误处理措施都是产品本身的组成部分,而非后续需要补充的内容。
import numpy as np
def retrieve(question, k=3):
q_vec = embedder.encode([question], normalize_embeddings=True)[0]
scores = chunk_vectors @ q_vec # cosine similarity to every chunk
top_idx = np.argsort(scores)[::-1][:k] # indices of the k highest scores
return [(chunks[i]["text"], float(scores[i])) for i in top_idx]for text, score in retrieve("How much does Nimbus Pro cost?"):
print(f"[{score:.3f}] {text[:80]}...")
第5步——生成:让模型根据上下文给出答案
在处理第5步的生成阶段时,首先需明确相关规范:所需的输入参数、成功标志,以及部分失败时的处理方式。这样的检查清单有助于确保后续代码修改的合理性。 建议采用小型、可测试的单元而非庞大的脚本。当某一步骤失败时,故障应能指向具体的责任模块,而非复杂的流程链。 应对稳定的系统指令和工具结构进行缓存。重复发送相同的开头信息是导致资源浪费的常见原因。
from transformers import pipeline
generator = pipeline("text2text-generation", model="google/flan-t5-base")def rag_answer(question, k=3):
retrieved = retrieve(question, k=k)
context = "\n".join(text for text, _ in retrieved) prompt = f"""Answer the question using only the context below.
If the answer is not in the context, say you don't know.Context:
{context}Question: {question}
Answer:""" out = generator(prompt, max_new_tokens=80)[0]["generated_text"]
return out.strip(), retrievedanswer, sources = rag_answer("How much does Nimbus Pro cost?")
print("ANSWER:", answer)
print("\nBased on:")
for text, score in sources:
print(f" [{score:.3f}] {text[:70]}...")
第6步——整合所有内容
在执行第6步“制定契约”阶段时,首先需写下相关约定:所需的输入参数、成功标志,以及部分失败时的处理方式。这样的清单能确保后续的代码修改始终符合初始约定。 将此阶段视为输入与验证后输出之间的契约。为相关成果命名,明确成功判定标准,杜绝默许部分完成的情况。 在调整提示词之前,先使用固定的问题集来测试召回率。仅仅更换提示词往往无法改善较差的检索效果。
# rag.py — a complete, local, no-API RAG system
import numpy as np
from sentence_transformers import SentenceTransformer
from transformers import pipeline
DOCUMENTS = [
"""Nimbus is a fictional note-taking app launched in 2023. The free plan,
called Nimbus Lite, allows up to 50 notes and 1 GB of storage. There are no
collaboration features on the free plan.""",
"""Nimbus Pro costs 8 dollars per month billed annually, or 10 dollars billed
monthly. Pro removes the note limit, gives 50 GB of storage, and unlocks
real-time collaboration with up to 5 people per note.""",
"""Nimbus stores all notes encrypted at rest using AES-256. End-to-end
encryption is only available on the Pro plan and must be enabled manually in
Settings > Security. Once enabled it cannot be turned off for that note.""",
"""The Nimbus mobile app supports offline editing. Changes made offline are
queued and sync automatically the next time the device is online. If two
devices edit the same note offline, Nimbus keeps both versions and flags a
conflict for the user to resolve.""",
"""Nimbus offers a 30-day refund policy on all paid plans, no questions asked.
Refunds are processed to the original payment method within 5 business days.
Annual plans cancelled after 30 days are not refundable but stay active until
the end of the billing period.""",
"""Nimbus support is available via email at help@nimbus.example and live chat.
Live chat is only staffed for Pro customers, Monday to Friday, 9am to 6pm UTC.
Free-plan users receive email support with a typical 48-hour response time.""",
]def chunk_text(text, chunk_size=60, overlap=15):
words = text.split()
chunks, start = [], 0
while start < len(words):
end = start + chunk_size
chunks.append(" ".join(words[start:end]))
if end >= len(words):
break
start = end - overlap
return chunksprint("Loading models (first run downloads them)...")
embedder = SentenceTransformer("all-MiniLM-L6-v2")
generator = pipeline("text2text-generation", model="google/flan-t5-base")# Index the documents once at startup
chunks = []
for doc_id, doc in enumerate(DOCUMENTS):
for c in chunk_text(doc):
chunks.append({"doc_id": doc_id, "text": c})
chunk_vectors = embedder.encode(
[c["text"] for c in chunks], normalize_embeddings=True
)def retrieve(question, k=3):
q_vec = embedder.encode([question], normalize_embeddings=True)[0]
scores = chunk_vectors @ q_vec
top_idx = np.argsort(scores)[::-1][:k]
return [(chunks[i]["text"], float(scores[i])) for i in top_idx]def rag_answer(question, k=3):
retrieved = retrieve(question, k=k)
context = "\n".join(text for text, _ in retrieved)
prompt = (
"Answer the question using only the context below. "
"If the answer is not in the context, say you don't know.\n\n"
f"Context:\n{context}\n\nQuestion: {question}\nAnswer:"
)
out = generator(prompt, max_new_tokens=80)[0]["generated_text"]
return out.strip()if __name__ == "__main__":
print("RAG ready. Ask about Nimbus (or type 'quit').\n")
while True:
q = input("You: ").strip()
if q.lower() in {"quit", "exit", ""}:
break
print("Nimbus bot:", rag_answer(q), "\n")
python rag.py
第7步 —— 证明RAG确实发挥了作用(A/B测试)
在完成第7步“验证RAG”阶段时,首先需明确相关规范:所需输入、成功标志以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持一致性。 在功能结果旁记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在系统从演示环境转向共享环境时出现意外费用。 在调整提示词之前,先使用固定的问题集测试召回率。仅仅更换提示词往往无法改善较差的检索效果。 在完成第7步“验证RAG”阶段时,首先需明确相关规范:所需输入、成功标志以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持一致性。 需同时记录正常流程和异常恢复流程。重试机制、人工审核环节以及错误处理措施都是产品本身的一部分,而非后续需要补充的内容。
def no_rag(question):
out = generator(f"Question: {question}\nAnswer:", max_new_tokens=80)
return out[0]["generated_text"].strip()
q = "Can free-plan Nimbus users use live chat support?"
print("WITHOUT context:", no_rag(q))
print("WITH context :", rag_answer(q))
第8步——进一步优化(选择你感兴趣的部分)
第8步的优化工作若能作为可衡量的指标来执行效果最佳。在扩大范围之前,先记录一份优秀的处理结果、一个失败案例以及回滚说明。 建议采用小型、可测试的单元,而非庞大的脚本。当某一步骤出现故障时,故障应指向单一的责任模块,而非复杂的流程链。 将分块策略与检索策略分开。当质量指标发生变化时,修改其中一项不应迫使重新编写另一项。
需牢记的心智模型
将阶段性的思维模型视为可测量的对象来处理,效果最佳。在扩大范围之前,先记录一份优秀的案例、一个失败案例以及回滚说明。 把这一阶段视为输入与已验证输出之间的契约。为相关成果命名,明确成功标准,绝不允许默默地只完成部分工作。 为每轮及每次会话设定token预算。智能工具往往会过度扩展上下文;设置上限可避免演示变成意外的费用账单。
故障排除
将故障排查阶段视为可度量的工作面,效果最佳。在扩大范围之前,需记录一份理想的处理流程、一个故障案例以及回滚说明。 在功能结果旁同时记录处理时间以及令牌或查询成本。提前了解成本情况,可避免在系统从演示环境过渡到共享环境时出现意外账单。 应将分块策略与检索策略分开。当质量指标发生变化时,调整其中一个不应迫使重新编写另一个。 将故障排查阶段视为可度量的工作面,效果最佳。在扩大范围之前,需记录一份理想的处理流程、一个故障案例以及回滚说明。 需同时记录正常处理路径和恢复路径。重试机制、人工审核环节以及死信处理都是产品本身的组成部分,而非后续需要补充的功能。
操作检查清单
在操作检查清单阶段,应在修改代码之前明确输入参数、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新执行该步骤,而无需猜测隐藏状态。
将配置信息与应用程序代码分开。环境文件、密钥存储以及功能标志应集中存放于一个位置,以便操作人员无需查看整个系统结构即可进行审计。
需注明实际作为答案依据的段落。如果没有引用,操作人员就无法区分是虚假信息还是索引缺失导致的错误。
编写简短的操作手册:包括如何轮换密钥、如何清空队列以及如何回滚上一次的数据导入操作。
同时记录正常流程和故障恢复流程。重试机制、人工审核环节以及死信处理都是产品功能的一部分,而非后续需要补充的内容。
请引用实际作为答案依据的段落。没有引用的话,操作人员就无法区分是幻觉还是索引缺失导致的错误。
在推广该技术栈之前,先冻结版本,为关键流程保存完整的记录,并确认回滚步骤。共享环境需要设置速率限制、进行租户检查,同时明确密钥轮换的负责人。与其展示花哨的一次性演示,不如追求扎实的可靠性。
关于5223acdafa84的批注:不要将提供商密钥放入代码仓库,为每个会话设置令牌使用上限,并将记录与评估用的固定文件放在一起,以便后续更换模型时仍能保持对比性。