首页 / 文章 / 实用说明:生产环境RAG的数据导入——构建可靠系统

实用说明:生产环境RAG的数据导入——构建可靠系统

《实用笔记》操作指南:面向生产环境RAG的数据导入——为采用该模式的团队构建可靠的解决方案,包括契约、校验机制以及可直接使用的代码模块。

4000 词

本指南将逐步构建从原材料到可运行系统的完整流程,适用于“面向生产环境的RAG数据摄取:打造可靠的企业级数据管道”这一主题。重点在于可操作的步骤、明确的检查点,以及可直接放入代码仓库的代码,无需猜测其用途。 在概览阶段,应在修改代码之前明确输入内容、各步骤的负责人以及完成标准。操作人员应能够从已知的检查点重新运行相应步骤,而无需推测隐藏的状态。 可将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,定义成功标准,杜绝默许的半完成状态。

数据摄取究竟是什么

在处理“实际数据摄取”阶段时,首先需明确相关约定:所需的输入参数、成功标志,以及部分失败时的处理方式。这样的清单能确保后续的代码修改始终符合预期。 在功能结果旁记录执行时间以及令牌或查询成本。提前了解成本情况,可避免从演示环境过渡到共享环境时出现意外费用。 在调整提示词之前,先使用固定的问题集测试召回率。仅仅更换提示词往往无法改善较差的检索效果。

格式问题

在处理“格式问题”阶段时,首先写下相关契约:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 将配置信息与应用程序代码分开存放。环境文件、密钥存储以及功能开关应集中管理,这样操作人员无需查看整个系统结构即可进行审计。 在调整提示词之前,先使用固定的问题集来测试召回率。仅仅更换提示词往往无法解决检索效果不佳的问题。

文档加载器:按格式分类

在分阶段实现文档加载器格式时,首先需明确相关约定:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 同时记录正常流程与异常恢复流程。重试机制、人工审核环节以及错误消息处理都是产品功能的一部分,而非后续需要补充的内容。 在调整提示词之前,先使用固定的问题集来评估检索效果。仅仅更换提示词很难解决检索能力不足的问题。

PDF文档

在处理PDF文档阶段时,首先写下合同的相关内容:所需输入、成功标志以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 优先选择小型、可测试的单元,而非庞大的脚本。当某个步骤失败时,故障应指向单一的责任模块,而非复杂的流程链。 在调整提示词之前,先使用固定的问题集来衡量召回率。仅仅更换提示词很难解决检索效果不佳的问题。

import fitz  # PyMuPDF
from pathlib import Path

def classify_pdf(pdf_path: str) -> str:
    """
    Classify a PDF as text-native, scanned, or hybrid.
    Returns: 'text', 'scanned', or 'hybrid'
    """
    doc = fitz.open(pdf_path)
    total_pages = len(doc)
    pages_with_text = 0
    total_text_length = 0
    for page in doc:
        text = page.get_text("text").strip()
        if text:
            pages_with_text += 1
            total_text_length += len(text)
    doc.close()
    # No text layer at all: scanned PDF
    if pages_with_text == 0:
        return "scanned"
    # Text exists but average is suspiciously short: likely hybrid with poor OCR
    avg_text_per_page = total_text_length / total_pages
    if avg_text_per_page < 100 and pages_with_text < total_pages * 0.5:
        return "hybrid"
    return "text"

def load_text_pdf(pdf_path: str) -> list[dict]:
    """
    Extract text from a text-native PDF using PyMuPDF.
    Returns a list of page dicts with text and metadata.
    """
    doc = fitz.open(pdf_path)
    pages = []
    for page_num, page in enumerate(doc, start=1):
        text = page.get_text("text")
        pages.append({
            "page_number": page_num,
            "text": text.strip(),
            "source": pdf_path,
            "total_pages": len(doc),
        })
    doc.close()
    return pages
from azure.ai.formrecognizer import DocumentAnalysisClient
from azure.core.credentials import AzureKeyCredential

def load_scanned_pdf_azure(
    pdf_path: str,
    endpoint: str,
    api_key: str,
) -> list[dict]:
    """
    Extract text from a scanned PDF using Azure Document Intelligence.
    Returns pages with text, confidence scores, and table data.
    """
    client = DocumentAnalysisClient(
        endpoint=endpoint,
        credential=AzureKeyCredential(api_key),
    )
    with open(pdf_path, "rb") as f:
        poller = client.begin_analyze_document("prebuilt-read", f)
    result = poller.result()
    pages = []
    for page in result.pages:
        page_text = ""
        low_confidence_words = []
        for line in page.lines:
            page_text += line.content + "\n"
        # Flag words with confidence below threshold
        for word in page.words:
            if word.confidence < 0.85:
                low_confidence_words.append({
                    "word": word.content,
                    "confidence": word.confidence,
                })
        pages.append({
            "page_number": page.page_number,
            "text": page_text.strip(),
            "low_confidence_words": low_confidence_words,
            "needs_review": len(low_confidence_words) > 5,
            "source": pdf_path,
        })
    return pages

DOCX:内部政策文件

在处理DOCX内部政策文档阶段时,首先需写下相关契约:所需的输入参数、成功标志以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改始终符合要求。 将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,明确成功标准,杜绝无声的半完成状态。 在调整提示词之前,先使用固定的问题集来测试召回率。仅仅更换提示词往往无法改善较差的检索效果。

from docx import Document as DocxDocument
from docx.oxml.ns import qn

def load_docx(docx_path: str) -> list[dict]:
    """
    Load a DOCX file preserving heading structure and tables.
    Returns a list of content blocks with type and text.
    """
    doc = DocxDocument(docx_path)
    blocks = []
    current_section = "Introduction"
    for element in doc.element.body:
        # Paragraphs
        if element.tag == qn("w:p"):
            para = element
            style = para.style.name if hasattr(para, "style") else ""
            text = para.text.strip()
            if not text:
                continue
            if style.startswith("Heading"):
                level = int(style.replace("Heading ", "")) if style != "Heading" else 1
                current_section = text
                blocks.append({
                    "type": "heading",
                    "level": level,
                    "text": text,
                    "section": current_section,
                    "source": docx_path,
                })
            else:
                blocks.append({
                    "type": "paragraph",
                    "text": text,
                    "section": current_section,
                    "source": docx_path,
                })
        # Tables
        elif element.tag == qn("w:tbl"):
            table_data = []
            for row in element.findall(f".//{qn('w:tr')}"):
                row_data = []
                for cell in row.findall(f".//{qn('w:tc')}"):
                    cell_text = " ".join(
                        p.text for p in cell.findall(f".//{qn('w:t')}")
                        if p.text
                    )
                    row_data.append(cell_text.strip())
                table_data.append(row_data)
            blocks.append({
                "type": "table",
                "data": table_data,
                "section": current_section,
                "source": docx_path,
            })
    return blocks

HTML:合规门户与内部维基

在处理 HTML 合规性门户及相关阶段时,首先需记录下合同要求:所需输入、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改始终符合要求。 在功能结果旁记录处理时间以及令牌或查询成本。提前了解成本情况,可避免在系统从演示环境过渡到共享环境时出现意外费用。 在调整提示词之前,先使用固定的问题集测试召回率。仅仅更换提示词往往无法改善较差的检索效果。

import httpx
from bs4 import BeautifulSoup

def load_html_page(url: str, headers: dict | None = None) -> dict:
    """
    Load an HTML page and extract main content, stripping boilerplate.
    """
    response = httpx.get(url, headers=headers or {}, follow_redirects=True)
    response.raise_for_status()
    soup = BeautifulSoup(response.text, "html.parser")
    # Remove boilerplate elements common in compliance portals
    for tag in soup(["nav", "header", "footer", "script", "style",
                     "aside", "advertisement", ".cookie-banner",
                     ".sidebar", ".breadcrumb"]):
        tag.decompose()
    # Find main content area (adjust selectors for your portal)
    main_content = (
        soup.find("main")
        or soup.find("article")
        or soup.find(id="content")
        or soup.find(class_="content")
        or soup.find("body")
    )
    text = main_content.get_text(separator="\n", strip=True) if main_content else ""
    return {
        "url": url,
        "title": soup.title.string if soup.title else "",
        "text": text,
        "source": url,
    }

JSON与关系型数据库

在处理 JSON 和关系型数据库阶段时,首先需明确相关规范:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改始终符合要求。 应将配置信息与应用程序代码分开。环境文件、密钥存储以及功能开关应集中存放于一处,这样操作人员无需查看整个系统结构即可进行审计。 在调整提示词之前,需先使用固定的问题集来测试召回率。仅仅更换提示词往往无法解决检索效果不佳的问题。

import json
import psycopg2
from typing import Any

def load_json_policies(json_path: str, text_fields: list[str]) -> list[dict]:
    """
    Load a JSON array of policy documents, extracting specified text fields.
    Banking example: a JSON export from a policy management system.
    """
    with open(json_path, "r", encoding="utf-8") as f:
        policies = json.load(f)
    documents = []
    for policy in policies:
        # Combine relevant text fields into a single document
        text_parts = []
        for field in text_fields:
            if field in policy and policy[field]:
                text_parts.append(f"{field.replace('_', ' ').title()}: {policy[field]}")
        documents.append({
            "text": "\n\n".join(text_parts),
            "policy_id": policy.get("id"),
            "policy_type": policy.get("type"),
            "effective_date": policy.get("effective_date"),
            "jurisdiction": policy.get("jurisdiction"),
            "source": json_path,
        })
    return documents

def load_from_postgres(
    connection_string: str,
    query: str,
    text_column: str,
    metadata_columns: list[str],
) -> list[dict]:
    """
    Load documents from a PostgreSQL database.
    Banking example: regulatory requirement records from a compliance database.
    """
    conn = psycopg2.connect(connection_string)
    cursor = conn.cursor()
    cursor.execute(query)
    columns = [desc[0] for desc in cursor.description]
    rows = cursor.fetchall()
    cursor.close()
    conn.close()
    documents = []
    for row in rows:
        row_dict = dict(zip(columns, row))
        documents.append({
            "text": str(row_dict.get(text_column, "")),
            **{col: row_dict.get(col) for col in metadata_columns},
            "source": "postgresql",
        })
    return documents

SharePoint 与 Google Docs:企业级文档存储解决方案

在处理 SharePoint 和 Google Docs 阶段时,首先需明确合同条款:所需输入、成功标志以及部分失败时的处理方式。这份清单能确保后续的代码修改保持一致性。 同时记录正常流程与恢复流程。重试机制、人工审核环节以及错误处理都属于产品功能的一部分,而非后续的优化工作。 在调整提示词之前,需先用固定的问题集来衡量检索效果。仅仅更换提示词很难解决检索能力薄弱的问题。 在处理 SharePoint 和 Google Docs 阶段时,首先需明确合同条款:所需输入、成功标志以及部分失败时的处理方式。这份清单能确保后续的代码修改保持一致性。 将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,定义成功检测标准,并杜绝无声的半完成状态。

import httpx
from typing import Generator

def load_sharepoint_library(
    tenant_id: str,
    client_id: str,
    client_secret: str,
    site_id: str,
    drive_id: str,
) -> Generator[dict, None, None]:
    """
    Iterate over all DOCX and PDF files in a SharePoint document library.
    Yields file metadata dicts; caller is responsible for downloading and parsing.
    """
    # Get access token
    token_url = f"https://login.microsoftonline.com/{tenant_id}/oauth2/v2.0/token"
    token_response = httpx.post(token_url, data={
        "grant_type": "client_credentials",
        "client_id": client_id,
        "client_secret": client_secret,
        "scope": "https://graph.microsoft.com/.default",
    })
    access_token = token_response.json()["access_token"]
    headers = {"Authorization": f"Bearer {access_token}"}
    base_url = f"https://graph.microsoft.com/v1.0/sites/{site_id}/drives/{drive_id}"
    # List files recursively
    next_link = f"{base_url}/root/children"
    while next_link:
        response = httpx.get(next_link, headers=headers)
        data = response.json()
        for item in data.get("value", []):
            if item.get("file") and item["name"].endswith((".pdf", ".docx")):
                yield {
                    "name": item["name"],
                    "download_url": item.get("@microsoft.graph.downloadUrl"),
                    "created": item.get("createdDateTime"),
                    "modified": item.get("lastModifiedDateTime"),
                    "size": item.get("size"),
                    "web_url": item.get("webUrl"),
                }
        next_link = data.get("@odata.nextLink")

流式处理与批量处理

将流式处理与批量导入阶段视为可度量的对象来管理,效果最佳。在扩大范围之前,先记录一份理想的处理结果、一个失败案例以及回滚说明。在功能结果旁同时记录处理时间以及令牌或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外账单。应将分块策略与检索策略分开,当质量指标发生变化时,修改其中一项无需强制重写另一项。

import time
import httpx
from pathlib import Path

class RateLimitedConnector:
    """
    A base connector with retry logic and rate limiting.
    Use this pattern for any external API that imposes rate limits.
    Banking example: connecting to an RBI document portal or a Bloomberg data feed.
    """
    def __init__(
        self,
        base_url: str,
        api_key: str,
        requests_per_minute: int = 60,
        max_retries: int = 3,
        backoff_factor: float = 2.0,
    ):
        self.base_url = base_url
        self.api_key = api_key
        self.min_interval = 60.0 / requests_per_minute
        self.last_request_time = 0.0
        self.max_retries = max_retries
        self.backoff_factor = backoff_factor
    def _wait_for_rate_limit(self):
        elapsed = time.monotonic() - self.last_request_time
        if elapsed < self.min_interval:
            time.sleep(self.min_interval - elapsed)
        self.last_request_time = time.monotonic()
    def fetch(self, endpoint: str, params: dict | None = None) -> dict:
        url = f"{self.base_url}/{endpoint}"
        headers = {"Authorization": f"Bearer {self.api_key}"}
        for attempt in range(self.max_retries):
            self._wait_for_rate_limit()
            try:
                response = httpx.get(url, headers=headers, params=params or {})
                if response.status_code == 429:
                    # Rate limited: back off and retry
                    wait_time = self.backoff_factor ** attempt
                    time.sleep(wait_time)
                    continue
                response.raise_for_status()
                return response.json()
            except httpx.HTTPError as e:
                if attempt == self.max_retries - 1:
                    raise
                time.sleep(self.backoff_factor ** attempt)
        raise RuntimeError(f"Failed to fetch {url} after {self.max_retries} attempts")

多模态数据导入

将多模态数据摄取阶段视为可度量的界面来处理,效果最佳。在扩大范围之前,先记录一份理想的转录结果、一个失败案例以及回滚说明。 将配置置于应用程序代码之外。环境文件、密钥存储和功能标志应集中存放于一个位置,以便操作人员无需查看整个系统结构即可进行审计。 将分块策略与检索策略分开。当质量指标发生变化时,修改其中一项不应迫使重新编写另一项。

基于表格的解析

将“基于表格的解析”阶段视为可度量的对象时,其效果最佳。在扩大范围之前,需记录一份理想的处理结果、一个失败案例以及回滚说明。 同时记录正常处理流程与故障恢复流程。重试机制、人工审核环节以及错误消息处理都是产品功能的一部分,而非后续需要补充的内容。 应将分块策略与检索策略分开。当质量指标发生变化时,修改其中一项不应强制要求重新编写另一项。 将“基于表格的解析”阶段视为输入与经过验证的输出之间的契约。为相关输出文件命名,明确成功标准,杜绝无声的半完成状态。

import fitz
import pdfplumber

def extract_tables_from_pdf(pdf_path: str) -> list[dict]:
    """
    Extract tables from a PDF using pdfplumber.
    Returns tables as both structured data and serialised text.
    Banking example: extracting capital ratio tables from Basel III compliance reports.
    """
    tables_output = []
    with pdfplumber.open(pdf_path) as pdf:
        for page_num, page in enumerate(pdf.pages, start=1):
            tables = page.extract_tables()
            for table_idx, table in enumerate(tables):
                if not table or len(table) < 2:
                    continue
                # First row as headers
                headers = [str(h).strip() if h else f"col_{i}"
                           for i, h in enumerate(table[0])]
                rows = table[1:]
                # Structured data representation
                structured = []
                for row in rows:
                    row_dict = {}
                    for header, cell in zip(headers, row):
                        row_dict[header] = str(cell).strip() if cell else ""
                    structured.append(row_dict)
                # Text serialisation (markdown table format for LLM consumption)
                header_row = " | ".join(headers)
                separator = " | ".join(["---"] * len(headers))
                data_rows = [
                    " | ".join(str(cell or "").strip() for cell in row)
                    for row in rows
                ]
                serialised = "\n".join([header_row, separator] + data_rows)
                tables_output.append({
                    "page_number": page_num,
                    "table_index": table_idx,
                    "headers": headers,
                    "structured_data": structured,
                    "serialised_text": serialised,
                    "source": pdf_path,
                })
    return tables_output

图片与图表导入

在图像与图表导入阶段,应在修改代码之前明确输入内容、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。除了功能结果外,还需记录处理时间以及令牌或查询成本。提前了解成本情况可以避免在从演示环境切换到共享环境时出现意外账单。必须注明实际作为答案依据的段落;没有引用的话,操作人员就无法区分是幻觉内容还是索引缺失导致的错误。

import fitz
import base64
import httpx
from io import BytesIO

def extract_and_describe_images(
    pdf_path: str,
    openai_api_key: str,
    min_image_size: int = 5000,  # bytes, to skip tiny decorative images
) -> list[dict]:
    """
    Extract images from a PDF and generate text descriptions using GPT-4V.
    Banking example: describing NPA trend charts and regulatory workflow diagrams.
    """
    doc = fitz.open(pdf_path)
    image_docs = []
    for page_num in range(len(doc)):
        page = doc[page_num]
        image_list = page.get_images(full=True)
        for img_idx, img_info in enumerate(image_list):
            xref = img_info[0]
            base_image = doc.extract_image(xref)
            image_bytes = base_image["image"]
            # Skip tiny decorative images
            if len(image_bytes) < min_image_size:
                continue
            # Encode to base64 for the vision API
            image_b64 = base64.b64encode(image_bytes).decode("utf-8")
            image_ext = base_image["ext"]
            # Generate description with GPT-4V
            response = httpx.post(
                "https://api.openai.com/v1/chat/completions",
                headers={"Authorization": f"Bearer {openai_api_key}"},
                json={
                    "model": "gpt-4o",
                    "messages": [
                        {
                            "role": "user",
                            "content": [
                                {
                                    "type": "text",
                                    "text": (
                                        "This image is from a banking regulatory document. "
                                        "Describe its content precisely, including any numerical "
                                        "values, trends, labels, axes, or process steps shown. "
                                        "If it is a chart, describe what metric is shown, the "
                                        "time period, and the key data points. If it is a "
                                        "diagram, describe the process or relationship shown."
                                    ),
                                },
                                {
                                    "type": "image_url",
                                    "image_url": {
                                        "url": f"data:image/{image_ext};base64,{image_b64}"
                                    },
                                },
                            ],
                        }
                    ],
                    "max_tokens": 400,
                },
            )
            description = response.json()["choices"][0]["message"]["content"]
            image_docs.append({
                "type": "image_description",
                "page_number": page_num + 1,
                "image_index": img_idx,
                "description": description,
                "source": pdf_path,
            })
    doc.close()
    return image_docs

音频与视频导入

在音频和视频摄取阶段,应在修改代码之前明确输入内容、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 配置应置于应用程序代码之外。环境文件、密钥存储以及功能标志应集中存放于一个位置,以便操作人员无需查看整个系统结构即可进行审计。 需引用实际作为答案依据的段落。如果没有引用,操作人员就无法区分是幻觉内容还是索引缺失导致的错误。

import httpx
import time

def transcribe_audio(
    audio_path: str,
    openai_api_key: str,
    language: str = "en",
) -> dict:
    """
    Transcribe an audio file using OpenAI Whisper.
    Banking example: transcribing an RBI monetary policy press conference recording.
    """
    with open(audio_path, "rb") as audio_file:
        response = httpx.post(
            "https://api.openai.com/v1/audio/transcriptions",
            headers={"Authorization": f"Bearer {openai_api_key}"},
            data={
                "model": "whisper-1",
                "language": language,
                "response_format": "verbose_json",  # includes timestamps per segment
                "timestamp_granularities[]": "segment",
            },
            files={"file": (audio_path, audio_file, "audio/mpeg")},
        )
    result = response.json()
    return {
        "transcript": result.get("text", ""),
        "segments": [
            {
                "start": seg["start"],
                "end": seg["end"],
                "text": seg["text"],
            }
            for seg in result.get("segments", [])
        ],
        "language": result.get("language", language),
        "source": audio_path,
    }

整合应用:生产环境摄取流程

在“整合实施”阶段,修改代码之前需明确输入参数、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 需同时记录正常流程与异常恢复路径。重试机制、人工审核环节以及错误处理都是产品本身的组成部分,而非后续需要补充的内容。 必须引用实际作为答案依据的段落。没有引用的话,操作人员就无法区分是虚假信息还是索引缺失导致的错误。 在“整合实施”阶段,修改代码之前需明确输入参数、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 应将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,明确成功判定标准,绝不允许出现无声无息的半完成状态。

from pathlib import Path
from enum import Enum
import logging

logger = logging.getLogger(__name__)

class DocumentType(Enum):
    PDF_TEXT = "pdf_text"
    PDF_SCANNED = "pdf_scanned"
    DOCX = "docx"
    HTML = "html"
    JSON = "json"
    AUDIO = "audio"
    UNKNOWN = "unknown"

def detect_document_type(file_path: str) -> DocumentType:
    path = Path(file_path)
    suffix = path.suffix.lower()
    if suffix == ".pdf":
        classification = classify_pdf(file_path)
        return DocumentType.PDF_SCANNED if classification == "scanned" else DocumentType.PDF_TEXT
    elif suffix == ".docx":
        return DocumentType.DOCX
    elif suffix in (".html", ".htm"):
        return DocumentType.HTML
    elif suffix == ".json":
        return DocumentType.JSON
    elif suffix in (".mp3", ".mp4", ".wav", ".m4a"):
        return DocumentType.AUDIO
    return DocumentType.UNKNOWN

def ingest_document(
    file_path: str,
    azure_endpoint: str | None = None,
    azure_api_key: str | None = None,
    openai_api_key: str | None = None,
) -> list[dict]:
    """
    Route a document to the correct parser and return a list of content blocks
    with a consistent schema regardless of input format.
    Every output block contains:
    - text: the extracted text content
    - source: origin path or URL
    - doc_type: the detected document type
    - page_number: where applicable
    - needs_review: flag for human review queue
    """
    doc_type = detect_document_type(file_path)
    results = []
    try:
        if doc_type == DocumentType.PDF_TEXT:
            pages = load_text_pdf(file_path)
            for page in pages:
                page["doc_type"] = "pdf_text"
                page["needs_review"] = False
                results.append(page)
        elif doc_type == DocumentType.PDF_SCANNED:
            if azure_endpoint and azure_api_key:
                pages = load_scanned_pdf_azure(file_path, azure_endpoint, azure_api_key)
                for page in pages:
                    page["doc_type"] = "pdf_scanned"
                    results.append(page)
            else:
                logger.warning(
                    f"Scanned PDF detected but no OCR credentials provided: {file_path}"
                )
                results.append({
                    "text": "",
                    "source": file_path,
                    "doc_type": "pdf_scanned",
                    "needs_review": True,
                    "error": "No OCR credentials available",
                })
        elif doc_type == DocumentType.DOCX:
            blocks = load_docx(file_path)
            for block in blocks:
                block["doc_type"] = "docx"
                block["needs_review"] = False
                results.append(block)
        elif doc_type == DocumentType.AUDIO:
            if openai_api_key:
                transcript = transcribe_audio(file_path, openai_api_key)
                for segment in transcript["segments"]:
                    results.append({
                        "text": segment["text"],
                        "source": file_path,
                        "doc_type": "audio_transcript",
                        "start_seconds": segment["start"],
                        "end_seconds": segment["end"],
                        "needs_review": False,
                    })
        else:
            logger.warning(f"Unknown document type, skipping: {file_path}")
    except Exception as e:
        logger.error(f"Ingestion failed for {file_path}: {e}")
        results.append({
            "text": "",
            "source": file_path,
            "doc_type": str(doc_type.value),
            "needs_review": True,
            "error": str(e),
        })
    return results

生产环境中会出现什么问题

在分析“可能出现什么问题”这一阶段时,首先需明确相关约定:所需的输入参数、成功信号,以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改始终符合预期。 在功能结果旁记录处理时间以及令牌或查询成本。提前了解成本情况,可避免从演示环境过渡到共享环境时出现意外的费用支出。 在调整提示词之前,先使用固定的问题集测试系统的召回率。仅仅更换提示词往往无法改善较差的检索效果。

生产环境中的权衡

在处理“生产环境权衡”阶段时,首先写下合同条款:所需的输入、成功标志以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持透明。 将配置信息与应用程序代码分开存放。环境文件、密钥存储以及功能开关应集中于一个位置,这样操作人员无需查看整个系统结构即可进行审计。 在调整提示词之前,先使用固定的问题集来衡量召回率。仅仅更换提示词往往无法解决检索效果不佳的问题。

下一步是什么

在处理“下一步该做什么”这一阶段时,首先需写下相关契约:所需的输入参数、成功标志以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持一致性。 同时记录正常流程与异常恢复路径。重试机制、人工审核环节以及错误消息处理都是产品功能的一部分,而非后续的优化工作。 在调整提示词之前,先使用固定的问题集来衡量检索效果。仅仅更换提示词很难解决检索能力薄弱的问题。 在处理“下一步该做什么”这一阶段时,首先需写下相关契约:所需的输入参数、成功标志以及部分失败时的处理方式。这样的清单能确保后续的代码修改保持一致性。 将这一阶段视为输入与经过验证的输出之间的契约。为相关成果命名,明确成功判定标准,绝不允许出现无声无息的部分完成情况。

操作检查清单

将操作检查清单阶段视为可度量的标准,效果最佳。在扩大范围之前,先记录一份完美的测试用例、一个故障案例以及回滚说明。

优先选择小型且可测试的单元,而非庞大的脚本。当某一步骤出错时,故障应能指向具体的责任方,而非复杂的流程链。

将分块策略与检索策略分开。当质量指标发生变化时,修改其中一项不应迫使重新编写另一项。

在预算允许的情况下,使用测试数据而非真实的付费 API,在持续集成过程中添加用于检测关键路径的冒烟测试。

将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,明确成功标准,拒绝默许部分完成的情况。

将分块策略与检索策略分开。当质量指标发生变化时,修改其中一项不应迫使重新编写另一项。

在推广该技术栈之前,应先冻结版本,为关键流程记录标准输出日志,并明确回滚步骤。共享环境需要设置速率限制、租户验证机制,以及负责密钥轮换的明确责任人。与其展示花哨的一次性演示,不如注重扎实的可靠性。

关于771b701b6010的批处理说明:请将提供商密钥存放在仓库之外,设定单会话令牌上限,并将日志存储在评估用示例文件旁,以便后续更换模型时仍能保持数据可比性。