Галоўная / Артыкулы / Практычныя прытамулкі: Абсалютный тэхнічны паўодл для інструментаў RAG з адкрытым кодам

Практычныя прытамулкі: Абсалютный тэхнічны паўодл для інструментаў RAG з адкрытым кодам

Практычныя прыказкі: Абсалютный тэхнічны паўнаклад для інструментаў RAG з адкрытым кодам: контракты, перакананні та слоты для коду, якія можна прыўязаць для команд, якіе выкарыстоўваюць гэты патэрн.

3416 слоў

У гэтым карыце параграфе перакладзена схема працы, яка практычна для выкорыстоўвання: «Абсалютны тэхнічны параднік з інструментамі RAG на адкрытым кодзе: Docling, LlamaIndex, LangChain, Haystack і RAGAS». Увага сфокусавана на практычных кроках, чысткіх пераканальніцах і кодзе, які можна проста дадаць у репазітарый, не прабуючы здогадвацца пра мету.

Стварэнне систем генеравання з падтрымкам разысквання дакументаў, гатовых да выкорыстоўвання, у 2026 годзе

Для стадіі «Адаптаваная генерэцыя з выкарыстоўваннем дадзеных парадку будавання» неабходна перад змянай коду адзначыць вхідныя даны, адпаведальнага за шаг і крэтыяяры завершэння. Аперацыёныя працавнікі должны магчымае запускаць шаг з вядомай точкі контролю, не прабуючы спадарожваць схованы стан. Запісваце час выконання і кост токенаў або запытаў разам з функцыйнальнымі рэзультатамі. Відразлівая видавальнасць костаў з самага пачатку запобегае неспакойным рахункам, калі працэс пераходзіць з дэмавай версіі ў спяльныя среды. Указваце тыя часткі тексту, якія фактычна ляглі в основу адпаведнай адказы. Без цых цітатаў аперацыёныя працавнікі не зможуць адразліваць галюцинацыі ад прасоў у індэксаванні.

Прыямыя прычыны

У стадії прабэрасупактаў неабходна з’явіць вхідныя даны, адпаведальнага за крок і крэтыры завершэння пры перадзмене коду. Аперацыйныя працавнікі павінны магчымае перазапускаць крок з вядомай точкі контролю, не падозрываючы прыхованы стан. Конфігурацыю трэба залічыць параду ад коду прыкладнення. Файлы сераўіса, хранільнікі секрэтных дадзеных і флагі функцыйяў павінны знаходзіцца ў аднам месцы, якое працавнікі можуць пераглядаць, не чытаючы весь структураны код. Паказваць трэба тыя часткі тексту, якія фактычна лежаць у падставе адпаведнай адказы. Без ціх цитатаў працавнікі не зможуць адразніць галюцинацыю ад працягу індэксавання.

1. Docling: Аднойчына для інтэлігентной обробкі дакументаў

Для стадіі 1 Docling Intelligent Document неабяжна ўзначыць вхідныя даны, адпаведальнага за крок і крэтыры завершэння пры змяне коду. Аперацыйныя працавнікі должны магчыма ўвайсці крок з вядомага пункта контролю, не спрабоўваючы здагадвацца пра схованы стан. Неабяжна задокументаваць як шлях успеху, так і шлях вярнення. Перапрыбуткі, людзкі контроль і обработка некоректных паведамленняў є часткай продукту, а не чымсь, што дадаецца пазней. Паказваць часткі тексту, якія фактычна лежаць у падставе адпаведнай адказы. Без ціх цитатаў аперацыйныя працавнікі не зможуць адразніць галюцинацыю ад прасоў у індэксаванні.

Інсталяцыя і базовая наработка

У стадії інсталяцыі та базовага викорыстання неабяжна практычна апранаванне: перад змінайом код неабяжна вызначыць параметры вводу, адпаведнага власніка крока та критэрыя завершэння. Аператары должны магчымаць перзапуск крока з вядомай точкі контролю, не падозрываючы прыхованы стан. Лепш выбіраць маленькія, тэставаныя елементы замест большых скрыптов. Калі крок не выконваецца, прычына неудачы должна вказываць на адзін конкрэтны элемент, а не на заплутаны ланцюг задач. Неабяжна цітаваць тые часткі тексту, якія фактычна служылі падставай для адказу. Без цітатаў аператары не зможаць разлічыць галюцинацію ад працягу праз прычыны, зв’язаныя з індэксаванням.

pip install docling
from docling.document_converter import DocumentConverter
# Convert a PDF document to structured Markdown
converter = DocumentConverter()
result = converter.convert("sample_document.pdf")
structured_markdown = result.document.export_to_markdown()
print(structured_markdown)

Развітая обробка дакументаў з выкарыстоўванням OCR

Для прынцыповай пераробкі дакументаў з выкарыстоўваннем стадзій неабходна пазначыць вхідныя даны, адпаведальнага за кожную стадзію і крэтыярыя завершэння пры змены коду. Аперацыйныя працавнікі должны магчымае перзапускаць стадзію з вядомага пункта контролю, не прымуджаючыся вычуваць схованы стан. Спрыяйце цій стадзіі як даговору межа вхіднымі данымі і перакананымі выходнымі рэзультатамі. Дайце назвы артыфактам, пазначыце крэтыярыі успеху і не прымайце часткова завершаныя рэзультаты без паведамлення. Указуйце тыя часткі тексту, якія фактычна лежалі в основе адпаведнай адказы. Без цых цітатаў аперацыйныя працавнікі не зможуць адразніць галюцинацыю ад прасоў у індэксаванні.

from docling.document_converter import DocumentConverter
from docling.datamodel.pipeline_options import PipelineOptions
# Configure pipeline with OCR for scanned documents
pipeline_options = PipelineOptions(
    do_ocr=True,  # Enable OCR for scanned documents
    ocr_engine="tesseract",  # Use Tesseract OCR engine
    ocr_language="eng"  # English language
)
converter = DocumentConverter(pipeline_options=pipeline_options)
result = converter.convert("scanned_document.pdf")
# Extract structured data including tables and images
document = result.document
tables = document.tables
images = document.images
print(f"Found {len(tables)} tables and {len(images)} images")

Інтэграцыя з іншымі фреймворкамі

У стадії інтеграцыі з іншымі фрэймворкамі неабяжна практычна адзначыць вхідныя даны, адпаведальнага за крок і критэрыі завершэння пры перадзеяванні коду. Аператары должны магчымаць перзапуск кроку з вядомай точкі контролю, не падозрываючы прыхованы стан. Запісваць час выконання і вартасць токенаў або запытак праза функцыйнальных рэзультаатаў. Відразлівасць вартасцей з самага пачатку запобегае неспакоўным рахункам, калі процес пераходзіць з дэмовай среды ў спакульную. Адзначаць тыя часткі тексту, якія фактычна ляглі в основу адпаведнай адказы. Без цых цітатаў аператары не можуць разлічыць галюцинацыю ад прасоўкі ў індэксаванні.

# Example: Integrating Docling with LangChain
from langchain_community.document_loaders import TextLoader
from docling.document_converter import DocumentConverter
def docling_to_langchain_docs(file_path):
    """Convert document using Docling and return as LangChain documents"""
    converter = DocumentConverter()
    result = converter.convert(file_path)
    content = result.document.export_to_markdown()
    # Create LangChain document
    loader = TextLoader(file_path)
    docs = loader.load()
    docs[0].page_content = content
    docs[0].metadata["source"] = file_path
    return docs
# Usage
langchain_docs = docling_to_langchain_docs("technical_manual.pdf")

2. LlamaIndex: Архітектура з адзыскам даных у працоўны режым

Для стадіі архітектуры LlamaIndex Retrieval-First 2 неабяжна прадзефінавацыя вхідных дадзеных, адпраўніка крока і крэтарыяў завершэння працы перад зменым коду. Аперацыйныя працавнікі павінны магчымае перзапускаць крок з вядомай точкі контролю, не падозрываючы прыхованы стан. Конфігурацыю трэба захаваць праз аддзел ад коду прыемліка. Файлы сераўіснага сераўісу, хранілішчы секрэтных дадзеных і флагі функцияў павінны знаходзіцца ў адном месцы, якое працавнікі можуць пераглядаць, не чытаючы весь граф. Неабяжна цітацыя тых частак, якія фактычна сталі падставай для адпаведзення. Без цітацый працавнікі не можуць адразніць галюцинацыю ад прасоўкі ў індэксаванні.

Базовая реалізацыя RAG

У стадії базовай рэалізацыі RAG неабходна прадзеяванне вхідных дадзеных, абяцкага, які керуе цым крокам, і крэтарыяў для завершэння працы перад зменым коду. Аперацыйныя працавнікі должны магчыма было перзапускаць крок з вядомай точкі контролю, не падозрываючы прыхованы стан системы. Неабходна аддзеяваць як шлях успеху, так і шлях вярнення да нормальнага стану. Перапрыбуткі, людзкія пераказы і обробка некоректных паведамленняў є часткай продукту, а не элементамі, якія дадзецца дапрацаваць пазней. Паказваць трэба тыя часткі тексту, якія фактычна сталі падставай для адпаведнага адказу. Без такіх цытатаў аперацыйныя працавнікі не зможу разлічыць галюцинацію ад прасоў у індэксаванні.

from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
from llama_index.llms.openai import OpenAI
# Load documents
documents = SimpleDirectoryReader("./data").load_data()
# Create index
index = VectorStoreIndex.from_documents(documents)
# Create query engine
query_engine = index.as_query_engine(
    similarity_top_k=3,
    response_mode="compact"
)
# Query the system
response = query_engine.query("What are the main features of the product?")
print(response)

Хіярархічна разбівка на часткі з аўтаматычным з’еднанням

Для стадіі іерархічнага разлікавання з аўтаматычным з’еднаннем неабходна пазначыць вхідныя даны, адпаведальнага за этап і крэтыры завершэння пры перадзеяванні коду. Аператары должны магчымаць перзапуск этапа з вядомай точкі контролю, не прабуючы спадарацца пра схованы стан. Лепш выбіраць маленькія, тэставаныя елементы замест большых скрыптов. Калі этап не выйшаў, прычына неудачы должна вказываць на адну конкрэтную адпаведальнасць, а не на заплутаны процес. Прытамульвайце тыя часткі тексту, якія фактычна лежаць у падставе адпаведнай адказы. Без ціх цытатаў аператары не зможуць разлічыць галюцинацію ад прасоўкі ў індэксаванні.

from llama_index.core import VectorStoreIndex
from llama_index.core.node_parser import HierarchicalNodeParser, get_leaf_nodes
from llama_index.core.retrievers import AutoMergingRetriever
from llama_index.core.query_engine import RetrieverQueryEngine
# Create hierarchical nodes: 2048 -> 512 -> 128 token chunks
node_parser = HierarchicalNodeParser.from_defaults(
    chunk_sizes=[2048, 512, 128]
)
nodes = node_parser.get_nodes_from_documents(documents)
leaf_nodes = get_leaf_nodes(nodes)
# Build index on leaf nodes only
index = VectorStoreIndex(leaf_nodes)
index.storage_context.docstore.add_documents(nodes)
# Auto-merging retriever replaces small chunks with parent context when relevant
retriever = AutoMergingRetriever(
    index.as_retriever(similarity_top_k=3),
    index.storage_context,
    merge_batch_size=5,
)
query_engine = RetrieverQueryEngine(retriever)
response = query_engine.query("Explain the technical specifications in detail")
print(response)

Адзінакавая ацэнка з LlamaIndex

Для стадіі адаптаванага ацэнкавання з LlamaIndex неабяжна ўзначыць вхідныя даны, адпаведальнага за крок і критэрыя завершэння прычыні змены коду. Аператары должны магчымаць перзапуск кроку з вядомай точкі контролю без неабяжнага вычыслення скрытых станоў. Што да гэтай стадіі, яе трэба спрыяць як кантракт межа вхіднымі данымі і падтвердзенымі выходнымі рэзультатамі. Назваць всі элементы, узначыць критэрыя успеху і не падтрымляць тых варыянтаў, калі робота завершаецца часткова без адпаведных падтверджэнняў. Неабяжна цітацыя тых частак тексту, якія фактычна ляглі в основу адпаведнай адказы. Без цітацый аператары не можуць разлічыць галюцинацію ад працягу процэсу індексавання.

from llama_index.core.evaluation import FaithfulnessEvaluator, RelevancyEvaluator
from llama_index.core import Settings
# Initialize evaluators
faithfulness_evaluator = FaithfulnessEvaluator(llm=Settings.llm)
relevancy_evaluator = RelevancyEvaluator(llm=Settings.llm)
# Evaluate response
eval_result = faithfulness_evaluator.evaluate_response(
    query="What are the system requirements?",
    response=response,
    contexts=[node.text for node in response.source_nodes]
)
print(f"Faithfulness score: {eval_result.score}")
print(f"Relevancy score: {relevancy_evaluator.evaluate_response(query='What are the system requirements?', response=response, contexts=[node.text for node in response.source_nodes]).score}")

3. LangChain: Фрэймворк, адзірваны на рабочых практыках

Для стадіі 3 фрэймворка LangChain, адзірванага на рабочых практыках, неабяжна ўзначыць вхідныя даны, адпаведальнага за крок і критэрыя завершэння прычыні змены коду. Аператары должны магчымаць перзапуск кроку з вядомай точкі контролю без неабяжнага вычыслення скрытых станоў.

Базовы прайсепт RAG

from langchain_community.document_loaders import PyPDFLoader
from langchain_text_splitters import RecursiveCharacterTextSplitter
from langchain_community.vectorstores import Chroma
from langchain_openai import OpenAIEmbeddings, ChatOpenAI
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.runnables import RunnablePassthrough
from langchain_core.output_parsers import StrOutputParser
# Load and split documents
loader = PyPDFLoader("sample.pdf")
docs = loader.load()
text_splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=200)
splits = text_splitter.split_documents(docs)
# Create vector store
vectorstore = Chroma.from_documents(documents=splits, embedding=OpenAIEmbeddings())
# Create retriever
retriever = vectorstore.as_retriever()
# Create prompt template
template = """Answer the question based only on the following context:
{context}
Question: {question}
"""
prompt = ChatPromptTemplate.from_template(template)
# Create chain
llm = ChatOpenAI(model_name="gpt-4o", temperature=0)
rag_chain = (
    {"context": retriever, "question": RunnablePassthrough()}
    | prompt
    | llm
    | StrOutputParser()
)
# Execute
response = rag_chain.invoke("What are the key benefits mentioned in the document?")
print(response)

Багатыякшы процес выкарыстоўвання з памятчукам

from langchain_core.messages import HumanMessage, AIMessage
from langchain_core.chat_history import BaseChatMessageHistory
from langchain_core.runnables.history import RunnableWithMessageHistory
from langchain_community.chat_message_histories import ChatMessageHistory
# Create chat history
class InMemoryHistory(BaseChatMessageHistory):
    def __init__(self):
        self.messages = []
    def add_user_message(self, message: str):
        self.messages.append(HumanMessage(content=message))
    def add_ai_message(self, message: str):
        self.messages.append(AIMessage(content=message))
    def clear(self):
        self.messages = []
# Store chat history
chat_histories = {}
def get_chat_history(session_id: str) -> BaseChatMessageHistory:
    if session_id not in chat_histories:
        chat_histories[session_id] = InMemoryHistory()
    return chat_histories[session_id]
# Create chain with memory
chain_with_memory = RunnableWithMessageHistory(
    rag_chain,
    get_chat_history,
    input_messages_key="question",
    history_messages_key="chat_history",
)
# Execute with memory
session_id = "user_123"
response = chain_with_memory.invoke(
    "What are the key benefits mentioned in the document?",
    config={"configurable": {"session_id": session_id}}
)
print(response)
# Follow-up question
follow_up_response = chain_with_memory.invoke(
    "Can you elaborate on the second benefit?",
    config={"configurable": {"session_id": session_id}}
)
print(follow_up_response)

4. Haystack: Пошук высокага стандарта

Базовая наладка каналу обработкі

from haystack import Pipeline
from haystack.components.embedders import SentenceTransformersTextEmbedder
from haystack.components.retrievers import InMemoryBM25Retriever, InMemoryEmbeddingRetriever
from haystack.components.joiners import JoinDocuments
from haystack.components.generators import OpenAIGenerator
from haystack.components.preprocessors import DocumentCleaner, DocumentSplitter
from haystack.document_stores.in_memory import InMemoryDocumentStore
from haystack.utils import Secret
# Initialize components
document_store = InMemoryDocumentStore()
cleaner = DocumentCleaner()
splitter = DocumentSplitter(split_by="word", split_length=1000)
embedder = SentenceTransformersTextEmbedder(model="sentence-transformers/all-MiniLM-L6-v2")
bm25_retriever = InMemoryBM25Retriever(document_store=document_store)
embedding_retriever = InMemoryEmbeddingRetriever(document_store=document_store)
joiner = JoinDocuments(join_mode="concatenate")
generator = OpenAIGenerator(api_key=Secret.from_env_var("OPENAI_API_KEY"), model="gpt-4o")
# Create pipeline
pipeline = Pipeline()
pipeline.add_component("cleaner", cleaner)
pipeline.add_component("splitter", splitter)
pipeline.add_component("embedder", embedder)
pipeline.add_component("bm25_retriever", bm25_retriever)
pipeline.add_component("embedding_retriever", embedding_retriever)
pipeline.add_component("joiner", joiner)
pipeline.add_component("generator", generator)
# Connect components
pipeline.connect("cleaner", "splitter")
pipeline.connect("splitter", "embedder")
pipeline.connect("embedder", "embedding_retriever")
pipeline.connect("bm25_retriever", "joiner")
pipeline.connect("embedding_retriever", "joiner")
pipeline.connect("joiner", "generator")
# Add documents to store
from haystack.dataclasses import Document
documents = [
    Document(content="The system requires 8GB RAM minimum"),
    Document(content="Supports Windows, macOS, and Linux"),
    Document(content="Network bandwidth should be at least 10Mbps")
]
document_store.write_documents(documents)
# Run pipeline
result = pipeline.run({
    "cleaner": {"documents": documents},
    "bm25_retriever": {"query": "system requirements"},
    "embedding_retriever": {"query": "system requirements"}
})
print(result["generator"]["replies"][0])

Адкалікаванне гібрыднага пошуку

from haystack import Pipeline
from haystack.components.embedders import SentenceTransformersTextEmbedder
from haystack.components.retrievers import InMemoryBM25Retriever, InMemoryEmbeddingRetriever
from haystack.components.rankers import TransformersRanker
from haystack.components.joiners import JoinDocuments
from haystack.components.generators import OpenAIGenerator
from haystack.document_stores.in_memory import InMemoryDocumentStore
# Create hybrid search pipeline
pipeline = Pipeline()
# Add components
document_store = InMemoryDocumentStore()
bm25_retriever = InMemoryBM25Retriever(document_store=document_store)
embedding_retriever = InMemoryEmbeddingRetriever(document_store=document_store)
ranker = TransformersRanker(model_name_or_path="BAAI/bge-reranker-base")
joiner = JoinDocuments(join_mode="concatenate")
generator = OpenAIGenerator(model="gpt-4o")
# Add components to pipeline
pipeline.add_component("bm25_retriever", bm25_retriever)
pipeline.add_component("embedding_retriever", embedding_retriever)
pipeline.add_component("ranker", ranker)
pipeline.add_component("joiner", joiner)
pipeline.add_component("generator", generator)
# Connect components
pipeline.connect("bm25_retriever", "ranker.query")
pipeline.connect("embedding_retriever", "ranker.documents")
pipeline.connect("ranker", "joiner")
pipeline.connect("joiner", "generator")
# Run hybrid search
result = pipeline.run({
    "bm25_retriever": {"query": "system requirements"},
    "embedding_retriever": {"query": "system requirements"}
})
print(result["generator"]["replies"][0])

5. RAGAS: Кантэкст для ацэнкі

Базовая наладка ацэнкі

import pandas as pd
from ragas import evaluate
from ragas.metrics import (
    faithfulness,
    answer_relevancy,
    context_precision,
    context_recall,
    context_relevancy,
    answer_similarity
)
from datasets import Dataset
# Create evaluation dataset
data = {
    "question": ["What are the system requirements?", "How does the authentication work?"],
    "answer": ["The system requires 8GB RAM minimum", "Authentication uses OAuth 2.0"],
    "contexts": [
        ["The system requires 8GB RAM minimum", "Supports Windows, macOS, and Linux"],
        ["Authentication uses OAuth 2.0", "Multi-factor authentication is optional"]
    ],
    "ground_truths": [
        ["The system requires 8GB RAM minimum"],
        ["Authentication uses OAuth 2.0 with JWT tokens"]
    ]
}
dataset = Dataset.from_dict(data)
# Evaluate
result = evaluate(
    dataset,
    metrics=[
        faithfulness,
        answer_relevancy,
        context_precision,
        context_recall,
        context_relevancy,
        answer_similarity
    ]
)
print(result)

Адкалікаванне па індывідуальнымя патрэнтамі з метрыкамі без апылу

from ragas import evaluate
from ragas.metrics import (
    faithfulness,
    answer_relevancy,
    context_precision,
    context_recall,
    context_relevancy,
    answer_similarity,
    answer_correctness
)
from datasets import Dataset
# Create dataset without ground truth (reference-free evaluation)
data = {
    "question": ["What are the system requirements?", "How does the authentication work?"],
    "answer": ["The system requires 8GB RAM minimum", "Authentication uses OAuth 2.0"],
    "contexts": [
        ["The system requires 8GB RAM minimum", "Supports Windows, macOS, and Linux"],
        ["Authentication uses OAuth 2.0", "Multi-factor authentication is optional"]
    ]
}
dataset = Dataset.from_dict(data)
# Evaluate without ground truth
result = evaluate(
    dataset,
    metrics=[
        faithfulness,
        answer_relevancy,
        context_precision,
        context_recall,
        context_relevancy,
        answer_similarity
    ]
)
print(result)

Інтэграцыя з існуючымі системамі RAG

from ragas import evaluate
from ragas.metrics import faithfulness, answer_relevancy, context_precision
from datasets import Dataset
import json
def evaluate_rag_system(rag_system, test_questions, expected_answers=None):
    """Evaluate a RAG system against test questions"""
    results = {
        "question": [],
        "answer": [],
        "contexts": [],
        "ground_truths": [] if expected_answers else None
    }
    for i, question in enumerate(test_questions):
        # Get answer from RAG system
        answer = rag_system(question)
        # Get contexts used by RAG system (assuming it returns contexts)
        contexts = rag_system.get_contexts(question) if hasattr(rag_system, 'get_contexts') else []
        results["question"].append(question)
        results["answer"].append(answer)
        results["contexts"].append(contexts)
        if expected_answers:
            results["ground_truths"].append([expected_answers[i]])
    # Create dataset
    dataset = Dataset.from_dict(results)
    # Evaluate
    metrics = [faithfulness, answer_relevancy, context_precision]
    if expected_answers:
        metrics.append(context_recall)
    evaluation_result = evaluate(dataset, metrics=metrics)
    return evaluation_result
# Example usage
# Assuming you have a RAG system object
# evaluation = evaluate_rag_system(your_rag_system, test_questions, expected_answers)
# print(evaluation)

Пораўнавальны аналіз: Спалучэнне інструментаў у архітектурах выкарыстоўвання

Рэкамендаваныя патрэнты архітектуры

# Complete production architecture combining all tools
from docling.document_converter import DocumentConverter
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
from langchain_core.prompts import ChatPromptTemplate
from haystack import Pipeline
from ragas import evaluate
from datasets import Dataset
class ProductionRAGSystem:
    def __init__(self):
        self.docling_converter = DocumentConverter()
        self.llamaindex_index = None
        self.langchain_chain = None
        self.haystack_pipeline = None
        self.evaluation_metrics = []
    def preprocess_documents(self, document_paths):
        """Use Docling for intelligent document processing"""
        processed_docs = []
        for path in document_paths:
            result = self.docling_converter.convert(path)
            markdown_content = result.document.export_to_markdown()
            processed_docs.append({
                "content": markdown_content,
                "metadata": {"source": path}
            })
        return processed_docs
    def build_llamaindex_index(self, documents):
        """Build index using LlamaIndex for efficient retrieval"""
        from llama_index.core import VectorStoreIndex
        from llama_index.core.node_parser import SentenceSplitter
        # Split documents
        splitter = SentenceSplitter(chunk_size=512, chunk_overlap=50)
        nodes = splitter.get_nodes_from_documents(documents)
        # Create index
        self.llamaindex_index = VectorStoreIndex(nodes)
        return self.llamaindex_index
    def create_langchain_chain(self, index):
        """Create LangChain chain for complex workflows"""
        from langchain_core.runnables import RunnablePassthrough
        from langchain_openai import ChatOpenAI
        from langchain_core.output_parsers import StrOutputParser
        # Create retriever
        retriever = index.as_retriever(similarity_top_k=3)
        # Create prompt
        template = """Answer the question based only on the following context:
        {context}
        Question: {question}
        """
        prompt = ChatPromptTemplate.from_template(template)
        # Create chain
        llm = ChatOpenAI(model_name="gpt-4o", temperature=0)
        self.langchain_chain = (
            {"context": retriever, "question": RunnablePassthrough()}
            | prompt
            | llm
            | StrOutputParser()
        )
        return self.langchain_chain
    def setup_haystack_pipeline(self, documents):
        """Set up Haystack pipeline for production deployment"""
        from haystack import Pipeline
        from haystack.components.embedders import SentenceTransformersTextEmbedder
        from haystack.components.retrievers import InMemoryEmbeddingRetriever
        from haystack.components.generators import OpenAIGenerator
        from haystack.document_stores.in_memory import InMemoryDocumentStore
        # Initialize components
        document_store = InMemoryDocumentStore()
        embedder = SentenceTransformersTextEmbedder(model="sentence-transformers/all-MiniLM-L6-v2")
        retriever = InMemoryEmbeddingRetriever(document_store=document_store)
        generator = OpenAIGenerator(model="gpt-4o")
        # Create pipeline
        pipeline = Pipeline()
        pipeline.add_component("embedder", embedder)
        pipeline.add_component("retriever", retriever)
        pipeline.add_component("generator", generator)
        # Connect components
        pipeline.connect("embedder", "retriever")
        pipeline.connect("retriever", "generator")
        # Add documents
        document_store.write_documents(documents)
        self.haystack_pipeline = pipeline
        return pipeline
    def evaluate_system(self, test_questions, test_answers=None):
        """Evaluate system using RAGAS"""
        data = {
            "question": test_questions,
            "answer": [],
            "contexts": []
        }
        if test_answers:
            data["ground_truths"] = []
        # Generate answers
        for question in test_questions:
            answer = self.langchain_chain.invoke(question)
            contexts = self.llamaindex_index.as_retriever().invoke(question)
            data["answer"].append(answer)
            data["contexts"].append([ctx.text for ctx in contexts])
            if test_answers:
                data["ground_truths"].append([test_answers[test_questions.index(question)]])
        # Create dataset
        dataset = Dataset.from_dict(data)
        # Evaluate
        from ragas.metrics import faithfulness, answer_relevancy, context_precision
        metrics = [faithfulness, answer_relevancy, context_precision]
        if test_answers:
            from ragas.metrics import context_recall
            metrics.append(context_recall)
        result = evaluate(dataset, metrics=metrics)
        self.evaluation_metrics.append(result)
        return result
# Example usage
rag_system = ProductionRAGSystem()
# 1. Preprocess documents with Docling
processed_docs = rag_system.preprocess_documents(["doc1.pdf", "doc2.pdf"])
# 2. Build index with LlamaIndex
index = rag_system.build_llamaindex_index(processed_docs)
# 3. Create LangChain chain
chain = rag_system.create_langchain_chain(index)
# 4. Set up Haystack pipeline for production
haystack_pipeline = rag_system.setup_haystack_pipeline(processed_docs)
# 5. Evaluate system
test_questions = ["What are the system requirements?", "How does authentication work?"]
test_answers = ["The system requires 8GB RAM minimum", "Authentication uses OAuth 2.0"]
evaluation = rag_system.evaluate_system(test_questions, test_answers)
print(evaluation)

Аспекты адкалікавання і стаўкі

1. Стаўкі обработкі дакументаў

# Best practices for document processing with Docling
from docling.document_converter import DocumentConverter
from docling.datamodel.pipeline_options import PipelineOptions
def optimize_docling_processing():
    """Optimize Docling for different document types"""
    # For scanned documents with OCR
    pipeline_options_ocr = PipelineOptions(
        do_ocr=True,
        ocr_engine="tesseract",
        ocr_language="eng"
    )
    # For clean digital documents
    pipeline_options_clean = PipelineOptions(
        do_ocr=False,
        extract_images=False  # Disable image extraction for text-only processing
    )
    # For documents with complex layouts
    pipeline_options_layout = PipelineOptions(
        do_ocr=True,
        ocr_engine="easyocr",  # More accurate but slower
        extract_tables=True,
        extract_images=True
    )
    return {
        "ocr": pipeline_options_ocr,
        "clean": pipeline_options_clean,
        "layout": pipeline_options_layout
    }
# Usage
optimizations = optimize_docling_processing()
converter = DocumentConverter(pipeline_options=optimizations["ocr"])

2. Оптымізацыя стратэгіі разбівання на часткі

# Advanced chunking strategies for different content types
from llama_index.core.node_parser import (
    SentenceSplitter,
    SemanticSplitterNodeParser,
    HierarchicalNodeParser
)
def create_optimal_chunking_strategy(content_type="general"):
    """Create optimal chunking strategy based on content type"""
    if content_type == "technical":
        # Technical documents need smaller chunks for precision
        return SentenceSplitter(
            chunk_size=512,
            chunk_overlap=64,
            paragraph_separator="\n\n",
            sentence_separator="\\n"
        )
    elif content_type == "legal":
        # Legal documents need to preserve entire clauses
        return SemanticSplitterNodeParser(
            buffer_size=1,
            breakpoint_percentile_threshold=95,
            embed_model="sentence-transformers/all-MiniLM-L6-v2"
        )
    elif content_type == "research":
        # Research papers benefit from hierarchical chunking
        return HierarchicalNodeParser.from_defaults(
            chunk_sizes=[2048, 512, 128]
        )
    else:
        # General purpose chunking
        return SentenceSplitter(
            chunk_size=1024,
            chunk_overlap=128
        )
# Usage
chunking_strategy = create_optimal_chunking_strategy("technical")

3. Аспекты развертвання у працэўным режыме

# Production deployment configuration for Haystack
from haystack import Pipeline
from haystack.components.embedders import SentenceTransformersTextEmbedder
from haystack.components.retrievers import InMemoryEmbeddingRetriever
from haystack.components.generators import OpenAIGenerator
from haystack.document_stores.in_memory import InMemoryDocumentStore
import os
def create_production_pipeline():
    """Create production-ready Haystack pipeline"""
    # Use environment variables for sensitive data
    api_key = os.getenv("OPENAI_API_KEY")
    model_name = os.getenv("LLM_MODEL_NAME", "gpt-4o")
    # Initialize components with production settings
    document_store = InMemoryDocumentStore()
    embedder = SentenceTransformersTextEmbedder(
        model="sentence-transformers/all-MiniLM-L6-v2",
        batch_size=32  # Optimize for throughput
    )
    retriever = InMemoryEmbeddingRetriever(
        document_store=document_store,
        top_k=5  # Return more results for ranking
    )
    generator = OpenAIGenerator(
        api_key=api_key,
        model=model_name,
        max_tokens=1024,
        temperature=0.7  # Balance creativity and accuracy
    )
    # Create pipeline
    pipeline = Pipeline()
    pipeline.add_component("embedder", embedder)
    pipeline.add_component("retriever", retriever)
    pipeline.add_component("generator", generator)
    # Connect components
    pipeline.connect("embedder", "retriever")
    pipeline.connect("retriever", "generator")
    return pipeline
# Add monitoring and logging
import logging
logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)
def monitored_pipeline_run(pipeline, query):
    """Run pipeline with monitoring"""
    logger.info(f"Processing query: {query}")
    try:
        result = pipeline.run({"retriever": {"query": query}})
        logger.info("Query processed successfully")
        return result
    except Exception as e:
        logger.error(f"Error processing query: {e}")
        raise

4. Ацэнка та постаўленне на вышэйшы рывак

# Continuous evaluation and improvement loop
from ragas import evaluate
from ragas.metrics import (
    faithfulness,
    answer_relevancy,
    context_precision,
    context_recall
)
from datasets import Dataset
import time
class RAGEvaluationSystem:
    def __init__(self):
        self.evaluation_history = []
        self.improvement_plan = {}
    def run_evaluation_cycle(self, rag_system, test_set, previous_results=None):
        """Run evaluation cycle and generate improvement plan"""
        # Evaluate current system
        evaluation_result = self.evaluate_system(rag_system, test_set)
        # Compare with previous results
        if previous_results:
            improvement = self.calculate_improvement(evaluation_result, previous_results)
            self.generate_improvement_plan(improvement, evaluation_result)
        # Record results
        self.evaluation_history.append({
            "timestamp": time.time(),
            "results": evaluation_result,
            "improvement_plan": self.improvement_plan.copy()
        })
        return evaluation_result
    def evaluate_system(self, rag_system, test_set):
        """Evaluate RAG system"""
        data = {
            "question": test_set["questions"],
            "answer": [],
            "contexts": [],
            "ground_truths": test_set["answers"]
        }
        # Generate answers
        for question in test_set["questions"]:
            answer = rag_system(question)
            contexts = rag_system.get_contexts(question) if hasattr(rag_system, 'get_contexts') else []
            data["answer"].append(answer)
            data["contexts"].append(contexts)
        # Create dataset
        dataset = Dataset.from_dict(data)
        # Evaluate
        metrics = [
            faithfulness,
            answer_relevancy,
            context_precision,
            context_recall
        ]
        result = evaluate(dataset, metrics=metrics)
        return result
    def calculate_improvement(self, current, previous):
        """Calculate improvement between evaluation cycles"""
        improvement = {}
        for metric in current.keys():
            if metric in previous:
                improvement[metric] = current[metric] - previous[metric]
        return improvement
    def generate_improvement_plan(self, improvement, current_results):
        """Generate improvement plan based on evaluation results"""
        # Identify areas needing improvement
        if improvement.get("faithfulness", 0) < 0.1:
            self.improvement_plan["faithfulness"] = "Improve retrieval quality by adjusting chunk size or using better embedding model"
        if improvement.get("answer_relevancy", 0) < 0.1:
            self.improvement_plan["answer_relevancy"] = "Improve prompt engineering or use more sophisticated answer generation techniques"
        if improvement.get("context_precision", 0) < 0.1:
            self.improvement_plan["context_precision"] = "Implement re-ranking or hybrid search to improve context selection"
        if improvement.get("context_recall", 0) < 0.1:
            self.improvement_plan["context_recall"] = "Increase top-k parameter or implement query expansion techniques"
# Example usage
evaluator = RAGEvaluationSystem()
# Define test set
test_set = {
    "questions": ["What are the system requirements?", "How does authentication work?"],
    "answers": [["The system requires 8GB RAM minimum"], ["Authentication uses OAuth 2.0"]]
}
# Run evaluation
evaluation_result = evaluator.run_evaluation_cycle(your_rag_system, test_set)
print(evaluation_result)
print("Improvement Plan:", evaluator.improvement_plan)

Вывык

Справакі та дадатковая літэратура

Чек-ліст для эксплуатацыі