Home / Articles / RAG From First Principles to Three Bedrock Agents

This article is published in English.

RAG From First Principles to Three Bedrock Agents

Why weights fail on private data, how embeddings and cosine fetch work, and three progressive agents from managed KB to self-managed ingest.

4464 words

Frozen foundation models cannot answer private business questions from weights alone. Training never saw internal HR policies, contracts, or runbooks; cutoffs freeze time; and when facts are missing the generator still produces fluent guesses. Teaching a network what a company knows therefore becomes a systems problem, not a bigger-checkpoint problem.

Why weights alone fail

Three structural limits block enterprise Q&A:

  1. Private corpora never entered training. Retention rules and ticket macros are not public web text.
  2. Time stops at the cutoff. Last month's legal rewrite is invisible to a frozen snapshot.
  3. Silence is not the default. Next-token prediction prefers a plausible continuation over an honest unknown.

Three ways to inject knowledge—and which scales

Option 1: Fine-tune on the corpus

Further training can imprint style or skills, yet smears facts across parameters without citations, and cannot enforce per-user visibility because everyone shares the same weights. For policy Q&A that must quote and gate access, fine-tuning is the wrong mechanism.

Option 2: Paste everything into the system instruction

Handing text at ask-time converts recall into reading—which is the right mechanism. For tiny stable packs it works. It collapses when estates exceed the window, when every token is billed on every ask, when buried mid-prompt answers degrade, and when one static paste cannot filter by role.

Option 3: Paste only what matters

Keep the read-at-ask-time mechanism; fix the scale. Fetch a handful of relevant paragraphs, then generate. That is retrieval-augmented generation in one sentence.

What RAG is, end to end

Offline: ingest files, split into segments, embed segments, store vectors (and often lexical indexes). Online: embed the question, fetch top segments, pack them into an instruction with the question, generate an answer grounded in those segments.

The hard part is not the last generate call—it is making fetching precise, cheap, and permission-aware.

Vectors and embeddings without mysticism

A vector here is a list of numbers placing a text snippet in a space where near-meaning items sit near each other. An embedding network maps text to that space. Cosine similarity compares direction more than length, which is why it dominates retrieval scoring. Distance metrics that ignore normalization behave worse across uneven segment lengths.

Choosing dimensionality and the embedder is a product decision: more dimensions can separate meanings but cost storage and latency; mixing embedders between index time and query time silently breaks nearest-neighbor search.

Arithmetic of similarity

Cosine is the dot product of unit vectors. Two segments about password resets with different wording still land close if the embedder captured intent. Keyword search misses that link; dense fetch catches it—and may miss exact SKU tokens unless hybrid lexical search joins the party.

Building three complete agents on managed cloud inference

The rest of this guide materializes the ideas as three progressive agents on AWS Bedrock-style managed models: a curated knowledge base, a LangGraph orchestrator, and a self-managed ingest pipeline. Each listing is a full agent sketch—read them as executable architecture, not snippets.

Agent 1 — Managed knowledge base path

"""
AGENT 1: Managed Knowledge Base. The 30-line version.
Bedrock does the chunking, embedding, indexing and searching. You wire it up.
"""
import boto3
from langchain_aws import AmazonKnowledgeBasesRetriever, ChatBedrockConverse
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser
from langchain_core.runnables import RunnableParallel, RunnablePassthrough

REGION = "eu-west-1"
KB_ID = "ABCD1234EF"

retriever = AmazonKnowledgeBasesRetriever(
    knowledge_base_id=KB_ID,
    region_name=REGION,
    retrieval_config={
        "vectorSearchConfiguration": {
            "numberOfResults": 5,
            "overrideSearchType": "HYBRID",
        }
    },
)

llm = ChatBedrockConverse(
    model="eu.anthropic.claude-sonnet-4-5-20250929-v1:0",
    region_name=REGION,
    temperature=0,
    max_tokens=1024,
)

PROMPT = ChatPromptTemplate.from_messages([
    ("system",
     "Answer only from the numbered context. Cite every claim as [n]. "
     "If the context does not contain the answer, say so plainly."),
    ("human", "Context:\n{context}\n\nQuestion: {question}"),
])


def format_docs(docs) -> str:
    return "\n\n".join(
        f"[{i}] source={d.metadata.get('location', {}).get('s3Location', {}).get('uri', '?')}\n"
        f"{d.page_content}"
        for i, d in enumerate(docs, start=1)
    )


chain = (
    RunnableParallel(context=retriever | format_docs, question=RunnablePassthrough())
    | PROMPT
    | llm
    | StrOutputParser()
)

if __name__ == "__main__":
    print(chain.invoke("How fast must we submit a shortlist for a niche engineering role?"))

This path leans on a hosted knowledge base: documents land in storage, a managed indexer builds embeddings, and the agent answers with citations returned by that service. Operational wins include less custom index code; tradeoffs include less control over segment boundaries and fusion.

Agent 2 — LangGraph orchestration over tools

"""
AGENT 2: LangGraph agent that decides WHETHER and WHICH index to search.
Retrieval is a tool. Identity travels in typed context, never in the prompt.
"""
from __future__ import annotations
from dataclasses import dataclass
from typing import Any

import boto3
from langchain.agents import create_agent
from langchain.agents.middleware import ModelCallLimitMiddleware, ToolCallLimitMiddleware
from langchain.tools import ToolRuntime
from langchain_aws import ChatBedrockConverse
from langchain_core.tools import tool

REGION = "eu-west-1"
POLICY_KB = "ABCD1234EF"
ENGINEERING_KB = "WXYZ5678GH"

agent_runtime = boto3.client("bedrock-agent-runtime", region_name=REGION)


@dataclass
class RequestContext:
    """Built from validated JWT claims. The model cannot read or forge it."""
    subject: str
    roles: tuple[str, ...]
    region: str


def _retrieve(kb_id: str, query: str, ctx: RequestContext, k: int = 5) -> str:
    """Bedrock Retrieve API with a server-side metadata filter for tenant + residency."""
    resp = agent_runtime.retrieve(
        knowledgeBaseId=kb_id,
        retrievalQuery={"text": query},
        retrievalConfiguration={
            "vectorSearchConfiguration": {
                "numberOfResults": k,
                "overrideSearchType": "HYBRID",
                "filter": {
                    "andAll": [
                        {"in": {"key": "acl_roles", "value": list(ctx.roles)}},
                        {"in": {"key": "region", "value": [ctx.region, "GLOBAL"]}},
                    ]
                },
            }
        },
    )
    results: list[dict[str, Any]] = resp.get("retrievalResults", [])
    if not results:
        return "NO_AUTHORISED_RESULTS: nothing in this index that you may read."
    return "\n\n".join(
        f"[{i}] score={r.get('score'):.3f} "
        f"source={r.get('location', {}).get('s3Location', {}).get('uri', '?')}\n"
        f"{r['content']['text']}"
        for i, r in enumerate(results, start=1)
    )


@tool
def search_policy_docs(query: str, runtime: ToolRuntime[RequestContext]) -> str:
    """Search HR, compliance and commercial policy.

    Use for shortlist SLAs and service credits, candidate data retention and consent,
    GDPR erasure, statement-of-work approval thresholds, background check rules and
    automated decision governance.
    """
    return _retrieve(POLICY_KB, query, runtime.context)


@tool
def search_engineering_docs(query: str, runtime: ToolRuntime[RequestContext]) -> str:
    """Search platform runbooks and integration documentation.

    Use for the requisition sync service, Workday and SuccessFactors webhooks, Kafka
    topics and event schemas, idempotency and retry behaviour, and vector index refresh.
    """
    return _retrieve(ENGINEERING_KB, query, runtime.context)


agent = create_agent(
    model=ChatBedrockConverse(
        model="eu.anthropic.claude-sonnet-4-5-20250929-v1:0",
        region_name=REGION,
        temperature=0,
    ),
    tools=[search_policy_docs, search_engineering_docs],
    system_prompt=(
        "You are an enterprise knowledge assistant.\n"
        "- Answer only from tool results and cite the [n] markers they return.\n"
        "- If a tool returns NO_AUTHORISED_RESULTS, say you cannot access a source and "
        "name the document owner. Never fill the gap from memory.\n"
        "- For questions spanning policy and platform, call both tools before answering."
    ),
    context_schema=RequestContext,
    middleware=[
        ModelCallLimitMiddleware(thread_limit=8, run_limit=6, exit_behavior="end"),
        ToolCallLimitMiddleware(thread_limit=6, run_limit=5, exit_behavior="end"),
    ],
)

if __name__ == "__main__":
    claims = {"sub": "asha@example.com", "roles": ["recruiter"], "region": "EU"}
    result = agent.invoke(
        {"messages": [{"role": "user", "content": "How is the vector index refreshed?"}]},
        context=RequestContext(
            subject=claims["sub"],
            roles=tuple(claims["roles"]),
            region=claims["region"],
        ),
    )
    print(result["messages"][-1].content)

Here control flow is an explicit graph: classify intent, call retrieval tools, optionally ask a human, then respond. Checkpointing and interrupts turn long-running support threads into durable sessions instead of in-memory dicts.

Agent 3 — Self-managed ingest and fetch

"""
AGENT 3: Self-managed pipeline: your chunking, your index, your fusion, your rerank.
Chosen when Bedrock KB's fixed options are not enough (custom fusion, per-stage timing,
an abstention gate, or a chunking rule the managed service will not express).
"""
from __future__ import annotations
import time
from typing import Annotated, Any, TypedDict

import boto3
from langchain_aws import BedrockEmbeddings, ChatBedrockConverse
from langchain_core.documents import Document
from langchain_core.prompts import ChatPromptTemplate
from langchain_text_splitters import MarkdownHeaderTextSplitter, RecursiveCharacterTextSplitter
from langgraph.graph import END, START, StateGraph

REGION = "eu-west-1"
CONFIDENCE_FLOOR = 0.34
CANDIDATE_K, CONTEXT_K = 20, 4

embeddings = BedrockEmbeddings(
    model_id="amazon.titan-embed-text-v2:0",
    region_name=REGION,
    model_kwargs={"dimensions": 1024, "normalize": True},
)
llm = ChatBedrockConverse(
    model="eu.anthropic.claude-sonnet-4-5-20250929-v1:0",
    region_name=REGION,
    temperature=0,
)
rerank_client = boto3.client("bedrock-agent-runtime", region_name=REGION)


# ---------------------------------------------------------------- ingest
def chunk_policy_document(markdown: str, source: str, acl: list[str], region: str):
    """Structure-aware first, size-capped second. Headings are the natural unit for
    policy text; the recursive splitter only intervenes on oversized sections."""
    by_heading = MarkdownHeaderTextSplitter(
        headers_to_split_on=[("#", "title"), ("##", "section"), ("###", "clause")]
    ).split_text(markdown)

    capper = RecursiveCharacterTextSplitter(chunk_size=2000, chunk_overlap=200)
    chunks: list[Document] = []
    for part in capper.split_documents(by_heading):
        heading = " > ".join(
            v for k, v in part.metadata.items() if k in ("title", "section", "clause")
        )
        chunks.append(
            Document(
                # contextual retrieval: the chunk carries its own scope
                page_content=f"{heading}\n{part.page_content}",
                metadata={**part.metadata, "source": source, "acl": acl, "region": region},
            )
        )
    return chunks


# ---------------------------------------------------------------- state
class State(TypedDict, total=False):
    question: str
    roles: list[str]
    region: str
    candidates: list[Document]
    context_docs: list[Document]
    confidence: float
    answer: str
    timings: Annotated[dict[str, float], lambda a, b: {**(a or {}), **(b or {})}]


def _ms(name: str, t0: float) -> dict[str, float]:
    return {name: round((time.perf_counter() - t0) * 1000, 1)}


# ---------------------------------------------------------------- nodes
def retrieve(state: State, *, store) -> State:
    """Hybrid search with an ACL pre-filter pushed into the index query."""
    t0 = time.perf_counter()
    docs = store.similarity_search(
        state["question"],
        k=CANDIDATE_K,
        filter={"acl": {"$in": state["roles"]}, "region": {"$in": [state["region"], "GLOBAL"]}},
    )
    return {"candidates": docs, "timings": _ms("retrieve", t0)}


def rerank(state: State) -> State:
    """Bedrock Rerank: a cross-encoder scoring each (query, chunk) pair jointly."""
    t0 = time.perf_counter()
    cands = state["candidates"]
    if not cands:
        return {"context_docs": [], "confidence": 0.0, "timings": _ms("rerank", t0)}

    resp = rerank_client.rerank(
        queries=[{"type": "TEXT", "textQuery": {"text": state["question"]}}],
        sources=[
            {"type": "INLINE", "inlineDocumentSource": {
                "type": "TEXT", "textDocument": {"text": d.page_content}}}
            for d in cands
        ],
        rerankingConfiguration={
            "type": "BEDROCK_RERANKING_MODEL",
            "bedrockRerankingConfiguration": {
                "numberOfResults": CONTEXT_K,
                "modelConfiguration": {
                    "modelArn": f"arn:aws:bedrock:{REGION}::foundation-model/amazon.rerank-v1:0"
                },
            },
        },
    )
    ranked = resp["results"]
    top = [cands[r["index"]] for r in ranked]
    return {
        "context_docs": top,
        "confidence": round(float(ranked[0]["relevanceScore"]), 3),
        "timings": _ms("rerank", t0),
    }


def generate(state: State) -> State:
    """Abstain below the floor rather than answering from weak context."""
    t0 = time.perf_counter()
    if not state["context_docs"] or state["confidence"] < CONFIDENCE_FLOOR:
        return {
            "answer": "I cannot ground an answer for that in a source you are permitted "
                      "to read. Routing to the document owner.",
            "timings": _ms("generate", t0),
        }
    context = "\n\n".join(
        f"[{i}] {d.metadata.get('source')}\n{d.page_content}"
        for i, d in enumerate(state["context_docs"], start=1)
    )
    prompt = ChatPromptTemplate.from_messages([
        ("system", "Answer only from the numbered context and cite each claim as [n]."),
        ("human", "Context:\n{context}\n\nQuestion: {question}"),
    ])
    reply = (prompt | llm).invoke({"context": context, "question": state["question"]})
    return {"answer": reply.content, "timings": _ms("generate", t0)}


def build_graph(store):
    g = StateGraph(State)
    g.add_node("retrieve", lambda s: retrieve(s, store=store))
    g.add_node("rerank", rerank)
    g.add_node("generate", generate)
    g.add_edge(START, "retrieve")
    g.add_edge("retrieve", "rerank")
    g.add_edge("rerank", "generate")
    g.add_edge("generate", END)
    return g.compile()

This variant owns chunking, embedding calls, vector storage, and prompt assembly. Teams choose it when segment policy, hybrid lexical+dense fusion, or tenancy filters must be first-class code rather than console toggles.

Choosing among the three

Start managed when speed-to-demo matters and corpora are modest. Move to LangGraph when branching, human review, and memory across turns dominate. Own the pipeline when retrieval quality on identifiers, ACLs, or domain segmentation becomes the product.

Evaluation that actually moves quality

Track Recall@k and Precision@k on a frozen question set, plus faithfulness of answers to fetched segments. Pair every segment-size or embedder change with before/after scores. Refuse when top segments are weak—calibrated silence beats fluent invention.

When not to use RAG

Skip retrieval for trivia the base network already knows, for pure creative tasks, for continuously changing numbers better served by APIs, and for ultra-low-latency paths that cannot afford a fetch hop. Sometimes a better instruction or a small fine-tune of format is enough.

Closing

RAG is plumbing: clean intake, honest fetching, tight generation instructions, and measurable gates. The three agents show the same idea at three ownership levels. Pick the level that matches control needs—not the one that sounds most advanced in a slide.

Operational checklist: (1) one embedder for index and query, (2) metadata for ACL and dates from day one, (3) hybrid search when identifiers appear, (4) re-ranking when precision matters, (5) golden questions in CI, (6) rebuild triggers when sources change. Treat the knowledge plane as a product with owners, SLOs, and rollback—not a weekend index that nobody monitors. When agent graphs grow, measure token cost beside answer quality; orchestration that burns fifteen times a simple chat without lifting faithfulness is a regression dressed as architecture. Keep runbooks for empty-result refusals, stale-index incidents, and embedder mismatches so on-call is not improvising from chat logs.

Document every corpus source with an owner and refresh cadence; orphaned buckets become silent rot.

Prefer segment overlap early; only adopt semantic splitters when evaluations prove fixed windows fail.

Log fetch ids with answers so bad replies can be diagnosed as fetch misses versus generation drift.

Separate eval judges from the answering network when possible to reduce self-preference bias.

Budget API spend for indexing separately from interactive serving so finance sees both lines.

Version prompt templates next to index builds; a prompt bump without re-eval is how regressions ship.

For multi-tenant estates, test that filters cannot be bypassed by prompt injection claiming admin roles.

Keep a kill switch that disables generation and returns citations-only when faithfulness monitors drop.

Revisit hybrid lexical weights whenever product codes dominate traffic mix.

Train support staff to read citation panels before escalating 'wrong bot' tickets.

Document every corpus source with an owner and refresh cadence; orphaned buckets become silent rot.

Prefer segment overlap early; only adopt semantic splitters when evaluations prove fixed windows fail.

Log fetch ids with answers so bad replies can be diagnosed as fetch misses versus generation drift.

Separate eval judges from the answering network when possible to reduce self-preference bias.

Budget API spend for indexing separately from interactive serving so finance sees both lines.

Version prompt templates next to index builds; a prompt bump without re-eval is how regressions ship.

For multi-tenant estates, test that filters cannot be bypassed by prompt injection claiming admin roles.

Keep a kill switch that disables generation and returns citations-only when faithfulness monitors drop.

Revisit hybrid lexical weights whenever product codes dominate traffic mix.

Train support staff to read citation panels before escalating 'wrong bot' tickets.

Document every corpus source with an owner and refresh cadence; orphaned buckets become silent rot.

Prefer segment overlap early; only adopt semantic splitters when evaluations prove fixed windows fail.

Log fetch ids with answers so bad replies can be diagnosed as fetch misses versus generation drift.

Separate eval judges from the answering network when possible to reduce self-preference bias.

Budget API spend for indexing separately from interactive serving so finance sees both lines.

Version prompt templates next to index builds; a prompt bump without re-eval is how regressions ship.

For multi-tenant estates, test that filters cannot be bypassed by prompt injection claiming admin roles.

Keep a kill switch that disables generation and returns citations-only when faithfulness monitors drop.

Revisit hybrid lexical weights whenever product codes dominate traffic mix.

Train support staff to read citation panels before escalating 'wrong bot' tickets.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.

Keep paired evaluations and refusal metrics beside latency so cost and quality move on one sheet.