Home / Articles / Practical notes: Data Cleaning and Preprocessing for Production RAG: Building

This article is published in English.

Practical notes: Data Cleaning and Preprocessing for Production RAG: Building

Operable walkthrough of Practical notes: Data Cleaning and Preprocessing for Production RAG: Building: contracts, checks, and drop-in code slots for teams shipping this pattern.

4234 words

Use this as an operator-facing rebuild of the ideas in “Data Cleaning and Preprocessing for Production RAG: Building Retrieval-Ready Documents”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

CONFIDENTIAL - INTERNAL USE ONLY
Group Financial Crime Compliance
Page 47 of 132

Where Data Cleaning Fits in the RAG Pipeline

For the Where Data Cleaning Fits stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

Encoding Normalization: Fix the Text Before You Analyze It

For the Encoding Normalization Fix the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

Customer’s identity must be verified before account opening.
The customer must provide proof of address.
Enhanced Due Diligence â€" High Risk Customers

Repairing Broken Unicode with ftfy

For the Repairing Broken Unicode with stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Repairing Broken Unicode with stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

import re
import unicodedata
from dataclasses import dataclass, field
import ftfy

@dataclass
class TextCleaningResult:
    original_text: str
    cleaned_text: str
    transformations: list[str] = field(default_factory=list)
    quality_flags: list[str] = field(default_factory=list)
def normalize_encoding(text: str) -> TextCleaningResult:
    """
    Repair common encoding problems while preserving meaningful Unicode.
    Suitable for policy, AML/KYC, payment, and regulatory documents where
    accented names, currency symbols, and multilingual text must survive.
    """
    transformations: list[str] = []
    quality_flags: list[str] = []
    if not text:
        return TextCleaningResult(original_text=text, cleaned_text=text)
    cleaned = text
    # Repair mojibake and common Unicode encoding mistakes.
    repaired = ftfy.fix_text(cleaned)
    if repaired != cleaned:
        transformations.append("ftfy_encoding_repair")
        cleaned = repaired
    # NFC preserves characters while producing a consistent Unicode
    # representation. Safer than aggressively stripping accents.
    normalized = unicodedata.normalize("NFC", cleaned)
    if normalized != cleaned:
        transformations.append("unicode_nfc_normalization")
        cleaned = normalized
    # Replace non-breaking spaces with normal spaces.
    if "\u00a0" in cleaned:
        cleaned = cleaned.replace("\u00a0", " ")
        transformations.append("non_breaking_space_normalization")
    # Remove zero-width characters that frequently leak from PDFs,
    # web pages, and copied Office content.
    zero_width_chars = {"\u200b", "\u200c", "\u200d", "\ufeff"}
    if any(char in cleaned for char in zero_width_chars):
        cleaned = "".join(char for char in cleaned if char not in zero_width_chars)
        transformations.append("zero_width_character_removal")
    # Remove control characters; preserve newline and tab because
    # they may still carry document structure needed downstream.
    cleaned_without_controls = "".join(
        char for char in cleaned
        if char in "\n\t" or unicodedata.category(char) != "Cc"
    )
    if cleaned_without_controls != cleaned:
        transformations.append("control_character_removal")
        cleaned = cleaned_without_controls
    # Normalize horizontal whitespace without flattening paragraphs.
    whitespace_normalized = re.sub(r"[ \t]+", " ", cleaned)
    whitespace_normalized = re.sub(r"\n{3,}", "\n\n", whitespace_normalized)
    if whitespace_normalized != cleaned:
        transformations.append("whitespace_normalization")
        cleaned = whitespace_normalized
    cleaned = cleaned.strip()
    # Keep suspicious replacement characters observable rather than silently deleting them.
    if "\ufffd" in cleaned:
        quality_flags.append("unicode_replacement_character_detected")
    return TextCleaningResult(
        original_text=text,
        cleaned_text=cleaned,
        transformations=transformations,
        quality_flags=quality_flags,
    )

Why Aggressive Text Cleaning Hurts RAG

When working through the Why Aggressive Text Cleaning stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Customer must NOT be classified as low risk.

Boilerplate Removal: Structure-Aware, Not Keyword-Aware

When working through the Boilerplate Removal Structure-Aware Not stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

KYC Policy | Version 7.2 | Page 41 of 132
KYC Policy | Version 7.2 | Page 42 of 132
import re
from collections import Counter
from dataclasses import dataclass

@dataclass
class BoilerplatePattern:
    normalized_text: str
    occurrences: int
    page_ratio: float
    position: str
def normalize_boilerplate_candidate(line: str) -> str:
    """
    Normalize variable fields so structurally identical headers and
    footers can be compared across pages.
    """
    normalized = line.strip()
    normalized = re.sub(
        r"\bpage\s+\d+\s+of\s+\d+\b", "page <n> of <n>",
        normalized, flags=re.IGNORECASE,
    )
    normalized = re.sub(r"\bpage\s+\d+\b", "page <n>", normalized, flags=re.IGNORECASE)
    normalized = re.sub(r"\s+", " ", normalized)
    return normalized.casefold()
def detect_repeated_marginal_text(
    pages: list[dict],
    margin_lines: int = 3,
    min_page_ratio: float = 0.6,
) -> list[BoilerplatePattern]:
    """
    Detect repeated text in page-header or page-footer regions.
    A candidate is considered boilerplate only when it appears in the same
    marginal position across a substantial fraction of the document.
    """
    if not pages:
        return []
    header_counts: Counter = Counter()
    footer_counts: Counter = Counter()
    for page in pages:
        lines = [line.strip() for line in page["text"].splitlines() if line.strip()]
        if not lines:
            continue
        header_counts.update(normalize_boilerplate_candidate(l) for l in lines[:margin_lines])
        footer_counts.update(normalize_boilerplate_candidate(l) for l in lines[-margin_lines:])
    total_pages = len(pages)
    patterns: list[BoilerplatePattern] = []
    for position, counts in (("header", header_counts), ("footer", footer_counts)):
        for normalized_text, occurrences in counts.items():
            page_ratio = occurrences / total_pages
            if page_ratio >= min_page_ratio:
                patterns.append(BoilerplatePattern(
                    normalized_text=normalized_text,
                    occurrences=occurrences,
                    page_ratio=page_ratio,
                    position=position,
                ))
    return patterns

Deduplication: Exact Matches Are the Easy Part

When working through the Deduplication Exact Matches Are stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Deduplication Exact Matches Are stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

AML_Policy_v8_Final.docx
AML_Policy_v8_Final_Approved.docx
AML_Policy_v8_Final_Copy.docx
AML_Policy_v8_Approved_2026.docx

Exact Deduplication with Content Hashing

The Exact Deduplication with Content stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

import hashlib
from dataclasses import dataclass

@dataclass
class DeduplicationRecord:
    document_id: str
    source: str
    content_hash: str
    duplicate_of: str | None = None
def canonicalize_for_hashing(text: str) -> str:
    """
    Deterministic normalization before hashing.
    Does not lowercase or remove punctuation because those transformations
    could collapse documents that are not actually identical.
    """
    lines = [" ".join(line.split()) for line in text.splitlines()]
    return "\n".join(lines).strip()
def calculate_content_hash(text: str) -> str:
    canonical_text = canonicalize_for_hashing(text)
    return hashlib.sha256(canonical_text.encode("utf-8")).hexdigest()
def find_exact_duplicates(
    documents: list[dict],
) -> tuple[list[dict], list[DeduplicationRecord]]:
    """
    Keep one canonical copy of each identical document while preserving
    duplicate lineage for auditability.
    """
    seen_hashes: dict[str, str] = {}
    unique_documents: list[dict] = []
    records: list[DeduplicationRecord] = []
    for document in documents:
        content_hash = calculate_content_hash(document["clean_text"])
        if content_hash in seen_hashes:
            records.append(DeduplicationRecord(
                document_id=document["document_id"],
                source=document["source"],
                content_hash=content_hash,
                duplicate_of=seen_hashes[content_hash],
            ))
            continue
        seen_hashes[content_hash] = document["document_id"]
        unique_documents.append(document)
        records.append(DeduplicationRecord(
            document_id=document["document_id"],
            source=document["source"],
            content_hash=content_hash,
        ))
    return unique_documents, records

Near-Duplicate Detection with MinHash LSH

The Near-Duplicate Detection with MinHash stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

import re
from datasketch import MinHash, MinHashLSH

def create_word_shingles(text: str, shingle_size: int = 5) -> set[str]:
    """
    Five-word shingles work well for long policy and regulatory documents
    because they capture local textual structure without being overly
    sensitive to isolated formatting changes.
    """
    tokens = re.findall(r"\b\w+\b", text.casefold())
    if len(tokens) < shingle_size:
        return {" ".join(tokens)} if tokens else set()
    return {
        " ".join(tokens[i:i + shingle_size])
        for i in range(len(tokens) - shingle_size + 1)
    }
def create_minhash(text: str, num_perm: int = 128) -> MinHash:
    shingles = create_word_shingles(text)
    minhash = MinHash(num_perm=num_perm)
    for shingle in shingles:
        minhash.update(shingle.encode("utf-8"))
    return minhash
def build_duplicate_index(
    documents: list[dict],
    threshold: float = 0.85,
    num_perm: int = 128,
) -> tuple[MinHashLSH, dict[str, MinHash]]:
    """
    Build an LSH index for candidate near-duplicate discovery.
    The threshold identifies candidates. It does not automatically
    determine whether a document is deleted.
    """
    lsh = MinHashLSH(threshold=threshold, num_perm=num_perm)
    signatures: dict[str, MinHash] = {}
    for document in documents:
        document_id = document["document_id"]
        signature = create_minhash(document["clean_text"], num_perm=num_perm)
        signatures[document_id] = signature
        lsh.insert(document_id, signature)
    return lsh, signatures

OCR Error Correction: When the Text Looks Right but Is Wrong

The OCR Error Correction When stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The OCR Error Correction When stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

import re
from dataclasses import dataclass

@dataclass
class OCRIssue:
    token: str
    issue_type: str
    risk_level: str
    suggested_value: str | None
    requires_review: bool
CURRENCY_PATTERN = re.compile(r"\b(EUR|USD|GBP|INR)\s+([A-Za-z0-9.,]+)\b")
PERCENTAGE_PATTERN = re.compile(r"\b([A-Za-z0-9.,]+)\s*%")
def detect_numeric_ocr_issues(text: str) -> list[OCRIssue]:
    """
    Detect suspicious OCR substitutions inside financial values.
    This function identifies candidates. It does not silently rewrite
    high-risk values in the source text.
    """
    issues: list[OCRIssue] = []
    substitutions = {"O": "0", "o": "0", "I": "1", "l": "1"}
    def inspect_numeric_token(token: str, issue_type: str) -> None:
        if token.replace(",", "").replace(".", "").isdigit():
            return
        corrected = "".join(substitutions.get(char, char) for char in token)
        numeric_candidate = corrected.replace(",", "").replace(".", "")
        if numeric_candidate.isdigit() and corrected != token:
            issues.append(OCRIssue(
                token=token,
                issue_type=issue_type,
                risk_level="high",
                suggested_value=corrected,
                requires_review=True,
            ))
    for match in CURRENCY_PATTERN.finditer(text):
        inspect_numeric_token(match.group(2), issue_type="currency_value")
    for match in PERCENTAGE_PATTERN.finditer(text):
        inspect_numeric_token(match.group(1), issue_type="percentage")
    return issues

Multilingual Preprocessing: One Document, Multiple Languages

For the Multilingual Preprocessing One Document stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

from dataclasses import dataclass
from langdetect import DetectorFactory, LangDetectException, detect_langs

DetectorFactory.seed = 0  # Makes pipeline output reproducible.
@dataclass
class LanguagePrediction:
    language: str
    confidence: float
    alternatives: list[tuple[str, float]]
    needs_review: bool
def detect_document_language(
    text: str,
    confidence_threshold: float = 0.85,
) -> LanguagePrediction:
    sample = text.strip()
    if not sample:
        return LanguagePrediction(language="unknown", confidence=0.0, alternatives=[], needs_review=True)
    try:
        predictions = detect_langs(sample)
    except LangDetectException:
        return LanguagePrediction(language="unknown", confidence=0.0, alternatives=[], needs_review=True)
    best = predictions[0]
    alternatives = [(p.lang, round(p.prob, 4)) for p in predictions]
    return LanguagePrediction(
        language=best.lang,
        confidence=best.prob,
        alternatives=alternatives,
        needs_review=best.prob < confidence_threshold,
    )
metadata = {
    "primary_language": "en",
    "languages": ["en", "de", "fr"],
    "language_distribution": {"en": 0.72, "de": 0.21, "fr": 0.07},
    "is_multilingual": True,
}

PII and Sensitive Data: Detection Before Embedding

For the PII and Sensitive Data stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

from dataclasses import dataclass, field
from presidio_analyzer import AnalyzerEngine

@dataclass
class SensitiveEntity:
    entity_type: str
    start: int
    end: int
    score: float
    original_value: str
@dataclass
class PIIScanResult:
    entities: list[SensitiveEntity] = field(default_factory=list)
    needs_review: bool = False
class BankingPIIDetector:
    def __init__(self) -> None:
        self.analyzer = AnalyzerEngine()
    def detect(
        self,
        text: str,
        language: str = "en",
        score_threshold: float = 0.70,
    ) -> PIIScanResult:
        """
        Detect sensitive entities without modifying the source text.
        Detection is deliberately separated from anonymization because
        different entity types require different handling policies.
        """
        results = self.analyzer.analyze(
            text=text, language=language, score_threshold=score_threshold
        )
        entities = [
            SensitiveEntity(
                entity_type=r.entity_type,
                start=r.start,
                end=r.end,
                score=r.score,
                original_value=text[r.start:r.end],
            )
            for r in results
        ]
        high_risk_types = {"CREDIT_CARD", "IBAN_CODE", "US_SSN", "IP_ADDRESS"}
        return PIIScanResult(
            entities=entities,
            needs_review=any(e.entity_type in high_risk_types for e in entities),
        )
Customer IBAN: <IBAN>
Customer Email: <EMAIL_ADDRESS>

Putting It Together: A Production-Grade Cleaning Pipeline

For the Putting It Together A stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Putting It Together A stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

from dataclasses import dataclass, field
from enum import Enum
from typing import Any

class QualityStatus(str, Enum):
    PASS = "pass"
    REVIEW = "review"
    REJECT = "reject"
@dataclass
class Transformation:
    stage: str
    operation: str
    details: dict[str, Any] = field(default_factory=dict)
@dataclass
class QualityFlag:
    code: str
    severity: str
    details: dict[str, Any] = field(default_factory=dict)
@dataclass
class RetrievalReadyDocument:
    document_id: str
    source: str
    raw_text: str
    clean_text: str
    language: str | None = None
    language_confidence: float | None = None
    content_hash: str | None = None
    duplicate_of: str | None = None
    transformations: list[Transformation] = field(default_factory=list)
    quality_flags: list[QualityFlag] = field(default_factory=list)
    metadata: dict[str, Any] = field(default_factory=dict)
    status: QualityStatus = QualityStatus.PASS
import logging
logger = logging.getLogger(__name__)

class ProductionDocumentCleaner:
    def __init__(
        self,
        pii_detector: BankingPIIDetector,
        pii_anonymizer,
        redact_pii_for_embeddings: bool = True,
    ) -> None:
        self.pii_detector = pii_detector
        self.pii_anonymizer = pii_anonymizer
        self.redact_pii_for_embeddings = redact_pii_for_embeddings
    def clean(self, document: dict) -> RetrievalReadyDocument:
        result = RetrievalReadyDocument(
            document_id=document["document_id"],
            source=document["source"],
            raw_text=document["text"],
            clean_text=document["text"],
            metadata=document.get("metadata", {}).copy(),
        )
        try:
            self._normalize_text(result)
            self._detect_language(result)
            self._validate_ocr(result)
            self._handle_sensitive_data(result)
            self._calculate_hash(result)
            self._apply_quality_gate(result)
        except Exception as exc:
            logger.exception("Preprocessing failed for document %s", result.document_id)
            result.quality_flags.append(QualityFlag(
                code="PREPROCESSING_EXCEPTION",
                severity="critical",
                details={"error": str(exc)},
            ))
            result.status = QualityStatus.REJECT
        return result
    def _apply_quality_gate(self, document: RetrievalReadyDocument) -> None:
        severities = {flag.severity for flag in document.quality_flags}
        if "critical" in severities:
            document.status = QualityStatus.REJECT
        elif "high" in severities:
            document.status = QualityStatus.REVIEW
        else:
            document.status = QualityStatus.PASS

Production Trade-offs

When working through the Production Trade-offs stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

What This Means in Practice

When working through the What This Means in stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

A Production Checklist Before You Start Chunking

When working through the A Production Checklist Before stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the A Production Checklist Before stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Operational checklist

When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.

Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Freeze a golden set before changing prompts or models. Moving both the system and the yardstick hides regressions.

Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.

Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for af00655fb2bc: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.

The hardening note 0 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Hardening detail 0/902: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 1 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Hardening detail 1/902: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

When working through the hardening note 2 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Hardening detail 2/902: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 3 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Hardening detail 3/902: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 4 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Hardening detail 4/902: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

When working through the hardening note 5 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Hardening detail 5/902: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 6 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Hardening detail 6/902: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 7 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Hardening detail 7/902: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

When working through the hardening note 8 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Hardening detail 8/902: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.