Home / Articles / Practical notes: Building a Modern ISR Pipeline for Code Modality in Large

This article is published in English.

Practical notes: Building a Modern ISR Pipeline for Code Modality in Large

Operable walkthrough of Practical notes: Building a Modern ISR Pipeline for Code Modality in Large: contracts, checks, and drop-in code slots for teams shipping this pattern.

4540 words

The following notes reconstruct a practical path around “Building a Modern ISR Pipeline for Code Modality in Large Scale Agent Harnesses”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

The Problem Nobody Really Solved

The The Problem Nobody Really stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

Phase 1: Ingestion — Building the Three Pillars

The Phase 1 Ingestion Building stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

Step 1: Cloning and Preparing the Repository

The Step 1 Cloning and stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

git.Repo.clone_from(
    f"https://{token}@github.com/{owner}/{repo}",
    target_dir=f"~/.isr/repos/{owner}/{repo}",
    depth=None  # Full history — crucial for incremental updates
)

The Step 1 Cloning and stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Step 2: File Enumeration and Language Detection

For the Step 2 File Enumeration stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

Step 3: AST Chunking — The Hardest Problem

For the Step 3 AST Chunking stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

assert "".join(chunk.raw_text for chunk in chunks) == file.content

Step 4: Symbol Extraction — Building the Call Graph

For the Step 4 Symbol Extraction stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the Step 4 Symbol Extraction stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

class_pattern = r'class\s+(\w+)'
func_pattern_python = r'def (\w+)\('
func_pattern_go = r'func.*\('

Step 5: Header Building — Adding Metadata Context

When working through the Step 5 Header Building stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

# File: decode.go
# Class/Module: decoder
# Function: unmarshal(n *Node, out reflect.Value) (good bool)
# Description: Handles type dispatch for YAML → Go struct conversion

Step 6: Context Generation (Optional) — LLM Enrichment

When working through the Step 6 Context Generation stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

INPUT: [large function body]
OUTPUT: "Handles type dispatch for YAML to Go struct conversion.
         Routes based on the target type reflection and handles nil values."
embedding_input = contextual_text + raw_text

Step 7: Code Embedding — Converting Code to Vectors

When working through the Step 7 Code Embedding stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Step 7 Code Embedding stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Step 8: The Three-Store System — The Heart of the Hybrid Approach

The Step 8 The Three-Store stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

Store 1: Qdrant — Vector + Full-Text

The Store 1 Qdrant Vector stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

{
  "chunk_id": "a1b2c3d4e5f6:1",
  "repo_name": "go-yaml/yaml",
  "rel_path": "decode.go",
  "language": "go",
  "raw_text": "func (d *decoder) unmarshal(...) { ... }",
  "contextual_text": "Handles type dispatch for YAML...",
  "header": "# File: decode.go\n# Function: unmarshal...",
  "start_line": 340,
  "end_line": 390
}
point_id = uuid.uuid5(NAMESPACE_DNS, chunk_id).int % (2**63)

Store 2: BM25 — Pure Keyword Ranking

The Store 2 BM25 Pure stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve. The Store 2 BM25 Pure stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

# Build the BM25 index with full text
corpus = [chunk.get_text_for_embedding() for chunk in all_chunks]
bm25.fit(corpus)
# Store only the chunk IDs in the doc store
doc_store = [chunk.chunk_id for chunk in all_chunks]# Save both
pickle.dump(bm25, "index.pkl")
pickle.dump(doc_store, "doc_store.pkl")

Store 3: KùzuDB — The Property Graph

For the Store 3 K zuDB stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

CREATE NODE TABLE Symbol (
  fqn STRING PRIMARY KEY,
  file STRING,
  language STRING,
  kind STRING
)
CREATE REL TABLE Calls (FROM Symbol TO Symbol, confidence FLOAT)
CREATE REL TABLE Imports (FROM Symbol TO Symbol, confidence FLOAT)
CREATE REL TABLE Inherits (FROM Symbol TO Symbol, confidence FLOAT)
MATCH (seed:Symbol WHERE seed.fqn IN [list_of_fqns])
      -[*1..hops]->(reachable:Symbol)
RETURN DISTINCT reachable.fqn

Store 4: Hash Tracker — Enabling Incremental Ingestion

For the Store 4 Hash Tracker stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

CREATE TABLE chunks (
  chunk_id TEXT PRIMARY KEY,
  repo_name TEXT,
  rel_path TEXT,
  content_sha256 TEXT,
  ingested_at TIMESTAMP
);
CREATE TABLE repos (
  repo_name TEXT PRIMARY KEY,
  last_sha TEXT,
  last_ingested TIMESTAMP
);

Step 9: Orchestration — Bringing It All Together

For the Step 9 Orchestration Bringing stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the Step 9 Orchestration Bringing stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Phase 2: Search — Three Signals, One Ranked List

When working through the Phase 2 Search Three stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

The Query Router — Zero-Cost Classification

When working through the The Query Router Zero-Cost stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

"how does ...", "explain how ...", "walk me through ...",
"trace ...", "step by step", "what happens when ..."
"what calls ...", "who calls ...", "callers of ...",
"what imports ...", "subclasses of ...", "where is X used ..."
camelCase like "parseTimestamp" or "NewDecoder"
snake_case like "parse_yaml"
quoted strings like "permission denied"
function calls like "handleErr()"

The EXACT_FIRST Short-Circuit — Grep

When working through the The EXACTFIRST Short-Circuit Grep stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours. When working through the The EXACTFIRST Short-Circuit Grep stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

rg --json --max-filesize 1M --context 5 --smart-case -- <query> <search_root>

Semantic Search — Vector Similarity

The Semantic Search Vector Similarity stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

Keyword Search — BM25 Lexical Matching

The Keyword Search BM25 Lexical stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

Graph Expansion — Structural Relationships

The Graph Expansion Structural Relationships stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve. The Graph Expansion Structural Relationships stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

MATCH (seed:Symbol WHERE seed.fqn IN [{seed_fqns}])
      -[*1..1]->(reachable:Symbol)
RETURN DISTINCT reachable.fqn
def _fqn_lookup_term(fqn: str) -> str:
    parts = fqn.split(".")
    # Use last two segments for specificity
    term = ".".join(parts[-2:])  # "PieceTree.applyDelta", not just "applyDelta"
    return term

RRF Fusion — Combining Signals

For the RRF Fusion Combining Signals stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

score(doc) = Σᵢ  1 / (k + rankᵢ(doc))
1/(60+1) + 1/(60+3) = 0.01639 + 0.01587 = 0.03226
1/(60+1) = 0.01639

Reranking — Deep Cross-Encoder Scoring

For the Reranking Deep Cross-Encoder Scoring stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

client.rerank(query, documents, model="rerank-english-v3.0", top_n=20)

The Agentic Loop — ORA (Observe → Reason → Act)

For the The Agentic Loop ORA stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the The Agentic Loop ORA stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

INPUT: "How does YAML error handling work?"
OUTPUT: "yaml error handling failf TypeError panic propagate"

SearchPipeline — The Orchestrator

When working through the SearchPipeline The Orchestrator stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

Phase 3: Retrieval — Context Assembly and Answer Synthesis

When working through the Phase 3 Retrieval Context stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

ContextAssembler — Building the Context Window

When working through the ContextAssembler Building the Context stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours. When working through the ContextAssembler Building the Context stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

<context id="1" file="decode.go" lines="340-390"
         repo="go-yaml/yaml" source="qdrant" expansion="function">
func newDecoder() *decoder {
    ...
}
</context>

AnswerSynthesizer — LLM Generation with Citations

The AnswerSynthesizer LLM Generation with stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

You are a code intelligence assistant. Answer questions about source code precisely.
Rules:
- Answer ONLY from the provided code context — never hallucinate code or APIs
- Cite every file you reference using [path/to/file.go:start-end] format
- Show working code examples from the context, not just descriptions
- If context is insufficient, say exactly: "Insufficient context: [what is missing]"
\[([^:\]\s][^:\]]*):(\d+)(?:-(\d+))?\]

RetrievalPipeline — The Final Orchestrator

The RetrievalPipeline The Final Orchestrator stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

retrieval = RetrievalPipeline(
    search_pipeline=SearchPipeline(...),
    context_assembler=ContextAssembler(search_pipeline.qdrant),
    answer_synthesizer=AnswerSynthesizer(...),
)

API and CLI

The API and CLI stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

HTTP API

The HTTP API stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

{
  "query": "parseTimestamp",
  "repo_name": "go-yaml/yaml",
  "top_k": 10,
  "use_graph": false,
  "use_rerank": true
}
{
  "question": "how does yaml error handling work?",
  "repo_name": "go-yaml/yaml",
  "mode": "answer",
  "max_context_chunks": 10
}

CLI

The CLI stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

# Ingest a repository
isr ingest go-yaml/yaml --with-context
# Search
isr search "what calls Unmarshal" --repo go-yaml/yaml# Retrieve with answer
isr retrieve "how does YAML error handling work?" \
  --repo go-yaml/yaml --mode answer --verbose# Start the server
isr server

Configuration and Tuning

The Configuration and Tuning stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

Evaluation and Testing

The Evaluation and Testing stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.

Conclusion: Why This Matters

The Conclusion Why This Matters stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve. The Conclusion Why This Matters stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Operational checklist

For the Operational checklist stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.

Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.

Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for 2f4a57bf7527: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.