This article is published in English.
Practical notes: Building a Modern ISR Pipeline for Code Modality in Large
Operable walkthrough of Practical notes: Building a Modern ISR Pipeline for Code Modality in Large: contracts, checks, and drop-in code slots for teams shipping this pattern.
The following notes reconstruct a practical path around “Building a Modern ISR Pipeline for Code Modality in Large Scale Agent Harnesses”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
The Problem Nobody Really Solved
The The Problem Nobody Really stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
Phase 1: Ingestion — Building the Three Pillars
The Phase 1 Ingestion Building stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
Step 1: Cloning and Preparing the Repository
The Step 1 Cloning and stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
git.Repo.clone_from(
f"https://{token}@github.com/{owner}/{repo}",
target_dir=f"~/.isr/repos/{owner}/{repo}",
depth=None # Full history — crucial for incremental updates
)
The Step 1 Cloning and stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Step 2: File Enumeration and Language Detection
For the Step 2 File Enumeration stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.
Step 3: AST Chunking — The Hardest Problem
For the Step 3 AST Chunking stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.
assert "".join(chunk.raw_text for chunk in chunks) == file.content
Step 4: Symbol Extraction — Building the Call Graph
For the Step 4 Symbol Extraction stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the Step 4 Symbol Extraction stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
class_pattern = r'class\s+(\w+)'
func_pattern_python = r'def (\w+)\('
func_pattern_go = r'func.*\('
Step 5: Header Building — Adding Metadata Context
When working through the Step 5 Header Building stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
# File: decode.go
# Class/Module: decoder
# Function: unmarshal(n *Node, out reflect.Value) (good bool)
# Description: Handles type dispatch for YAML → Go struct conversion
Step 6: Context Generation (Optional) — LLM Enrichment
When working through the Step 6 Context Generation stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
INPUT: [large function body]
OUTPUT: "Handles type dispatch for YAML to Go struct conversion.
Routes based on the target type reflection and handles nil values."
embedding_input = contextual_text + raw_text
Step 7: Code Embedding — Converting Code to Vectors
When working through the Step 7 Code Embedding stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Step 7 Code Embedding stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Step 8: The Three-Store System — The Heart of the Hybrid Approach
The Step 8 The Three-Store stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
Store 1: Qdrant — Vector + Full-Text
The Store 1 Qdrant Vector stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
{
"chunk_id": "a1b2c3d4e5f6:1",
"repo_name": "go-yaml/yaml",
"rel_path": "decode.go",
"language": "go",
"raw_text": "func (d *decoder) unmarshal(...) { ... }",
"contextual_text": "Handles type dispatch for YAML...",
"header": "# File: decode.go\n# Function: unmarshal...",
"start_line": 340,
"end_line": 390
}
point_id = uuid.uuid5(NAMESPACE_DNS, chunk_id).int % (2**63)
Store 2: BM25 — Pure Keyword Ranking
The Store 2 BM25 Pure stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve. The Store 2 BM25 Pure stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
# Build the BM25 index with full text
corpus = [chunk.get_text_for_embedding() for chunk in all_chunks]
bm25.fit(corpus)
# Store only the chunk IDs in the doc store
doc_store = [chunk.chunk_id for chunk in all_chunks]# Save both
pickle.dump(bm25, "index.pkl")
pickle.dump(doc_store, "doc_store.pkl")
Store 3: KùzuDB — The Property Graph
For the Store 3 K zuDB stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.
CREATE NODE TABLE Symbol (
fqn STRING PRIMARY KEY,
file STRING,
language STRING,
kind STRING
)
CREATE REL TABLE Calls (FROM Symbol TO Symbol, confidence FLOAT)
CREATE REL TABLE Imports (FROM Symbol TO Symbol, confidence FLOAT)
CREATE REL TABLE Inherits (FROM Symbol TO Symbol, confidence FLOAT)
MATCH (seed:Symbol WHERE seed.fqn IN [list_of_fqns])
-[*1..hops]->(reachable:Symbol)
RETURN DISTINCT reachable.fqn
Store 4: Hash Tracker — Enabling Incremental Ingestion
For the Store 4 Hash Tracker stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.
CREATE TABLE chunks (
chunk_id TEXT PRIMARY KEY,
repo_name TEXT,
rel_path TEXT,
content_sha256 TEXT,
ingested_at TIMESTAMP
);
CREATE TABLE repos (
repo_name TEXT PRIMARY KEY,
last_sha TEXT,
last_ingested TIMESTAMP
);
Step 9: Orchestration — Bringing It All Together
For the Step 9 Orchestration Bringing stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the Step 9 Orchestration Bringing stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Phase 2: Search — Three Signals, One Ranked List
When working through the Phase 2 Search Three stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
The Query Router — Zero-Cost Classification
When working through the The Query Router Zero-Cost stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
"how does ...", "explain how ...", "walk me through ...",
"trace ...", "step by step", "what happens when ..."
"what calls ...", "who calls ...", "callers of ...",
"what imports ...", "subclasses of ...", "where is X used ..."
camelCase like "parseTimestamp" or "NewDecoder"
snake_case like "parse_yaml"
quoted strings like "permission denied"
function calls like "handleErr()"
The EXACT_FIRST Short-Circuit — Grep
When working through the The EXACTFIRST Short-Circuit Grep stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours. When working through the The EXACTFIRST Short-Circuit Grep stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
rg --json --max-filesize 1M --context 5 --smart-case -- <query> <search_root>
Semantic Search — Vector Similarity
The Semantic Search Vector Similarity stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
Keyword Search — BM25 Lexical Matching
The Keyword Search BM25 Lexical stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
Graph Expansion — Structural Relationships
The Graph Expansion Structural Relationships stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve. The Graph Expansion Structural Relationships stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
MATCH (seed:Symbol WHERE seed.fqn IN [{seed_fqns}])
-[*1..1]->(reachable:Symbol)
RETURN DISTINCT reachable.fqn
def _fqn_lookup_term(fqn: str) -> str:
parts = fqn.split(".")
# Use last two segments for specificity
term = ".".join(parts[-2:]) # "PieceTree.applyDelta", not just "applyDelta"
return term
RRF Fusion — Combining Signals
For the RRF Fusion Combining Signals stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.
score(doc) = Σᵢ 1 / (k + rankᵢ(doc))
1/(60+1) + 1/(60+3) = 0.01639 + 0.01587 = 0.03226
1/(60+1) = 0.01639
Reranking — Deep Cross-Encoder Scoring
For the Reranking Deep Cross-Encoder Scoring stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.
client.rerank(query, documents, model="rerank-english-v3.0", top_n=20)
The Agentic Loop — ORA (Observe → Reason → Act)
For the The Agentic Loop ORA stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the The Agentic Loop ORA stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
INPUT: "How does YAML error handling work?"
OUTPUT: "yaml error handling failf TypeError panic propagate"
SearchPipeline — The Orchestrator
When working through the SearchPipeline The Orchestrator stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
Phase 3: Retrieval — Context Assembly and Answer Synthesis
When working through the Phase 3 Retrieval Context stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
ContextAssembler — Building the Context Window
When working through the ContextAssembler Building the Context stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours. When working through the ContextAssembler Building the Context stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
<context id="1" file="decode.go" lines="340-390"
repo="go-yaml/yaml" source="qdrant" expansion="function">
func newDecoder() *decoder {
...
}
</context>
AnswerSynthesizer — LLM Generation with Citations
The AnswerSynthesizer LLM Generation with stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
You are a code intelligence assistant. Answer questions about source code precisely.
Rules:
- Answer ONLY from the provided code context — never hallucinate code or APIs
- Cite every file you reference using [path/to/file.go:start-end] format
- Show working code examples from the context, not just descriptions
- If context is insufficient, say exactly: "Insufficient context: [what is missing]"
\[([^:\]\s][^:\]]*):(\d+)(?:-(\d+))?\]
RetrievalPipeline — The Final Orchestrator
The RetrievalPipeline The Final Orchestrator stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
retrieval = RetrievalPipeline(
search_pipeline=SearchPipeline(...),
context_assembler=ContextAssembler(search_pipeline.qdrant),
answer_synthesizer=AnswerSynthesizer(...),
)
API and CLI
The API and CLI stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
HTTP API
The HTTP API stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
{
"query": "parseTimestamp",
"repo_name": "go-yaml/yaml",
"top_k": 10,
"use_graph": false,
"use_rerank": true
}
{
"question": "how does yaml error handling work?",
"repo_name": "go-yaml/yaml",
"mode": "answer",
"max_context_chunks": 10
}
CLI
The CLI stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
# Ingest a repository
isr ingest go-yaml/yaml --with-context
# Search
isr search "what calls Unmarshal" --repo go-yaml/yaml# Retrieve with answer
isr retrieve "how does YAML error handling work?" \
--repo go-yaml/yaml --mode answer --verbose# Start the server
isr server
Configuration and Tuning
The Configuration and Tuning stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
Evaluation and Testing
The Evaluation and Testing stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
Conclusion: Why This Matters
The Conclusion Why This Matters stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve. The Conclusion Why This Matters stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Operational checklist
For the Operational checklist stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.
Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.
Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 2f4a57bf7527: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.