Home / Articles / Practical notes: Graph RAG in Action: Why Standard RAG Fails at Complex Queries

This article is published in English.

Practical notes: Graph RAG in Action: Why Standard RAG Fails at Complex Queries

Operable walkthrough of Practical notes: Graph RAG in Action: Why Standard RAG Fails at Complex Queries: contracts, checks, and drop-in code slots for teams shipping this pattern.

3051 words

Use this as an operator-facing rebuild of the ideas in “Graph RAG in Action: Why Standard RAG Fails at Complex Queries (And How Graph RAG Fixes It)”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

The real problem: disconnected fragments

For the The real problem disconnected stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

What Graph RAG does differently

For the What Graph RAG does stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

Feature            | Vector RAG          | Graph RAG
-------------------|---------------------|-------------------------
Storage unit       | Text chunks         | Entities + relationships
Retrieval method   | Semantic similarity | Graph traversal
Best for           | Direct lookup       | Multi-hop reasoning
Context scope      | Local fragment      | Connected network

The Extraction Bottleneck

For the The Extraction Bottleneck stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the The Extraction Bottleneck stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

How the Graph RAG pipeline works

When working through the How the Graph RAG stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

flowchart LR
    Q[User query] --> E[Entity extraction]
    E --> G[Graph construction]
    G --> T[Traversal + path ranking]
    T --> C[Path context]
    C --> L[LLM answer generation]
    L --> R[Final response]

The Implementation Blueprint (POC)

When working through the The Implementation Blueprint POC stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

1. Extract structured facts

When working through the 1 Extract structured facts stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the 1 Extract structured facts stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

import networkx as nx

SAMPLE_TRIPLES = [
    ("John Doe", "is CEO of", "Acme Corp"),
    ("Jane Smith", "sits on board of", "Acme Corp"),
    ("Jane Smith", "mentors", "John Doe"),
]

def build_graph(triples):
    graph = nx.DiGraph()
    for source, relation, target in triples:
        graph.add_node(source)
        graph.add_node(target)
        graph.add_edge(source, target, relation=relation)
    return graph

2. Find candidate paths

The 2 Find candidate paths stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

def traverse_graph(graph, seeds, depth=2):
    paths = []
    seen = set()
    undirected = graph.to_undirected()

    for seed in seeds:
        for target in graph.nodes:
            if seed == target:
                continue
            for path in nx.all_simple_paths(undirected, source=seed, target=target, cutoff=depth):
                canonical = tuple(path) if tuple(path) <= tuple(reversed(path)) else tuple(reversed(path))
                if canonical in seen:
                    continue
                seen.add(canonical)
                paths.append(path)
    return paths

3. Score the paths

The 3 Score the paths stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

def score_path(graph, path, query):
    query_tokens = set(tokenize(query))
    path_nodes = [node.lower() for node in path]
    path_rels = []

    for i in range(len(path) - 1):
        rel, _ = get_edge_relation(graph, path[i], path[i + 1])
        path_rels.append(rel.lower())

    path_text = " ".join(path_nodes + path_rels)

    overlap_score = len(query_tokens.intersection(set(tokenize(path_text)))) * 10
    node_score = sum(1 for node in path_nodes if any(token in node for token in query_tokens)) * 5
    rel_score = sum(1 for rel in path_rels if any(token in rel for token in query_tokens)) * 8

    length_penalty = max(0, len(path) - 2) * 2
    connection_bonus = sum(len(node.split()) for node in path_nodes)

    return overlap_score + node_score + rel_score + connection_bonus - length_penalty

4. Convert the best path into context

The 4 Convert the best stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The 4 Convert the best stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

def path_to_text(graph, path):
    lines = []
    for i in range(len(path) - 1):
        source = path[i]
        target = path[i + 1]
        relation, reversed_edge = get_edge_relation(graph, source, target)
        if reversed_edge:
            lines.append(f"{target} {relation} {source}.")
        else:
            lines.append(f"{source} {relation} {target}.")
    return " ".join(lines)

5. Ask the LLM with the chosen path

For the 5 Ask the LLM stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

def generate_llm_answer(query, context):
    prompt = (
        "You are a helpful assistant. Use only the graph facts below to answer the query clearly. "
        "Do not introduce any new information. "
        f"If the answer is not directly supported by these facts, say you don't know.\n\n"
        f"Question: {query}\n\n"
        "Graph facts:\n"
        f"{context}\n\n"
        "Answer with a short explanation of the supporting facts:"
    )

    response = client.responses.create(
        model=config["deployment_name"],
        input=prompt,
        max_output_tokens=250,
        temperature=0.1,
    )
    return response.output_text.strip()

6. Expose it as a demo API

For the 6 Expose it as stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

@app.post("/query")
def query_graph_rag(request: QueryRequest):
    query = request.query.strip()
    graph = build_graph(SAMPLE_TRIPLES)
    answer, path, context, ranked_paths = answer_query(query, graph)
    return {
        "query": query,
        "answer": answer,
        "reasoning_path": context,
        "path_nodes": path,
        "ranked_paths": ranked_paths,
    }

Bridging Prototype to Production: Scaling Up

For the Bridging Prototype to Production stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Bridging Prototype to Production stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

1. The Automated Extraction Pipeline (The Real Bottleneck)

When working through the 1 The Automated Extraction stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

2. Persistent Enterprise Graph Storage

When working through the 2 Persistent Enterprise Graph stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

3. Optimized Retrieval & Latency Management

When working through the 3 Optimized Retrieval Latency stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the 3 Optimized Retrieval Latency stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

4. Operationalization & Security

The 4 Operationalization Security stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

Summary Architecture

The Summary Architecture stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

POC System design at a glance

The POC System design at stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The POC System design at stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Results

For the Results stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

When to use vector RAG vs Graph RAG

For the When to use vector stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

Why this matters

For the Why this matters stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Why this matters stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Resources

When working through the Resources stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

What’s your take?

When working through the What s your take stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Operational checklist

When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.

Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.

Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for 8e81aec03ffd: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.