This article is published in English.
RAG, vectorless RAG, and GraphRAG compared
Classic vector RAG, lexical vectorless retrieval, and GraphRAG: chunking, embeddings, BM25, multi-hop graphs, and when each approach earns its keep.
How today’s language models answer from material that never appeared in training.
RAG in one sentence
Retrieval Augmented Generation is exactly what the acronym spells. First you fetch material related to the user’s question. Then a model writes an answer using that material plus the prompt. The pattern is retrieval first, generation second.
Why bother?
Models only know what their training cut-off contained. Everything else is invisible.
Imagine a scarce 3,000-page book that barely appears online and never entered any training corpus. Ask a model about it and you get blanks. Hand it scrapers and you still get blanks — there is nothing discoverable to scrape.
RAG closes that gap. Load the PDF into the pipeline. At query time, pull the most relevant sections and attach them as context. The model’s job shrinks to: answer from this text, using language skill for form.
No fine-tune required. Just the right context at the right time.
Pipeline skeleton
Indexing comes first, and indexing starts with chunking.
Chunking strategies
A multi-thousand-page PDF does not belong in a vector store as one blob. Split it so each slice can be embedded and fetched on its own.
Common splits:
Per page. One page → one chunk (3,000 pages → 3,000 chunks). Strong when you want tight, precise hits.
Per paragraph. Finer cuts on paragraph boundaries. Can feel snappier, but sizes swing wildly (200 tokens next to 2,000). Uneven lengths make uneven embeddings and quietly damage retrieval.
Fixed windows. Constant token budgets — often 512 — wherever the cut falls. Uniform sizes yield more comparable embeddings. When unsure, most teams default here.
from langchain_text_splitters import CharacterTextSplitter
from langchain_core.documents import Document
def perform_fixed_size_chunking(document, chunk_size=1000, chunk_overlap=200😞
"""
Performs fixed-size chunking on a document with specified overlap.
Args:
document (str): The text document to process
chunk_size (int): The target size of each chunk in characters
chunk_overlap (int): The number of characters of overlap between chunks
Returns:
list: The chunked documents with metadata
"""
# Create the text splitter with optimal parameters
text_splitter = CharacterTextSplitter(
separator="\n\n",
chunk_size=chunk_size,
chunk_overlap=chunk_overlap,
length_function=len
)
# Split the text into chunks
chunks = text_splitter.split_text(document)
print(f"Document split into {len(chunks)} chunks")
# Convert to Document objects with metadata
documents = []
for i, chunk in enumerate(chunks):
doc = Document(
page_content=chunk,
metadata={
"chunk_id": i,
"total_chunks": len(chunks),
"chunk_size": len(chunk),
"chunk_type": "fixed-size"
}
)
documents.append(doc)
return documents
# Example usage
if __name__ == "__main__":
# Create the dummy document
document = create_dummy_document()
# Process with fixed-size chunking
chunked_docs = perform_fixed_size_chunking(
document,
chunk_size=1000,
chunk_overlap=200
)
# Display results
print("\n----- CHUNKING RESULTS -----")
print(f"Total chunks: {len(chunked_docs)}")
# Print an example chunk
print("\n----- EXAMPLE CHUNK -----")
middle_chunk_idx = len(chunked_docs) // 2
example_chunk = chunked_docs[middle_chunk_idx]
print(f"Chunk {middle_chunk_idx} content ({len(example_chunk.page_content)} characters):")
print("-" * 40)
print(example_chunk.page_content)
print("-" * 40)
print(f"Metadata: {example_chunk.metadata}")
# For integration with Databricks Vector Search
print("\nThese documents are ready for embedding and storage in Databricks Vector Search")
print("Example next steps:")
print("1. Create embeddings using the Databricks embedding endpoint")
print("2. Store documents and embeddings in Delta table")
print("3. Create Vector Search index for retrieval")
Embeddings
An embedding maps text (word, sentence, page) into a dense high-dimensional vector that encodes meaning so similar ideas cluster.
Four phases inside RAG
- Corpus side. Embed each chunk; store index + vector.
- Query side. Embed the user prompt so the request has a semantic fingerprint.
- Lookup. Search the store for nearest neighbors — typically cosine similarity.
- Augment. Attach retrieved passages as extra context; the model answers from prompt + context.
Where vectors live
This is a specialized store for embeddings, not a general relational or object database. AlloyDB, Pinecone, and Qdrant are common; many teams use pgvector on PostgreSQL.
Where vanilla vector RAG struggles
- Cuts can ignore meaning. Related windows may separate; larger overlap helps sometimes but is not a cure-all.
- Similarity can miss paraphrases — “sales decreased” versus “company in decline” may not land nearby in vector space.
- Multi-hop facts break when cause and effect sit in different chunks and only one is retrieved.
- Building, storing, indexing, and reindexing embeddings is costly.
Also: a vector database is not synonymous with RAG. It is one retrieval backend. Vectorless RAG keeps “fetch then generate” while dropping embedding search.
Why skip vectors? Embedding build/reindex cost, weak exact-match behavior for IDs/numbers/error codes, and extra infrastructure to operate.
Vectorless is a family, not one recipe:
- Lexical search. BM25, Postgres
tsvector, Elasticsearch — exact terms beat fuzzy semantics for SKUs, citations, and log lines.
from rank_bm25 import BM25Okapi
def vectorless_retrieve(query, corpus_chunks, top_k=3):
"""
Lexical retrieval over raw text chunks - no embeddings, no vector DB.
"""
tokenized_corpus = [chunk.lower().split() for chunk in corpus_chunks]
bm25 = BM25Okapi(tokenized_corpus)
tokenized_query = query.lower().split()
scores = bm25.get_scores(tokenized_query)
ranked = sorted(zip(corpus_chunks, scores), key=lambda x: x[1], reverse=True)
return [chunk for chunk, score in ranked[:top_k]]
- Agentic / tool retrieval. No pre-index; the model greps, calls search APIs, or opens sections on demand the way a coding agent walks a repo. Live, reasoning-driven lookup.
- Long-context stuffing. With huge context windows, small corpora can ride along in the prompt. Not classic retrieval, but similar outcomes for modest corpora.
- Hybrid re-rank. Cheap lexical shortlist, then a model reorders for relevance — keyword speed with some semantic nuance, without a full embedding index upfront.
Vectorless limits
Keywords still miss paraphrases — sometimes worse than embeddings. Agentic loops add latency and tokens per query. Huge corpora still favor a well-built vector index. Vectorless tends to win at small/medium scale or when exactness beats fuzziness.
GraphRAG
Classic RAG can tear cause from effect across chunks. Vectorless swaps embeddings for keywords or long context. Neither models how ideas relate. Graph RAG targets that gap.
Instead of “which chunk is nearest?”, ask “how do these concepts connect?” The answer is a knowledge graph.
Indexing — grow the graph
Skip chunk/embed first. Run the document through a model that extracts entities and relationships. Output: nodes and edges.
On that long theory book you might get nodes such as Marx, Engels, Capital, The Communist Manifesto, surplus value, dialectical materialism, and edges such as authored, co-wrote, introduces-concept.
Entities become nodes; relations become edges. You are mapping meaning, not slicing pages.
Cluster tightly linked nodes into communities (Leiden is popular), then summarize each community with another model pass.
Operate on two levels:
- Nodes for specific facts and direct links
- Communities for themes and summaries
Narrow questions → nodes. Broad “core ideas?” questions → community summaries. One index, two retrieval modes.
Querying — walk the graph
Question: “Who co-wrote with Marx, and what did they write together?”
Vanilla RAG struggles: co-authorship in one chunk, works hundreds of pages away, hoping embeddings agree.
Graph RAG walks:
- Detect Marx as the anchor
- Load Marx’s node and edges
- Follow co-wrote → Engels
- Follow Engels authored → Manifesto, Condition of the Working Class, related works
Typed directional edges keep relationships intact. Cause and effect that classic RAG split become linked nodes. That multi-hop pattern is hard for both basic and vectorless RAG to do cleanly.
Assemble nodes, edges, and community summaries into context; generate as usual.
from graphrag import GraphRAGPipeline
# Indexing — runs once
pipeline = GraphRAGPipeline(llm="claude-3", graph_store="neo4j")
pipeline.index(documents=["book.pdf"])
# Under the hood: entity extraction → graph build → community detection → summaries
# Querying
result = pipeline.query(
"Who co-wrote with Marx and what did they write together?",
mode="global" # uses community summaries for broad questions
# mode="local" # uses node-level traversal for specific facts
)
print(result.answer)
print(result.sources) # returns actual nodes + edges used, fully traceable
mode is not cosmetic. Implementations (including Microsoft’s open-source stack) expose global vs local because they are different strategies.
GraphRAG costs
Extraction is expensive. The whole document must be read; long books burn tokens; implicit references (“as discussed earlier”) may never become edges.
Graph quality caps quality. Bad extraction → bad graph → bad retrieval. Fixes often mean re-reading everything, which hurts at scale.
In practice, pick the retriever by query type rather than forcing one path on every question.
Summary: use classic RAG for semantic lookup, vectorless when you need exactness or less infra, and Graph RAG when relationships are the product.
Choosing a retrieval style without dogma
A useful decision sequence looks like this.
Start with classic vector RAG when your corpus is large, language varies a lot, and approximate semantic neighbors are usually what you want. Invest in chunking quality and embedding refresh processes up front; those are the levers that dominate quality.
Reach for vectorless techniques when exact identifiers matter more than paraphrase, when the corpus is small enough for lexical search or long context, or when you refuse the cost of embedding infrastructure. BM25 and friends are not “old-fashioned” — they are the right tool for SKUs, citations, and error strings.
Reach for Graph RAG when the product question is relational: who connected to whom, which concept introduced which idea, which community summarizes a theme. Expect higher indexing cost and treat extraction quality as a first-class dependency.
Many teams end up hybrid: lexical first pass, vector second pass, graph for a subset of multi-hop domains. The point is not to pick a tribe. The point is to match the retriever to the failure mode you actually see.
Operational notes that demos skip
Reindexing schedules matter. Stale embeddings quietly degrade vector RAG even when the model is unchanged. Overlap settings matter when adjacent windows share meaning. Metadata filters matter when tenants must never see each other’s chunks. Evaluation sets matter when you claim “better” without labels.
For Graph RAG, plan for extraction retries, partial graph updates, and community re-summarization when documents change. For agentic retrieval, plan for tool budgets and timeout policies so a curious model cannot burn your token wallet on a single query.
None of these notes are glamorous. They are the difference between a diagram and a system that survives a month of real traffic.