This article is published in English.
Spring AI vs LangChain4j: same RAG, different hidden choices
Twin Java RAG builds on one handbook showed line-count gaps shrink with Spring integration—and a wrong leave answer exposed chunking defaults.
The same RAG pipeline was built twice. The first scoreboard looked decisive. One experiment flipped the conclusion.
Deleting 85 lines of Java from a RAG service changed nothing users could see — and ended a two-week team argument about frameworks.
Spring AI fans and LangChain4j fans each brought slide decks. What nobody brought was a twin implementation. Both versions were written under identical controls.
Controls stayed fixed: one 240-page HR PDF, one embedding model, Postgres plus pgvector, one chat model, and a 200-question set frozen before either codebase existed.
Goal: if behavior differed, the framework should be the cause — not architecture drift between two hand-rolled designs.
Shared pipeline shape
Nothing exotic on the ask path:
Question
|
v
Embed Query
|
v
Vector Search
|
v
Top 4 Chunks
|
v
Build Context
|
v
Chat Model
|
v
Answer
Ingestion was equally boring on purpose:
PDF
↓
Extract Text
↓
Split Into Chunks
↓
Generate Embeddings
↓
Store In pgvector
The point was a framework comparison, not a novel architecture that would muddy the results.
Spring AI version
Configuration stayed tiny:
spring:
ai:
openai.api-key: ${OPENAI_KEY}
vectorstore.pgvector:
initialize-schema: true
Ingestion as a Spring component:
@Component
class Ingest {
Ingest(VectorStore store) {
var docs =
new TikaDocumentReader(
"classpath:/docs/handbook.pdf").get();
store.add(new TokenTextSplitter().apply(docs));
}
}
HTTP surface for questions:
@RestController
class AskApi {
private final ChatClient ai;
AskApi(ChatClient.Builder b, VectorStore store) {
this.ai = b.defaultAdvisors(
new QuestionAnswerAdvisor(store)).build();
}
@GetMapping("/ask")
String ask(@RequestParam String q) {
return ai.prompt().user(q).call().content();
}
}
QuestionAnswerAdvisor hid retrieval, context assembly, and prompt injection. The retrieve path barely needed custom code — which felt like a win until later.
LangChain4j version
The explicit build showed more machinery up front:
var embed =
OpenAiEmbeddingModel.builder()
.apiKey(key)
.build();
var store =
PgVectorEmbeddingStore.builder()
.host("localhost")
.port(5432)
.database("rag")
.user("app")
.password(pw)
.table("chunks")
.dimension(1536)
.build();
EmbeddingStoreIngestor.builder()
.documentSplitter(
DocumentSplitters.recursive(500, 60))
.embeddingModel(embed)
.embeddingStore(store)
.build()
.ingest(
FileSystemDocumentLoader.loadDocument(path));
Retriever construction was equally open:
var retriever =
EmbeddingStoreContentRetriever.builder()
.embeddingStore(store)
.embeddingModel(embed)
.maxResults(4)
.minScore(0.6)
.build();
Readable, yes. Compact compared with Spring AI? Not at first. Application Java LOC landed near 41 vs 126 — roughly three times more on the LangChain4j side. The debate looked settled on brevity.
Then the system answered a paternity-leave question.
The answer that was not in the handbook
Asked how many paternity days the handbook grants, the model replied confidently:
Employees are entitled to 12 days of paternity leave,
subject to the conditions listed in the policy.
The handbook says 15 days. The number 12 never appears in that policy. Retrieved context told the real story: policy tables had been shredded by the default splitter. Related rows split across chunks; the retriever returned individually plausible fragments; the model glued them into a fluent wrong answer.
The bug lived before the chat model. Framework line counts had not predicted quality.
The 3× gap was not really the framework
The line that mattered most:
new TokenTextSplitter().apply(docs)
One call smuggled decisions that had not been reviewed: table handling, chunk boundaries, overlap, surviving metadata, and what retrieval actually sees. Spring AI made those choices easy to ignore. LangChain4j made more of them visible in application code.
The first comparison also mismatched Spring Boot–integrated Spring AI with a more manual LangChain4j setup. Rebuilding LangChain4j with its Spring integration dropped the app to 58 lines. The gap shrank from about 3× to about 1.4×. Complexity had moved into the framework, not vanished from the system.
What to pick in practice
- Existing Spring Boot services → Spring AI for conventions and less ceremony on ordinary RAG.
- Standalone Java or heavy custom retrieval → LangChain4j when explicit pipeline pieces and tunable chunking/retrieval matter.
Do not choose on line count alone. After building both, the useful filter became: how many pipeline choices are you willing to leave inside the framework’s defaults? Start there — before anyone deletes eighty-five lines and declares victory.