实用笔记:掌握 Neo4j 与 LangChain4j:GraphRAG、持久化 AI 内存
《实用笔记》操作指南:掌握 Neo4j 与 LangChain4j——GraphRAG、持久化 AI 内存,以及适用于采用该架构的团队的合约、校验机制与即插即用代码模块。
以下笔记为“掌握 Neo4j 和 LangChain4j:GraphRAG、持久化 AI 内存及更多功能”提供了一条实用的学习路径。重点在于契约定义、校验机制以及可直接插入的代码占位符,而非激励性表述。 在完成概览阶段时,首先明确契约内容:所需输入、成功标志以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 建议采用小型、可测试的单元而非庞大的脚本。当某一步骤失败时,故障应能指向单一责任点,而非复杂的流程链。
<dependencies>
<dependency>
<groupId>dev.langchain4j</groupId>
<artifactId>langchain4j-community-neo4j-retriever</artifactId>
<version>${langchain.version}</version>
</dependency>
<dependency>
<groupId>dev.langchain4j</groupId>
<artifactId>langchain4j</artifactId>
<version>${langchain.version}</version>
</dependency>
<dependency>
<groupId>dev.langchain4j</groupId>
<artifactId>langchain4j-community-neo4j</artifactId>
<version>${langchain.version}</version>
</dependency>
<!-- other deps -->
</dependencies>
动态模式抽象:Neo4jGraph
将动态模式抽象 Neo4jGraph 阶段视为可度量的处理环节时,其效果最佳。在扩大范围之前,先记录一份理想的输出结果、一个失败案例以及回滚说明。 把这一阶段视为输入与经过验证的输出之间的契约。为相关成果命名,明确成功标准,绝不允许出现无声无息的半完成状态。 将分块策略与检索策略分开。当质量指标发生变化时,修改其中一项不应强制要求重新编写另一项。
单查询模式检索
将单次查询模式检索阶段视为可度量的对象来处理,其效果最佳。在扩大范围之前,先记录一份理想的查询结果、一个失败案例以及回滚说明。在功能结果旁同时记录处理时间以及令牌或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外费用。应将分块策略与检索策略分开,当质量指标发生变化时,修改其中一项无需强制重写另一项。
import dev.langchain4j.store.graph.neo4j.Neo4jGraph;
// If I want to initialize the graph abstraction and load the schema:
Neo4jGraph graph = Neo4jGraph.builder()
.driver(driver)
.build();
// Under the hood, this executes a single, consolidated APOC query that yields
// labels, element types, and properties all at once, avoiding multiple DB calls.
graph.refreshSchema();
Neo4jGraph.StructuredSchema schema = graph.getStructuredSchema();
// We expect a well-formatted string logically divided into three sections,
// exactly as formatted by the new Neo4jGraphSchemaUtils class:
//
// `schema.nodesProperties()` is the following:
// :Person {name: STRING}, :Company {name: STRING}
//
// `schema.relationshipsProperties()` is the following:
// :WORKS_FOR {since: INTEGER}
//
// `schema.patterns()` is the following:
// (:Person)-[:WORKS_FOR]->(:Company)
System.out.println("Current Database Schema Context:\n" + schema);
配置模式采样与增强功能
将“配置模式采样”与相关阶段视为可度量的对象来处理时,其效果最佳。在扩大范围之前,需记录一份理想状态下的转录内容、一个故障案例以及回滚说明。 应将配置信息与应用程序代码分开。环境文件、密钥存储和功能标志应集中存放于一处,以便操作人员无需查看整个系统结构即可进行审计。 需将分块策略与检索策略区分开来。当质量指标发生变化时,修改其中一项不应迫使重新编写另一项。 将“配置模式采样”与相关阶段视为可度量的对象来处理时,其效果最佳。在扩大范围之前,需记录一份理想状态下的转录内容、一个故障案例以及回滚说明。 相较于庞大的脚本,应优先使用小型且可测试的单元。当某个步骤出现故障时,故障原因应能明确指向单一责任模块,而非复杂的流程链。
// If I want to scan a very large database efficiently by configuring the APOC sampling,
// and I want to enhance the LLM's understanding with sample property values:
Neo4jGraph optimizedGraph = Neo4jGraph.builder()
.driver(driver)
// We can configure the underlying apoc.meta.data parameters.
// 'sample' limits the number of nodes inspected per label to speed up execution.
// 'maxRels' limits the number of relationships inspected per node.
// (Note: These are passed internally to the getSchemaFromMetadata utility)
.build();
// When the schema is refreshed, the underlying query runs:
// CALL apoc.meta.data({maxRels: $maxRels, sample: $sample})
optimizedGraph.refreshSchema();
// The LLM now receives a fast, accurately sampled schema representation,
// protecting database performance during application startup or schema refreshes.
System.out.println("Optimized Schema loaded successfully.");
自动化知识图谱构建与源链接关联
在自动化知识图谱构建阶段,修改代码之前需明确输入内容、该步骤的负责人以及完成标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 应将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,定义成功判定标准,杜绝无声的半完成状态。 必须注明支撑答案的具体依据。若没有引用,操作人员就无法区分幻觉内容与索引缺失问题。
初始数据集与少样本提示
在处理初始数据集及相应阶段时,应在修改代码之前明确输入内容、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。除了功能结果外,还需记录执行时间以及令牌或查询成本。提前了解这些成本信息,可避免在从演示环境过渡到共享环境时出现意外费用。当下一步操作为编写代码或调用工具时,应优先选择具有架构验证的结构化输出,而非自由形式的文本。
[
{
"tail": "Microsoft",
"head": "Adam",
"head_type": "Person",
"text": "Adam is a software engineer in Microsoft since 2009...",
"relation": "WORKS_FOR",
"tail_type": "Company"
},
{
"tail": "Microsoft Word",
"head": "Microsoft",
"head_type": "Company",
"text": "Microsoft is a tech company that provides several products...",
"relation": "PRODUCED_BY",
"tail_type": "Product"
}
]
import dev.langchain4j.community.data.document.transformer.graph.LLMGraphTransformer;
import dev.langchain4j.community.data.document.graph.GraphDocument;
import dev.langchain4j.data.document.Document;
import dev.langchain4j.data.document.DefaultDocument;
import dev.langchain4j.data.document.Metadata;
import java.util.List;
ChatModel chatModel = /* dev.langchain4j.model.chat instance */
Driver driver = /* org.neo4j.driver.Driver instance */
// If I want to guide the extraction by providing a structured set of examples:
LLMGraphTransformer transformer = LLMGraphTransformer.builder()
.model(chatModel)
.examples(EXAMPLES_PROMPT) // Injects the above JSON dataset into the system prompt
.build();
Document docKeanu = new DefaultDocument(
"Keanu Reeves acted in Matrix",
Metadata.from("key33", "value3")
);
// The LLM will transform the text, structuring nodes and relationships based on the examples
List<GraphDocument> graphDocs = transformer.transformAll(List.of(docKeanu));
/*
The above `graphDocs` returns this result:
GraphDocument
├─ Nodes
│ ├─ GraphNode
│ │ ├─ id: Matrix
│ │ ├─ type: Movie
│ │ └─ properties: {}
│ │
│ └─ GraphNode
│ ├─ id: Keanu Reeves
│ ├─ type: Person
│ └─ properties: {}
│
├─ Relationships
│ └─ GraphEdge
│ ├─ type: ACTED_IN
│ ├─ sourceNode
│ │ ├─ id: Keanu Reeves
│ │ └─ type: Person
│ ├─ targetNode
│ │ ├─ id: Matrix
│ │ └─ type: Movie
│ └─ properties: {}
│
└─ Source
├─ text: "Keanu Reeves acted in Matrix"
└─ metadata
└─ key33: value3
*/
可重试的图数据持久化
在幂等图持久化阶段,修改代码之前需明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 配置信息应置于应用程序代码之外。环境文件、密钥存储和功能标志应集中存放于一个位置,以便操作人员无需查看整个系统结构即可进行审计。 需引用实际作为答案依据的段落。没有引用的话,操作人员就无法区分是虚假信息还是索引缺失导致的错误。 在幂等图持久化阶段,修改代码之前需明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 相较于庞大的脚本,应优先使用小型且可测试的单元。当某个步骤失败时,故障原因应能指向单一责任模块,而非多个部分共同导致的问题。
带角度的管道。import dev.langchain4j.community.rag.content.retriever.neo4j.KnowledgeGraphWriter;
Neo4jGraph neo4jGraph = /* dev.langchain4j.store.graph.neo4j.Neo4jGraph instance */;
LLMGraphTransformer graphTransformer = /* dev.langchain4j.community.data.document.transformer.graph.LLMGraphTransformer instance */;
Document docKeanu = new DefaultDocument(
"Keanu Reeves acted in Matrix",
Metadata.from("key33", "value3")
);
List<GraphDocument> graphDocs = graphTransformer.transformAll(List.of(docKeanu));
// If I want to persist the extracted entities safely:
KnowledgeGraphWriter writer = KnowledgeGraphWriter.builder()
.graph(neo4jGraph)
.build();
// The first write populates the database
writer.addGraphDocuments(graphDocs, false);
// Executing the exact same command again is safe.
// The internal UNWIND and MERGE logic generated by the writer guarantees
// that no duplicate entities or relationships are created.
writer.addGraphDocuments(graphDocs, false);
// Expected resulting topology in the database:
// (:__Entity__ {id: 'keanu'})-[:ACTED]->(:__Entity__ {id: 'matrix'})
System.out.println("Entities persisted successfully. No duplicates created.");
源文档链接(includeSource)
在处理源文档链接的includeSource阶段时,首先需明确相关约定:所需的输入参数、成功信号以及部分失败时的处理方式。这样的清单能确保后续的代码修改始终符合约定。 将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,定义成功判定标准,杜绝无声的半完成状态。 在调整提示词之前,先使用固定的问题集来测试召回率。仅仅更换提示词往往无法改善较差的检索效果。
// If I want to maintain data provenance and link entities back to their source:
writer.addGraphDocuments(graphDocs, true); // true = includeSource
// Behind the scenes, the writer executes three crucial operations:
// 1. It creates the Document node, copying the original metadata (e.g., key33: value3) and the text.
// 2. If the original document lacks an ID, the writer automatically generates
// an MD5 hash of the text to use as a unique identifier.
// 3. It links the document to the extracted entities using a relationship (default: HAS_ENTITY).
// Expected resulting topology in the database:
// (:Document {id: '<MD5_hash>', text: 'Keanu Reeves...', key33: 'value3'})-[:HAS_ENTITY]->(:__Entity__ {id: 'keanu'})
System.out.println("Source document successfully linked to the extracted entities.");
深度模式定制与安全性
在处理深度模式定制阶段时,首先需明确相关约定:所需的输入参数、成功标识以及部分失败时的处理方式。这样的清单能确保后续的代码修改始终符合预期。 在功能结果旁记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外费用。 在调整提示词之前,先使用固定的问题集测试召回率。仅仅更换提示词往往无法改善较差的检索效果。
// If I want to adapt the ingestion to a pre-existing enterprise schema,
// and customize the relationship that links the document to the entities:
KnowledgeGraphWriter customWriter = KnowledgeGraphWriter.builder()
.graph(neo4jGraph)
.label("ActorOrMovie") // Replaces "__Entity__"
.idProperty("customId") // Replaces "id"
.textProperty("customText") // Replaces "text" for the Document node
.relType("MENTIONED_IN_SOURCE") // Replaces "HAS_ENTITY"
.constraintName("unique_custom") // Sets a specific name for the Neo4j CONSTRAINT
.build();
// Inserting the data with includeSource set to true will now use the new nomenclature:
customWriter.addGraphDocuments(graphDocs, true);
// Example of the resulting Cypher pattern generated by the writer:
// (:Document {customText: '...'})-[:MENTIONED_IN_SOURCE]->(:ActorOrMovie {customId: 'keanu'})
通过 Cypher DSL 集成实现安全保障
在处理“通过 Cypher DSL 实现安全保障”这一阶段时,首先需明确相关规范:所需的输入参数、成功标识,以及部分失败时的处理方式。这样的清单能确保后续的代码修改始终符合既定要求。 应将配置信息与应用程序代码分开存放。环境文件、密钥存储以及功能开关应集中管理,以便操作人员无需查看整个系统结构即可进行审计。 在调整提示词之前,先使用固定的问题集来测试召回率。仅仅更换提示词往往无法解决检索效果不佳的问题。 在处理“通过 Cypher DSL 实现安全保障”这一阶段时,首先需明确相关规范:所需的输入参数、成功标识,以及部分失败时的处理方式。这样的清单能确保后续的代码修改始终符合既定要求。 相比庞大的脚本,更应优先使用小型且可测试的单元。当某个步骤出现故障时,故障点应能明确指向某一个具体的功能模块,而非整个复杂的流程。
import org.neo4j.cypherdsl.core.Cypher;
import org.neo4j.cypherdsl.core.Statement;
import org.neo4j.cypherdsl.core.Node;
Driver driver = /* org.neo4j.driver.Driver instance */
// Demonstrating how internal queries are constructed safely via the DSL
// We define a Node representation first
Node documentNode = Cypher.node("Document").named("d");
// We build the query programmatically using the fluent API.
// Notice how literal values are wrapped safely, preventing injection.
Statement statement = Cypher.match(documentNode)
.where(Cypher.property("d", "id").isEqualTo(Cypher.literalOf("doc-123")))
.returning(Cypher.property("d", "text"))
.build();
// The DSL engine traverses the AST and compiles it into a syntactically safe string.
String safeCypherQuery = statement.getCypher();
// Expected output: MATCH (d:`Document`) WHERE d.id = 'doc-123' RETURN d.text
System.out.println("Generated safe Cypher via DSL: " + safeCypherQuery);
使用2026.01语法进行索引内预过滤
将“使用2026版本进行索引内预过滤”这一阶段视为可度量的流程最为有效。在扩大范围之前,需收集一份理想的处理结果、一个失败案例以及回滚说明。 应将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,明确成功标准,杜绝默许的半完成状态。 需将分块策略与检索策略分开处理。当质量指标发生变化时,修改其中一项不应强制要求重新编写另一项。
import dev.langchain4j.store.embedding.neo4j.Neo4jEmbeddingStore;
import dev.langchain4j.store.embedding.filter.Filter;
import static dev.langchain4j.store.embedding.filter.MetadataFilterBuilder.metadataKey;
Driver driver = /* org.neo4j.driver.Driver instance */
Embedding embedding = /* dev.langchain4j.data.embedding.Embedding instance */
MatchSearchClauseStrategy matchSearchClauseStrategy = new MatchSearchClauseStrategy();
// Configure the store with the new syntax enabled
Neo4jEmbeddingStore store = Neo4jEmbeddingStore.builder()
.driver(driver)
.dimension(1536)
.searchStrategy(matchSearchClauseStrategy) // Crucial flag: Enables the 2026.01 optimized native vector search syntax
.filterMetadata(Arrays.asList("year", "department")) // Enable filtering for 'year' and 'department', which translates to the `WITH [indexName.year, indexName.department]` clause during index creation
.build();
// Build a metadata filter combining multiple boolean conditions
Filter filter = metadataKey("year").isEqualTo(2024)
.and(metadataKey("department").isEqualTo("Engineering"));
// Execute the search request
EmbeddingSearchRequest request = EmbeddingSearchRequest.builder()
.queryEmbedding(embedding)
.maxResults(5)
.filter(filter) // Filter is executed natively inside the Neo4j Vector Index block
.build();
SearchResult<TextSegment> results = store.search(request);
// The results will natively exclude any documents not matching the criteria specified in the `filter` instance
// returning the final Top-K, ensuring you always get 5 highly relevant segments if they exist.
System.out.println("Search executed with in-index filtering.");
// Execute results.matches() to verify in-index filtering
GraphRAG检索概念:高级结构化检索器与数据导入工具
将GraphRAG检索概念的进阶阶段视为可度量的对象来处理,效果最佳。在扩大应用范围之前,先记录一个成功的案例、一个失败案例以及回滚说明。在功能结果旁同时记录处理时间以及令牌或查询成本。提前了解成本情况,就能避免在从演示环境过渡到共享环境时出现意外费用。应将分块策略与检索策略分开,当质量指标发生变化时,修改其中一项无需强制重写另一项。
父子模式
将“父子模式”阶段视为可度量的对象来处理时,其效果最佳。在扩大范围之前,需记录一份理想案例、一个故障实例以及回滚说明。 应将配置与应用程序代码分开。环境文件、密钥存储和功能标志应集中存放于一处,以便操作人员无需查看整个架构即可进行审计。 应将分块策略与检索策略分开。当质量指标发生变化时,修改其中一项不应迫使重新编写另一项。 将“父子模式”阶段视为可度量的对象来处理时,其效果最佳。在扩大范围之前,需记录一份理想案例、一个故障实例以及回滚说明。 相比庞大的脚本,更应采用小型且可测试的单元。当某个步骤出现故障时,故障应指向单一责任模块,而非复杂的流程链。
import dev.langchain4j.store.graph.neo4j.Neo4jParentChildIngestor;
import dev.langchain4j.data.document.Document;
import dev.langchain4j.data.document.splitter.DocumentSplitters;
// embeddingModel, embeddingStore, childSplitter instances...
// If I want to automatically ingest a document into a Parent-Child graph topology:
Neo4jEmbeddingStoreIngestor ingestor = ParentChildGraphIngestor.builder()
.driver(driver)
.embeddingModel(embeddingModel)
// We define how the document should be chunked before ingestion
.documentSplitter(DocumentSplitters.recursive(200, 20))
.documentChildSplitter(childSplitter)
.build();
Document document = Document.from( """Artificial Intelligence (AI) is a field of computer science. It focuses on creating intelligent agents capable of performing tasks that require human intelligence.
Machine Learning (ML) is a subset of AI. It uses data to learn patterns and make predictions. Deep Learning is a specialized form of ML based on neural networks.
""");
ingestor.ingest(List.of(document));
// The graph now contains one 'Document' node connected via 'HAS_CHILD'
// to multiple embedded 'DocumentChunk' nodes.
System.out.println("Parent and child nodes successfully ingested and linked.");
import dev.langchain4j.rag.content.retriever.EmbeddingStoreContentRetriever;
// If I want to search against chunks but retrieve the rich parent document:
final EmbeddingStoreContentRetriever retriever = EmbeddingStoreContentRetriever.builder()
.embeddingStore(embeddingStore)
.maxResults(1)
.minScore(0.4)
.build();
List<Content> contents = retriever.retrieve(Query.from("specific configuration detail"));
// The retriever hits the small 'DocumentChunk' index, traverses the 'HAS_CHILD'
// relationship, and returns the entire 'Document' node.
System.out.println("Retrieved full parent document context.");
/*
The result of `contents` is :
DefaultContent {
textSegment = TextSegment {
text = """
Machine Learning (ML) is a subset of AI. It uses data to learn patterns and make predictions.
Deep Learning is a specialized form of ML based on neural networks.
Machine Learning (ML) is a subset of AI. It uses data to learn patterns and make predictions.
Deep Learning is a specialized form of ML based on neural networks.
Artificial Intelligence (AI) is a field of computer science. It focuses on creating intelligent agents
capable of performing tasks that require human intelligence.
""",
metadata = {
index = 1,
source = Wikipedia link,
title = AI Basics,
url = https://example.com/ai,
parentId = parent_1af3e080-5029-40ab-b3d7-0829064d800d
}
},
metadata = {
EMBEDDING_ID = null,
SCORE = 0.8560410737991333
}
}
*/
总结模式
在“总结模式”阶段,应在修改代码之前明确输入内容、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,定义成功判定标准,并拒绝默许的半完成状态。 需引用实际作为答案依据的段落。没有引用的话,操作人员就无法区分幻觉内容与索引缺失的情况。
import dev.langchain4j.community.store.embedding.neo4j.Neo4jEmbeddingStoreIngestor;
import dev.langchain4j.community.store.embedding.neo4j.SummaryGraphIngestor;
// If I want the LLM to summarize my document during ingestion and store the summary:
/* Neo4jEmbeddingStore, ChatModel and DocumentSplitter instances... */
final Neo4jEmbeddingStoreIngestor ingestor = SummaryGraphIngestor.builder()
.driver(driver)
.embeddingModel(embeddingModel)
.questionModel(chatModel)
.documentSplitter(parentSplitter)
.build();
ingestor.ingest(List.of(document));
// The graph now has a 'Summary' node (containing the LLM-generated summary)
// linked to the specific 'DocumentChunk' nodes.
System.out.println("Document ingested and summarized successfully.");
import dev.langchain4j.community.store.embedding.neo4j.Neo4jEmbeddingStoreIngestor;
import dev.langchain4j.community.store.embedding.neo4j.SummaryGraphIngestor;
/* required instances */
Document document = Document.from("""
Artificial Intelligence (AI) is a field of computer science. It focuses on creating intelligent agents capable of performing tasks that require human intelligence.
Machine Learning (ML) is a subset of AI. It uses data to learn patterns and make predictions. Deep Learning is a specialized form of ML based on neural networks.
""");
ingestor.ingest(document);
final EmbeddingStoreContentRetriever retriever = EmbeddingStoreContentRetriever.builder()
.embeddingModel(embeddingModel)
.maxResults(5)
.minScore(0.6)
.embeddingStore(ingestor.getEmbeddingStore())
.build();
/*
The result is something like this:
DefaultContent {
textSegment = TextSegment {
text = "Machine Learning (ML) is a subset of AI",
metadata = {
index = 0,
source = Wikipedia link,
title = Quantum Mechanics,
url = https://example.com/ai
}
},
metadata = {
EMBEDDING_ID = null,
SCORE = 0.8425111770629883
}
}
*/
假设性问题模式
在“假设性问题模式”阶段,应在修改代码之前明确输入内容、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。应在功能结果旁记录执行时间以及令牌或查询成本。提前显示成本信息,可避免在从演示环境过渡到共享环境时出现意外费用。需注明实际作为答案依据的段落;没有引用的话,操作人员就无法区分是幻觉内容还是索引缺失导致的错误。
import dev.langchain4j.community.store.embedding.neo4j.Neo4jEmbeddingStoreIngestor;
import dev.langchain4j.community.store.embedding.neo4j.HypotheticalQuestionGraphIngestor;
/* ... Neo4jEmbeddingStore, ChatModel and DocumentSplitter instances.. */
Neo4jEmbeddingStoreIngestor ingestor = HypotheticalQuestionGraphIngestor.builder()
.embeddingModel(embeddingModel)
.driver(driver)
.documentSplitter(splitter)
.questionModel(chatModel)
.embeddingStore(embeddingStore)
.build();
Document document = Document.from("""
Quantum mechanics studies how particles behave. It is a fundamental theory in physics.
Gradient descent and backpropagation algorithms.
Spaghetti carbonara and Italian dishes.
John Doe is a Super Saiyan.
""");
ingestor.ingest(document);
EmbeddingStoreContentRetriever retriever = EmbeddingStoreContentRetriever.builder()
.embeddingModel(embeddingModel)
.maxResults(2)
.minScore(0.5)
.embeddingStore(ingestor.getEmbeddingStore())
.build();
List<Content> results = retriever.retrieve(Query.from("Who is John Doe?"));
System.out.println("Retrieved Hypothetical context: " + results);
/*
The result is something like this:
DefaultContent {
textSegment = TextSegment {
text = "John Doe is a Super Saiyan.",
metadata = {
index = 2,
source = Wikipedia link,
title = Quantum Mechanics,
url = https://example.com/ai
}
},
metadata = {
SCORE = 0.8234479427337646,
EMBEDDING_ID = 2002cabe-2a3e-4c6e-96ec-a0292e26e817
}
}
*/
通用模式
在“通用模式”阶段,应在修改代码之前明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测其中的隐藏状态。 配置信息应置于应用程序代码之外。环境文件、密钥存储以及功能标志应集中存放于一个位置,以便操作人员无需查看整个流程即可进行审计。 需引用实际作为答案依据的段落。如果没有引用,操作人员就无法区分是虚假信息还是索引缺失导致的错误。 在“通用模式”阶段,应在修改代码之前明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测其中的隐藏状态。 相较于庞大的脚本,更应优先使用小型且可测试的单元。当某个步骤出现故障时,故障原因应能指向单一责任模块,而非复杂的流程链。
final Neo4jEmbeddingStore neo4jEmbeddingStore = /* Neo4jEmbeddingStore instance */
final EmbeddingStoreContentRetriever retriever = EmbeddingStoreContentRetriever.builder()
.embeddingModel(embeddingModel)
.maxResults(5)
.minScore(0.4)
.embeddingStore(neo4jEmbeddingStore)
.build();
// other required instances ...
Document doc = Document.from("""
Quantum mechanics studies how particles behave. It is a fundamental theory in physics.
Gradient descent and backpropagation algorithms.
Spaghetti carbonara and Italian dishes.
John Doe is a Super Saiyan.
""");
// Ingest the document into Neo4j as parent-child nodes
final Neo4jEmbeddingStoreIngestor ingestor = Neo4jEmbeddingStoreIngestor.builder()
.documentSplitter(parentSplitter)
.documentChildSplitter(childSplitter)
.driver(driver)
.query("CREATE (:MainDoc $metadata)") // a Cypher query template used for storing the processed segment data in Neo4j
.embeddingStore(neo4jEmbeddingStore)
.embeddingModel(embeddingModel)
.build();
ingestor.ingest(doc);
final String retrieveQuery = "Machine Learning";
List<Content> results = retriever.retrieve(Query.from(retrieveQuery));
System.out.println("Retrieved Generic context: " + results);
与数据库类型无关的父子关系检索工具
在处理“与数据库类型无关的父子关系”这一阶段时,首先需明确相关约定:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 将这一阶段视为输入与经过验证的输出之间的契约。为相关成果命名,定义成功判定标准,杜绝无声的半完成状态。 在调整提示词之前,先使用固定的问题集来衡量检索的召回率。仅仅更换提示词很难改善较差的检索效果。
import dev.langchain4j.community.store.embedding.ParentChildEmbeddingStoreIngestor;
/* Required instances */
ParentChildEmbeddingStoreIngestor ingestor = ParentChildEmbeddingStoreIngestor.builder()
.documentTransformer(documentTransformer)
.documentSplitter(documentSplitter)
.textSegmentTransformer(textSegmentTransformer)
.embeddingModel(embeddingModel)
.embeddingStore(embeddingStore)
.documentChildSplitter(documentChildSplitter)
.childTextSegmentTransformer(childTextSegmentTransformer)
.build();
有状态人工智能:图结构中的持久对话记忆
在处理有状态AI持久对话阶段时,首先需列出相关规范:所需输入、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 在功能结果旁记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外费用。 在调整提示词之前,先使用固定的问题集测试召回率。仅仅更换提示词往往无法改善较差的检索效果。
配置
在配置阶段工作时,首先需明确相关规范:所需的输入参数、成功信号以及部分失败时的处理方式。这份清单能确保后续的代码修改始终符合预期。 应将配置信息与应用程序代码分开存放。环境文件、密钥存储以及功能开关应集中于一个位置,这样操作人员无需查看整个系统结构即可进行审计。 在调整提示词之前,先使用固定的问题集来测试系统的召回率。仅仅更换提示词往往无法解决检索效果不佳的问题。 在配置阶段工作时,首先需明确相关规范:所需的输入参数、成功信号以及部分失败时的处理方式。这份清单能确保后续的代码修改始终符合预期。 相比庞大的脚本,更应优先使用小型且可测试的单元。当某个步骤出现故障时,故障点应能明确指向某个具体的功能模块,而非整个复杂的处理流程。
管理多租户聊天记录与多模态输入
将“管理多租户聊天记录”这一阶段视为可度量的工作对象会更为有效。在扩大范围之前,先记录一份最佳处理案例、一个失败案例以及回滚说明。 应将此阶段视为输入与经过验证的输出之间的契约。为相关文档命名,明确成功标准,杜绝默许的半完成状态。 需将分块策略与检索策略分开处理。当质量指标发生变化时,调整其中一项不应强制要求重新编写另一项。
import dev.langchain4j.store.memory.chat.neo4j.Neo4jChatMemoryStore;
import dev.langchain4j.data.message.UserMessage;
import dev.langchain4j.data.message.AiMessage;
import dev.langchain4j.data.message.ImageContent;
import java.util.List;
// 1. Initialize the store using an existing driver
Neo4jChatMemoryStore memoryStore = Neo4jChatMemoryStore.builder()
.driver(driver)
.build();
// 2. Identify the specific user sessions
String sessionId1 = "user-alice-123";
String sessionId2 = "user-bob-456";
// 3. Append standard text messages for Alice
List<ChatMessage> aliceMessages = List.of(
new UserMessage("Hi, I'm Alice."),
new AiMessage("Hello Alice!")
);
memoryStore.updateMessages(sessionId1, aliceMessages);
// 4. Append multimodal messages (text + images) for Bob
List<ChatMessage> bobMessages = List.of(
new UserMessage("What do you see in this image?", List.of(new ImageContent("https://...")))
);
memoryStore.updateMessages(sessionId2, bobMessages);
// When we retrieve or delete messages using sessionId1,
// the graph guarantees that Bob's linked list of messages remains completely untouched.
System.out.println("Isolated memory chains created for both Alice and Bob.");
// 5. Optionally delete messages
// memoryStore.deleteMessages(sessionId1);
// memoryStore.deleteMessages(sessionId2);
定制图结构模式
将“定制图表架构”阶段视为可度量的对象来处理效果最佳。在扩大范围之前,先记录一个成功的案例、一个失败案例以及回滚说明。在功能结果旁同时记录处理时间以及令牌或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外费用。应将分块策略与检索策略分开,当质量指标发生变化时,修改其中一项不应迫使重新编写另一项。
// If I want to align the memory storage with my specific domain ontology:
Neo4jChatMemoryStore customMemoryStore = Neo4jChatMemoryStore.builder()
.driver(driver)
.memoryLabel("UserSession") // Overrides the default "Memory" label
.messageLabel("ChatTurn") // Overrides the default "Message" label
.lastMessageRelType("LATEST_CHAT") // Overrides the default "LAST_MESSAGE" rel
.nextMessageRelType("FOLLOWED_BY") // Overrides the default "NEXT" rel
.idProperty("sessionKey") // Overrides the default "id" property
.messageProperty("textContent") // Overrides the default "message" property
.build();
// Now, when the system persists a chat, it will execute domain-specific Cypher queries like:
// MERGE (m:UserSession {sessionKey: 'user-alice-123'})
// CREATE (msg:ChatTurn {textContent: 'Hi...'})
// MERGE (m)-[:LATEST_CHAT]->(msg)
System.out.println("Custom memory store initialized with domain-specific schema.");
管理令牌限制(上下文窗口大小)
将“管理令牌限制上下文”阶段视为可度量的对象来处理,效果最佳。在扩大范围之前,需记录一份理想的操作日志、一个故障案例以及回滚说明。 配置应置于应用程序代码之外。环境文件、密钥存储和功能标志应集中存放于一处,以便操作人员无需查看整个系统结构即可进行审计。 为每轮对话和每次会话设定令牌预算。智能工具往往会大量消耗上下文资源,设置上限可避免演示过程突然产生额外费用。 将“管理令牌限制上下文”阶段视为可度量的对象来处理,效果最佳。在扩大范围之前,需记录一份理想的操作日志、一个故障案例以及回滚说明。 相较于庞大的脚本,应优先使用小型且可测试的单元。当某个步骤出现故障时,故障原因应能明确指向某个特定功能,而非整个复杂的流程。
// If I want to heavily restrict the context window to only the most recent interactions:
Neo4jChatMemoryStore slidingWindowStore = Neo4jChatMemoryStore.builder()
.driver(driver)
.size(3) // The default is 10. We limit it to the 3 most recent messages.
.build();
// Assuming the user has sent 20 messages in this session over the past month.
List<ChatMessage> recentHistory = slidingWindowStore.getMessages("user-alice-123");
// The system efficiently traverses the graph starting from the LATEST_CHAT relationship
// and walks backwards via the FOLLOWED_BY relationships, stopping after it collects the
// limited batch of recent messages. The older historical messages remain safely in the
// database, but are not loaded into memory, saving precious tokens.
System.out.println("Loaded only the " + recentHistory.size() + " most recent messages.");
// If I want to extract the complete history for analytics or summarization:
Neo4jChatMemoryStore completeHistoryStore = Neo4jChatMemoryStore.builder()
.driver(driver)
.size(0) // 0 disables the sliding window limit
.build();
List<ChatMessage> fullHistory = completeHistoryStore.getMessages("user-alice-123");
System.out.println("Extracted the complete session history containing " + fullHistory.size() + " messages.");
简化版连接处理
在采用简化版连接处理流程时,应在修改代码之前明确输入内容、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 应将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,明确成功判定标准,并杜绝无声的半完成状态。 需引用实际作为答案依据的段落。若没有引用,操作人员就无法区分是幻觉内容还是索引缺失导致的错误。
// If I want to instantiate the store directly without managing an external Driver instance:
Neo4jChatMemoryStore standaloneStore = Neo4jChatMemoryStore.builder()
.withBasicAuth("bolt://localhost:7687", "neo4j", "password")
.build();
System.out.println("Memory store connected directly via Basic Auth.");
结论
在结论阶段,应在修改代码之前明确输入参数、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。除了功能结果外,还需记录执行时间以及令牌或查询成本。提前显示成本信息,可避免在从演示环境过渡到共享环境时出现意外费用。必须注明支撑答案的具体内容段落;没有引用的话,操作人员就无法区分是幻觉内容还是索引缺失导致的错误。
资源
在资源准备阶段,应在修改代码之前明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测其中的隐藏状态。 配置信息应置于应用程序代码之外。环境文件、密钥存储以及功能开关应集中存放于一个位置,以便操作人员无需查看整个流程即可进行审计。 需引用实际作为答案依据的段落。如果没有引用,操作人员就无法区分是虚假信息还是索引缺失导致的错误。 在资源准备阶段,应在修改代码之前明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测其中的隐藏状态。 相较于庞大的脚本,更应优先使用小型且可测试的单元。当某个步骤失败时,故障原因应能明确指向某个特定功能,而非整个复杂的流程。
操作检查清单
在制定操作检查清单时,需明确输入参数、各步骤的负责人以及代码修改后的退出标准。操作人员应能够从已知的检查点重新执行相应步骤,而无需猜测隐藏状态。
需同时记录正常流程与故障恢复流程。重试机制、人工审核环节以及错误处理方式都是产品功能的一部分,而非后续需要补充的内容。
必须注明支撑答案的具体依据。若没有引用,操作人员就无法区分是虚假信息还是索引缺失导致的错误。
编写简短的操作手册:包括如何轮换密钥、如何清空队列以及如何回滚最近的导入操作。
优先选择小型、可测试的单元,而非庞大的脚本。当某个步骤出现故障时,故障原因应能明确指向某个具体责任模块,而非整个复杂的处理流程。
请引用那些真正作为答案依据的段落。没有引用的话,操作人员就无法区分是幻觉还是索引缺失导致的错误。
在推广该技术栈之前,先冻结版本,为关键流程保存完整的记录,并确认回滚步骤。共享环境需要设置速率限制、进行租户检查,同时明确密钥轮换的负责人。与其展示花哨的一次性演示,不如追求扎实的可靠性。
关于 9f23f8fe623e 的批量说明:不要将提供商密钥放入代码仓库,为每个会话设置令牌使用上限,并将记录与评估用的固定文件放在一起,以便后续更换模型时仍能保持对比性。