首页 / 文章 / 实用笔记:掌握 Neo4j 与 LangChain4j:GraphRAG、持久化 AI 内存

实用笔记:掌握 Neo4j 与 LangChain4j:GraphRAG、持久化 AI 内存

《实用笔记》操作指南:掌握 Neo4j 与 LangChain4j——GraphRAG、持久化 AI 内存,以及适用于采用该架构的团队的合约、校验机制与即插即用代码模块。

5031 词

以下笔记为“掌握 Neo4j 和 LangChain4j:GraphRAG、持久化 AI 内存及更多功能”提供了一条实用的学习路径。重点在于契约定义、校验机制以及可直接插入的代码占位符,而非激励性表述。 在完成概览阶段时,首先明确契约内容:所需输入、成功标志以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 建议采用小型、可测试的单元而非庞大的脚本。当某一步骤失败时,故障应能指向单一责任点,而非复杂的流程链。

<dependencies>
  <dependency>
    <groupId>dev.langchain4j</groupId>
    <artifactId>langchain4j-community-neo4j-retriever</artifactId>
    <version>${langchain.version}</version>
  </dependency>
  <dependency>
    <groupId>dev.langchain4j</groupId>
    <artifactId>langchain4j</artifactId>
    <version>${langchain.version}</version>
  </dependency>
  <dependency>
    <groupId>dev.langchain4j</groupId>
    <artifactId>langchain4j-community-neo4j</artifactId>
    <version>${langchain.version}</version>
  </dependency>
  <!-- other deps -->
</dependencies>

动态模式抽象:Neo4jGraph

将动态模式抽象 Neo4jGraph 阶段视为可度量的处理环节时,其效果最佳。在扩大范围之前,先记录一份理想的输出结果、一个失败案例以及回滚说明。 把这一阶段视为输入与经过验证的输出之间的契约。为相关成果命名,明确成功标准,绝不允许出现无声无息的半完成状态。 将分块策略与检索策略分开。当质量指标发生变化时,修改其中一项不应强制要求重新编写另一项。

单查询模式检索

将单次查询模式检索阶段视为可度量的对象来处理,其效果最佳。在扩大范围之前,先记录一份理想的查询结果、一个失败案例以及回滚说明。在功能结果旁同时记录处理时间以及令牌或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外费用。应将分块策略与检索策略分开,当质量指标发生变化时,修改其中一项无需强制重写另一项。

import dev.langchain4j.store.graph.neo4j.Neo4jGraph;
// If I want to initialize the graph abstraction and load the schema:
Neo4jGraph graph = Neo4jGraph.builder()
    .driver(driver)
    .build();
// Under the hood, this executes a single, consolidated APOC query that yields
// labels, element types, and properties all at once, avoiding multiple DB calls.
graph.refreshSchema();
Neo4jGraph.StructuredSchema schema = graph.getStructuredSchema();
// We expect a well-formatted string logically divided into three sections,
// exactly as formatted by the new Neo4jGraphSchemaUtils class:
//
// `schema.nodesProperties()` is the following:
// :Person {name: STRING}, :Company {name: STRING}
//
// `schema.relationshipsProperties()` is the following:
// :WORKS_FOR {since: INTEGER}
//
// `schema.patterns()` is the following:
// (:Person)-[:WORKS_FOR]->(:Company)
System.out.println("Current Database Schema Context:\n" + schema);

配置模式采样与增强功能

将“配置模式采样”与相关阶段视为可度量的对象来处理时,其效果最佳。在扩大范围之前,需记录一份理想状态下的转录内容、一个故障案例以及回滚说明。 应将配置信息与应用程序代码分开。环境文件、密钥存储和功能标志应集中存放于一处,以便操作人员无需查看整个系统结构即可进行审计。 需将分块策略与检索策略区分开来。当质量指标发生变化时,修改其中一项不应迫使重新编写另一项。 将“配置模式采样”与相关阶段视为可度量的对象来处理时,其效果最佳。在扩大范围之前,需记录一份理想状态下的转录内容、一个故障案例以及回滚说明。 相较于庞大的脚本,应优先使用小型且可测试的单元。当某个步骤出现故障时,故障原因应能明确指向单一责任模块,而非复杂的流程链。

// If I want to scan a very large database efficiently by configuring the APOC sampling,
// and I want to enhance the LLM's understanding with sample property values:
Neo4jGraph optimizedGraph = Neo4jGraph.builder()
    .driver(driver)
    // We can configure the underlying apoc.meta.data parameters.
    // 'sample' limits the number of nodes inspected per label to speed up execution.
    // 'maxRels' limits the number of relationships inspected per node.
    // (Note: These are passed internally to the getSchemaFromMetadata utility)
    .build();

// When the schema is refreshed, the underlying query runs:
// CALL apoc.meta.data({maxRels: $maxRels, sample: $sample})
optimizedGraph.refreshSchema();
// The LLM now receives a fast, accurately sampled schema representation,
// protecting database performance during application startup or schema refreshes.
System.out.println("Optimized Schema loaded successfully.");

自动化知识图谱构建与源链接关联

在自动化知识图谱构建阶段,修改代码之前需明确输入内容、该步骤的负责人以及完成标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 应将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,定义成功判定标准,杜绝无声的半完成状态。 必须注明支撑答案的具体依据。若没有引用,操作人员就无法区分幻觉内容与索引缺失问题。

初始数据集与少样本提示

在处理初始数据集及相应阶段时,应在修改代码之前明确输入内容、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。除了功能结果外,还需记录执行时间以及令牌或查询成本。提前了解这些成本信息,可避免在从演示环境过渡到共享环境时出现意外费用。当下一步操作为编写代码或调用工具时,应优先选择具有架构验证的结构化输出,而非自由形式的文本。

[
   {
      "tail": "Microsoft",
      "head": "Adam",
      "head_type": "Person",
      "text": "Adam is a software engineer in Microsoft since 2009...",
      "relation": "WORKS_FOR",
      "tail_type": "Company"
   },
   {
      "tail": "Microsoft Word",
      "head": "Microsoft",
      "head_type": "Company",
      "text": "Microsoft is a tech company that provides several products...",
      "relation": "PRODUCED_BY",
      "tail_type": "Product"
   }
]
import dev.langchain4j.community.data.document.transformer.graph.LLMGraphTransformer;
import dev.langchain4j.community.data.document.graph.GraphDocument;
import dev.langchain4j.data.document.Document;
import dev.langchain4j.data.document.DefaultDocument;
import dev.langchain4j.data.document.Metadata;
import java.util.List;

ChatModel chatModel = /* dev.langchain4j.model.chat instance */
Driver driver = /* org.neo4j.driver.Driver instance */

// If I want to guide the extraction by providing a structured set of examples:
LLMGraphTransformer transformer = LLMGraphTransformer.builder()
    .model(chatModel)
    .examples(EXAMPLES_PROMPT) // Injects the above JSON dataset into the system prompt
    .build();
Document docKeanu = new DefaultDocument(
    "Keanu Reeves acted in Matrix",
    Metadata.from("key33", "value3")
);
// The LLM will transform the text, structuring nodes and relationships based on the examples
List<GraphDocument> graphDocs = transformer.transformAll(List.of(docKeanu));

/*
The above `graphDocs` returns this result:

GraphDocument
├─ Nodes
│  ├─ GraphNode
│  │  ├─ id: Matrix
│  │  ├─ type: Movie
│  │  └─ properties: {}
│  │
│  └─ GraphNode
│     ├─ id: Keanu Reeves
│     ├─ type: Person
│     └─ properties: {}
│
├─ Relationships
│  └─ GraphEdge
│     ├─ type: ACTED_IN
│     ├─ sourceNode
│     │  ├─ id: Keanu Reeves
│     │  └─ type: Person
│     ├─ targetNode
│     │  ├─ id: Matrix
│     │  └─ type: Movie
│     └─ properties: {}
│
└─ Source
   ├─ text: "Keanu Reeves acted in Matrix"
   └─ metadata
      └─ key33: value3
*/

可重试的图数据持久化

在幂等图持久化阶段,修改代码之前需明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 配置信息应置于应用程序代码之外。环境文件、密钥存储和功能标志应集中存放于一个位置,以便操作人员无需查看整个系统结构即可进行审计。 需引用实际作为答案依据的段落。没有引用的话,操作人员就无法区分是虚假信息还是索引缺失导致的错误。 在幂等图持久化阶段,修改代码之前需明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 相较于庞大的脚本,应优先使用小型且可测试的单元。当某个步骤失败时,故障原因应能指向单一责任模块,而非多个部分共同导致的问题。

带角度的管道。

import dev.langchain4j.community.rag.content.retriever.neo4j.KnowledgeGraphWriter;

Neo4jGraph neo4jGraph = /* dev.langchain4j.store.graph.neo4j.Neo4jGraph instance */;
LLMGraphTransformer graphTransformer = /* dev.langchain4j.community.data.document.transformer.graph.LLMGraphTransformer instance */;

Document docKeanu = new DefaultDocument(
        "Keanu Reeves acted in Matrix",
        Metadata.from("key33", "value3")
);
List<GraphDocument> graphDocs = graphTransformer.transformAll(List.of(docKeanu));


// If I want to persist the extracted entities safely:
KnowledgeGraphWriter writer = KnowledgeGraphWriter.builder()
    .graph(neo4jGraph)
    .build();
// The first write populates the database
writer.addGraphDocuments(graphDocs, false);
// Executing the exact same command again is safe.
// The internal UNWIND and MERGE logic generated by the writer guarantees
// that no duplicate entities or relationships are created.
writer.addGraphDocuments(graphDocs, false);
// Expected resulting topology in the database:
// (:__Entity__ {id: 'keanu'})-[:ACTED]->(:__Entity__ {id: 'matrix'})
System.out.println("Entities persisted successfully. No duplicates created.");

源文档链接(includeSource)

在处理源文档链接的includeSource阶段时,首先需明确相关约定:所需的输入参数、成功信号以及部分失败时的处理方式。这样的清单能确保后续的代码修改始终符合约定。 将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,定义成功判定标准,杜绝无声的半完成状态。 在调整提示词之前,先使用固定的问题集来测试召回率。仅仅更换提示词往往无法改善较差的检索效果。

// If I want to maintain data provenance and link entities back to their source:
writer.addGraphDocuments(graphDocs, true); // true = includeSource

// Behind the scenes, the writer executes three crucial operations:
// 1. It creates the Document node, copying the original metadata (e.g., key33: value3) and the text.
// 2. If the original document lacks an ID, the writer automatically generates
//    an MD5 hash of the text to use as a unique identifier.
// 3. It links the document to the extracted entities using a relationship (default: HAS_ENTITY).
// Expected resulting topology in the database:
// (:Document {id: '<MD5_hash>', text: 'Keanu Reeves...', key33: 'value3'})-[:HAS_ENTITY]->(:__Entity__ {id: 'keanu'})
System.out.println("Source document successfully linked to the extracted entities.");

深度模式定制与安全性

在处理深度模式定制阶段时,首先需明确相关约定:所需的输入参数、成功标识以及部分失败时的处理方式。这样的清单能确保后续的代码修改始终符合预期。 在功能结果旁记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外费用。 在调整提示词之前,先使用固定的问题集测试召回率。仅仅更换提示词往往无法改善较差的检索效果。

// If I want to adapt the ingestion to a pre-existing enterprise schema,
// and customize the relationship that links the document to the entities:
KnowledgeGraphWriter customWriter = KnowledgeGraphWriter.builder()
    .graph(neo4jGraph)
    .label("ActorOrMovie")             // Replaces "__Entity__"
    .idProperty("customId")            // Replaces "id"
    .textProperty("customText")        // Replaces "text" for the Document node
    .relType("MENTIONED_IN_SOURCE")    // Replaces "HAS_ENTITY"
    .constraintName("unique_custom")   // Sets a specific name for the Neo4j CONSTRAINT
    .build();

// Inserting the data with includeSource set to true will now use the new nomenclature:
customWriter.addGraphDocuments(graphDocs, true);

// Example of the resulting Cypher pattern generated by the writer:
// (:Document {customText: '...'})-[:MENTIONED_IN_SOURCE]->(:ActorOrMovie {customId: 'keanu'})

通过 Cypher DSL 集成实现安全保障

在处理“通过 Cypher DSL 实现安全保障”这一阶段时,首先需明确相关规范:所需的输入参数、成功标识,以及部分失败时的处理方式。这样的清单能确保后续的代码修改始终符合既定要求。 应将配置信息与应用程序代码分开存放。环境文件、密钥存储以及功能开关应集中管理,以便操作人员无需查看整个系统结构即可进行审计。 在调整提示词之前,先使用固定的问题集来测试召回率。仅仅更换提示词往往无法解决检索效果不佳的问题。 在处理“通过 Cypher DSL 实现安全保障”这一阶段时,首先需明确相关规范:所需的输入参数、成功标识,以及部分失败时的处理方式。这样的清单能确保后续的代码修改始终符合既定要求。 相比庞大的脚本,更应优先使用小型且可测试的单元。当某个步骤出现故障时,故障点应能明确指向某一个具体的功能模块,而非整个复杂的流程。

import org.neo4j.cypherdsl.core.Cypher;
import org.neo4j.cypherdsl.core.Statement;
import org.neo4j.cypherdsl.core.Node;

Driver driver = /* org.neo4j.driver.Driver instance */

// Demonstrating how internal queries are constructed safely via the DSL
// We define a Node representation first
Node documentNode = Cypher.node("Document").named("d");
// We build the query programmatically using the fluent API.
// Notice how literal values are wrapped safely, preventing injection.
Statement statement = Cypher.match(documentNode)
    .where(Cypher.property("d", "id").isEqualTo(Cypher.literalOf("doc-123")))
    .returning(Cypher.property("d", "text"))
    .build();
// The DSL engine traverses the AST and compiles it into a syntactically safe string.
String safeCypherQuery = statement.getCypher();
// Expected output: MATCH (d:`Document`) WHERE d.id = 'doc-123' RETURN d.text
System.out.println("Generated safe Cypher via DSL: " + safeCypherQuery);

使用2026.01语法进行索引内预过滤

将“使用2026版本进行索引内预过滤”这一阶段视为可度量的流程最为有效。在扩大范围之前,需收集一份理想的处理结果、一个失败案例以及回滚说明。 应将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,明确成功标准,杜绝默许的半完成状态。 需将分块策略与检索策略分开处理。当质量指标发生变化时,修改其中一项不应强制要求重新编写另一项。

import dev.langchain4j.store.embedding.neo4j.Neo4jEmbeddingStore;
import dev.langchain4j.store.embedding.filter.Filter;
import static dev.langchain4j.store.embedding.filter.MetadataFilterBuilder.metadataKey;

Driver driver = /* org.neo4j.driver.Driver instance */
Embedding embedding = /* dev.langchain4j.data.embedding.Embedding instance */

MatchSearchClauseStrategy matchSearchClauseStrategy = new MatchSearchClauseStrategy();

// Configure the store with the new syntax enabled
Neo4jEmbeddingStore store = Neo4jEmbeddingStore.builder()
    .driver(driver)
    .dimension(1536)
    .searchStrategy(matchSearchClauseStrategy) // Crucial flag: Enables the 2026.01 optimized native vector search syntax
    .filterMetadata(Arrays.asList("year", "department")) // Enable filtering for 'year' and 'department', which translates to the `WITH [indexName.year, indexName.department]` clause during index creation
    .build();

// Build a metadata filter combining multiple boolean conditions
Filter filter = metadataKey("year").isEqualTo(2024)
    .and(metadataKey("department").isEqualTo("Engineering"));
// Execute the search request
EmbeddingSearchRequest request = EmbeddingSearchRequest.builder()
    .queryEmbedding(embedding)
    .maxResults(5)
    .filter(filter) // Filter is executed natively inside the Neo4j Vector Index block
    .build();
SearchResult<TextSegment> results = store.search(request);
// The results will natively exclude any documents not matching the criteria specified in the `filter` instance
// returning the final Top-K, ensuring you always get 5 highly relevant segments if they exist.
System.out.println("Search executed with in-index filtering.");
// Execute results.matches() to verify in-index filtering

GraphRAG检索概念:高级结构化检索器与数据导入工具

将GraphRAG检索概念的进阶阶段视为可度量的对象来处理,效果最佳。在扩大应用范围之前,先记录一个成功的案例、一个失败案例以及回滚说明。在功能结果旁同时记录处理时间以及令牌或查询成本。提前了解成本情况,就能避免在从演示环境过渡到共享环境时出现意外费用。应将分块策略与检索策略分开,当质量指标发生变化时,修改其中一项无需强制重写另一项。

父子模式

将“父子模式”阶段视为可度量的对象来处理时,其效果最佳。在扩大范围之前,需记录一份理想案例、一个故障实例以及回滚说明。 应将配置与应用程序代码分开。环境文件、密钥存储和功能标志应集中存放于一处,以便操作人员无需查看整个架构即可进行审计。 应将分块策略与检索策略分开。当质量指标发生变化时,修改其中一项不应迫使重新编写另一项。 将“父子模式”阶段视为可度量的对象来处理时,其效果最佳。在扩大范围之前,需记录一份理想案例、一个故障实例以及回滚说明。 相比庞大的脚本,更应采用小型且可测试的单元。当某个步骤出现故障时,故障应指向单一责任模块,而非复杂的流程链。

import dev.langchain4j.store.graph.neo4j.Neo4jParentChildIngestor;
import dev.langchain4j.data.document.Document;
import dev.langchain4j.data.document.splitter.DocumentSplitters;

// embeddingModel, embeddingStore, childSplitter instances...

// If I want to automatically ingest a document into a Parent-Child graph topology:
Neo4jEmbeddingStoreIngestor ingestor = ParentChildGraphIngestor.builder()
          .driver(driver)
          .embeddingModel(embeddingModel)
          // We define how the document should be chunked before ingestion
          .documentSplitter(DocumentSplitters.recursive(200, 20))
          .documentChildSplitter(childSplitter)
          .build();

Document document = Document.from( """Artificial Intelligence (AI) is a field of computer science. It focuses on creating intelligent agents capable of performing tasks that require human intelligence.
                        Machine Learning (ML) is a subset of AI. It uses data to learn patterns and make predictions. Deep Learning is a specialized form of ML based on neural networks.
                        """);

ingestor.ingest(List.of(document));
// The graph now contains one 'Document' node connected via 'HAS_CHILD'
// to multiple embedded 'DocumentChunk' nodes.
System.out.println("Parent and child nodes successfully ingested and linked.");
import dev.langchain4j.rag.content.retriever.EmbeddingStoreContentRetriever;


// If I want to search against chunks but retrieve the rich parent document:
final EmbeddingStoreContentRetriever retriever = EmbeddingStoreContentRetriever.builder()
        .embeddingStore(embeddingStore)
        .maxResults(1)
        .minScore(0.4)
        .build();

List<Content> contents = retriever.retrieve(Query.from("specific configuration detail"));
// The retriever hits the small 'DocumentChunk' index, traverses the 'HAS_CHILD'
// relationship, and returns the entire 'Document' node.
System.out.println("Retrieved full parent document context.");
/*
The result of `contents` is :

DefaultContent {
  textSegment = TextSegment {
    text = """
      Machine Learning (ML) is a subset of AI. It uses data to learn patterns and make predictions.
      Deep Learning is a specialized form of ML based on neural networks.

      Machine Learning (ML) is a subset of AI. It uses data to learn patterns and make predictions.
      Deep Learning is a specialized form of ML based on neural networks.

      Artificial Intelligence (AI) is a field of computer science. It focuses on creating intelligent agents
      capable of performing tasks that require human intelligence.
    """,
    metadata = {
      index = 1,
      source = Wikipedia link,
      title = AI Basics,
      url = https://example.com/ai,
      parentId = parent_1af3e080-5029-40ab-b3d7-0829064d800d
    }
  },
  metadata = {
    EMBEDDING_ID = null,
    SCORE = 0.8560410737991333
  }
}
*/

总结模式

在“总结模式”阶段,应在修改代码之前明确输入内容、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,定义成功判定标准,并拒绝默许的半完成状态。 需引用实际作为答案依据的段落。没有引用的话,操作人员就无法区分幻觉内容与索引缺失的情况。

import dev.langchain4j.community.store.embedding.neo4j.Neo4jEmbeddingStoreIngestor;
import dev.langchain4j.community.store.embedding.neo4j.SummaryGraphIngestor;

// If I want the LLM to summarize my document during ingestion and store the summary:
/* Neo4jEmbeddingStore, ChatModel and DocumentSplitter instances... */

final Neo4jEmbeddingStoreIngestor ingestor = SummaryGraphIngestor.builder()
        .driver(driver)
        .embeddingModel(embeddingModel)
        .questionModel(chatModel)
        .documentSplitter(parentSplitter)
        .build();

ingestor.ingest(List.of(document));
// The graph now has a 'Summary' node (containing the LLM-generated summary)
// linked to the specific 'DocumentChunk' nodes.
System.out.println("Document ingested and summarized successfully.");
import dev.langchain4j.community.store.embedding.neo4j.Neo4jEmbeddingStoreIngestor;
import dev.langchain4j.community.store.embedding.neo4j.SummaryGraphIngestor;

/* required instances */

Document document = Document.from("""
        Artificial Intelligence (AI) is a field of computer science. It focuses on creating intelligent agents capable of performing tasks that require human intelligence.
        Machine Learning (ML) is a subset of AI. It uses data to learn patterns and make predictions. Deep Learning is a specialized form of ML based on neural networks.
        """);

ingestor.ingest(document);

final EmbeddingStoreContentRetriever retriever = EmbeddingStoreContentRetriever.builder()
        .embeddingModel(embeddingModel)
        .maxResults(5)
        .minScore(0.6)
        .embeddingStore(ingestor.getEmbeddingStore())
        .build();

/*
The result is something like this:

DefaultContent {
  textSegment = TextSegment {
    text = "Machine Learning (ML) is a subset of AI",
    metadata = {
      index  = 0,
      source = Wikipedia link,
      title  = Quantum Mechanics,
      url    = https://example.com/ai
    }
  },
  metadata = {
    EMBEDDING_ID = null,
    SCORE        = 0.8425111770629883
  }
}
*/

假设性问题模式

在“假设性问题模式”阶段,应在修改代码之前明确输入内容、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。应在功能结果旁记录执行时间以及令牌或查询成本。提前显示成本信息,可避免在从演示环境过渡到共享环境时出现意外费用。需注明实际作为答案依据的段落;没有引用的话,操作人员就无法区分是幻觉内容还是索引缺失导致的错误。

import dev.langchain4j.community.store.embedding.neo4j.Neo4jEmbeddingStoreIngestor;
import dev.langchain4j.community.store.embedding.neo4j.HypotheticalQuestionGraphIngestor;

/* ... Neo4jEmbeddingStore, ChatModel and DocumentSplitter instances.. */

Neo4jEmbeddingStoreIngestor ingestor = HypotheticalQuestionGraphIngestor.builder()
        .embeddingModel(embeddingModel)
        .driver(driver)
        .documentSplitter(splitter)
        .questionModel(chatModel)
        .embeddingStore(embeddingStore)
        .build();

Document document = Document.from("""
                        Quantum mechanics studies how particles behave. It is a fundamental theory in physics.
                        Gradient descent and backpropagation algorithms.
                        Spaghetti carbonara and Italian dishes.
                        John Doe is a Super Saiyan.
                        """);
ingestor.ingest(document);

EmbeddingStoreContentRetriever retriever = EmbeddingStoreContentRetriever.builder()
                .embeddingModel(embeddingModel)
                .maxResults(2)
                .minScore(0.5)
                .embeddingStore(ingestor.getEmbeddingStore())
                .build();

List<Content> results = retriever.retrieve(Query.from("Who is John Doe?"));

System.out.println("Retrieved Hypothetical context: " + results);

/*
The result is something like this:

DefaultContent {
  textSegment = TextSegment {
    text = "John Doe is a Super Saiyan.",
    metadata = {
      index  = 2,
      source = Wikipedia link,
      title  = Quantum Mechanics,
      url    = https://example.com/ai
    }
  },
  metadata = {
    SCORE        = 0.8234479427337646,
    EMBEDDING_ID = 2002cabe-2a3e-4c6e-96ec-a0292e26e817
  }
}

*/

通用模式

在“通用模式”阶段,应在修改代码之前明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测其中的隐藏状态。 配置信息应置于应用程序代码之外。环境文件、密钥存储以及功能标志应集中存放于一个位置,以便操作人员无需查看整个流程即可进行审计。 需引用实际作为答案依据的段落。如果没有引用,操作人员就无法区分是虚假信息还是索引缺失导致的错误。 在“通用模式”阶段,应在修改代码之前明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测其中的隐藏状态。 相较于庞大的脚本,更应优先使用小型且可测试的单元。当某个步骤出现故障时,故障原因应能指向单一责任模块,而非复杂的流程链。

final Neo4jEmbeddingStore neo4jEmbeddingStore = /* Neo4jEmbeddingStore instance */

final EmbeddingStoreContentRetriever retriever = EmbeddingStoreContentRetriever.builder()
        .embeddingModel(embeddingModel)
        .maxResults(5)
        .minScore(0.4)
        .embeddingStore(neo4jEmbeddingStore)
        .build();

// other required instances ...

Document doc = Document.from("""
                        Quantum mechanics studies how particles behave. It is a fundamental theory in physics.
                        Gradient descent and backpropagation algorithms.
                        Spaghetti carbonara and Italian dishes.
                        John Doe is a Super Saiyan.
                        """);

// Ingest the document into Neo4j as parent-child nodes
final Neo4jEmbeddingStoreIngestor ingestor = Neo4jEmbeddingStoreIngestor.builder()
        .documentSplitter(parentSplitter)
        .documentChildSplitter(childSplitter)
        .driver(driver)
        .query("CREATE (:MainDoc $metadata)") // a Cypher query template used for storing the processed segment data in Neo4j
        .embeddingStore(neo4jEmbeddingStore)
        .embeddingModel(embeddingModel)
        .build();

ingestor.ingest(doc);

final String retrieveQuery = "Machine Learning";
List<Content> results = retriever.retrieve(Query.from(retrieveQuery));

System.out.println("Retrieved Generic context: " + results);

与数据库类型无关的父子关系检索工具

在处理“与数据库类型无关的父子关系”这一阶段时,首先需明确相关约定:所需的输入参数、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 将这一阶段视为输入与经过验证的输出之间的契约。为相关成果命名,定义成功判定标准,杜绝无声的半完成状态。 在调整提示词之前,先使用固定的问题集来衡量检索的召回率。仅仅更换提示词很难改善较差的检索效果。

import dev.langchain4j.community.store.embedding.ParentChildEmbeddingStoreIngestor;

/* Required instances */

ParentChildEmbeddingStoreIngestor ingestor = ParentChildEmbeddingStoreIngestor.builder()
                .documentTransformer(documentTransformer)
                .documentSplitter(documentSplitter)
                .textSegmentTransformer(textSegmentTransformer)
                .embeddingModel(embeddingModel)
                .embeddingStore(embeddingStore)
                .documentChildSplitter(documentChildSplitter)
                .childTextSegmentTransformer(childTextSegmentTransformer)
                .build();

有状态人工智能:图结构中的持久对话记忆

在处理有状态AI持久对话阶段时,首先需列出相关规范:所需输入、成功信号以及部分失败时的处理方式。这样的检查清单能确保后续的代码修改保持一致性。 在功能结果旁记录执行时间以及令牌或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外费用。 在调整提示词之前,先使用固定的问题集测试召回率。仅仅更换提示词往往无法改善较差的检索效果。

配置

在配置阶段工作时,首先需明确相关规范:所需的输入参数、成功信号以及部分失败时的处理方式。这份清单能确保后续的代码修改始终符合预期。 应将配置信息与应用程序代码分开存放。环境文件、密钥存储以及功能开关应集中于一个位置,这样操作人员无需查看整个系统结构即可进行审计。 在调整提示词之前,先使用固定的问题集来测试系统的召回率。仅仅更换提示词往往无法解决检索效果不佳的问题。 在配置阶段工作时,首先需明确相关规范:所需的输入参数、成功信号以及部分失败时的处理方式。这份清单能确保后续的代码修改始终符合预期。 相比庞大的脚本,更应优先使用小型且可测试的单元。当某个步骤出现故障时,故障点应能明确指向某个具体的功能模块,而非整个复杂的处理流程。

管理多租户聊天记录与多模态输入

将“管理多租户聊天记录”这一阶段视为可度量的工作对象会更为有效。在扩大范围之前,先记录一份最佳处理案例、一个失败案例以及回滚说明。 应将此阶段视为输入与经过验证的输出之间的契约。为相关文档命名,明确成功标准,杜绝默许的半完成状态。 需将分块策略与检索策略分开处理。当质量指标发生变化时,调整其中一项不应强制要求重新编写另一项。

import dev.langchain4j.store.memory.chat.neo4j.Neo4jChatMemoryStore;
import dev.langchain4j.data.message.UserMessage;
import dev.langchain4j.data.message.AiMessage;
import dev.langchain4j.data.message.ImageContent;
import java.util.List;

// 1. Initialize the store using an existing driver
Neo4jChatMemoryStore memoryStore = Neo4jChatMemoryStore.builder()
    .driver(driver)
    .build();
// 2. Identify the specific user sessions
String sessionId1 = "user-alice-123";
String sessionId2 = "user-bob-456";
// 3. Append standard text messages for Alice
List<ChatMessage> aliceMessages = List.of(
    new UserMessage("Hi, I'm Alice."),
    new AiMessage("Hello Alice!")
);
memoryStore.updateMessages(sessionId1, aliceMessages);
// 4. Append multimodal messages (text + images) for Bob
List<ChatMessage> bobMessages = List.of(
    new UserMessage("What do you see in this image?", List.of(new ImageContent("https://...")))
);
memoryStore.updateMessages(sessionId2, bobMessages);
// When we retrieve or delete messages using sessionId1,
// the graph guarantees that Bob's linked list of messages remains completely untouched.
System.out.println("Isolated memory chains created for both Alice and Bob.");


// 5. Optionally delete messages
// memoryStore.deleteMessages(sessionId1);
// memoryStore.deleteMessages(sessionId2);

定制图结构模式

将“定制图表架构”阶段视为可度量的对象来处理效果最佳。在扩大范围之前,先记录一个成功的案例、一个失败案例以及回滚说明。在功能结果旁同时记录处理时间以及令牌或查询成本。提前了解成本情况,可避免在从演示环境过渡到共享环境时出现意外费用。应将分块策略与检索策略分开,当质量指标发生变化时,修改其中一项不应迫使重新编写另一项。

// If I want to align the memory storage with my specific domain ontology:
Neo4jChatMemoryStore customMemoryStore = Neo4jChatMemoryStore.builder()
    .driver(driver)
    .memoryLabel("UserSession")         // Overrides the default "Memory" label
    .messageLabel("ChatTurn")           // Overrides the default "Message" label
    .lastMessageRelType("LATEST_CHAT")  // Overrides the default "LAST_MESSAGE" rel
    .nextMessageRelType("FOLLOWED_BY")  // Overrides the default "NEXT" rel
    .idProperty("sessionKey")           // Overrides the default "id" property
    .messageProperty("textContent")     // Overrides the default "message" property
    .build();

// Now, when the system persists a chat, it will execute domain-specific Cypher queries like:
// MERGE (m:UserSession {sessionKey: 'user-alice-123'})
// CREATE (msg:ChatTurn {textContent: 'Hi...'})
// MERGE (m)-[:LATEST_CHAT]->(msg)
System.out.println("Custom memory store initialized with domain-specific schema.");

管理令牌限制(上下文窗口大小)

将“管理令牌限制上下文”阶段视为可度量的对象来处理,效果最佳。在扩大范围之前,需记录一份理想的操作日志、一个故障案例以及回滚说明。 配置应置于应用程序代码之外。环境文件、密钥存储和功能标志应集中存放于一处,以便操作人员无需查看整个系统结构即可进行审计。 为每轮对话和每次会话设定令牌预算。智能工具往往会大量消耗上下文资源,设置上限可避免演示过程突然产生额外费用。 将“管理令牌限制上下文”阶段视为可度量的对象来处理,效果最佳。在扩大范围之前,需记录一份理想的操作日志、一个故障案例以及回滚说明。 相较于庞大的脚本,应优先使用小型且可测试的单元。当某个步骤出现故障时,故障原因应能明确指向某个特定功能,而非整个复杂的流程。

// If I want to heavily restrict the context window to only the most recent interactions:
Neo4jChatMemoryStore slidingWindowStore = Neo4jChatMemoryStore.builder()
    .driver(driver)
    .size(3) // The default is 10. We limit it to the 3 most recent messages.
    .build();

// Assuming the user has sent 20 messages in this session over the past month.
List<ChatMessage> recentHistory = slidingWindowStore.getMessages("user-alice-123");
// The system efficiently traverses the graph starting from the LATEST_CHAT relationship
// and walks backwards via the FOLLOWED_BY relationships, stopping after it collects the
// limited batch of recent messages. The older historical messages remain safely in the
// database, but are not loaded into memory, saving precious tokens.
System.out.println("Loaded only the " + recentHistory.size() + " most recent messages.");
// If I want to extract the complete history for analytics or summarization:
Neo4jChatMemoryStore completeHistoryStore = Neo4jChatMemoryStore.builder()
    .driver(driver)
    .size(0) // 0 disables the sliding window limit
    .build();

List<ChatMessage> fullHistory = completeHistoryStore.getMessages("user-alice-123");
System.out.println("Extracted the complete session history containing " + fullHistory.size() + " messages.");

简化版连接处理

在采用简化版连接处理流程时,应在修改代码之前明确输入内容、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。 应将此阶段视为输入与经过验证的输出之间的契约。为相关成果命名,明确成功判定标准,并杜绝无声的半完成状态。 需引用实际作为答案依据的段落。若没有引用,操作人员就无法区分是幻觉内容还是索引缺失导致的错误。

// If I want to instantiate the store directly without managing an external Driver instance:
Neo4jChatMemoryStore standaloneStore = Neo4jChatMemoryStore.builder()
    .withBasicAuth("bolt://localhost:7687", "neo4j", "password")
    .build();
System.out.println("Memory store connected directly via Basic Auth.");

结论

在结论阶段,应在修改代码之前明确输入参数、该步骤的负责人以及结束标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测隐藏状态。除了功能结果外,还需记录执行时间以及令牌或查询成本。提前显示成本信息,可避免在从演示环境过渡到共享环境时出现意外费用。必须注明支撑答案的具体内容段落;没有引用的话,操作人员就无法区分是幻觉内容还是索引缺失导致的错误。

资源

在资源准备阶段,应在修改代码之前明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测其中的隐藏状态。 配置信息应置于应用程序代码之外。环境文件、密钥存储以及功能开关应集中存放于一个位置,以便操作人员无需查看整个流程即可进行审计。 需引用实际作为答案依据的段落。如果没有引用,操作人员就无法区分是虚假信息还是索引缺失导致的错误。 在资源准备阶段,应在修改代码之前明确输入参数、该步骤的负责人以及终止标准。操作人员应能够从已知的检查点重新运行该步骤,而无需猜测其中的隐藏状态。 相较于庞大的脚本,更应优先使用小型且可测试的单元。当某个步骤失败时,故障原因应能明确指向某个特定功能,而非整个复杂的流程。

操作检查清单

在制定操作检查清单时,需明确输入参数、各步骤的负责人以及代码修改后的退出标准。操作人员应能够从已知的检查点重新执行相应步骤,而无需猜测隐藏状态。

需同时记录正常流程与故障恢复流程。重试机制、人工审核环节以及错误处理方式都是产品功能的一部分,而非后续需要补充的内容。

必须注明支撑答案的具体依据。若没有引用,操作人员就无法区分是虚假信息还是索引缺失导致的错误。

编写简短的操作手册:包括如何轮换密钥、如何清空队列以及如何回滚最近的导入操作。

优先选择小型、可测试的单元,而非庞大的脚本。当某个步骤出现故障时,故障原因应能明确指向某个具体责任模块,而非整个复杂的处理流程。

请引用那些真正作为答案依据的段落。没有引用的话,操作人员就无法区分是幻觉还是索引缺失导致的错误。

在推广该技术栈之前,先冻结版本,为关键流程保存完整的记录,并确认回滚步骤。共享环境需要设置速率限制、进行租户检查,同时明确密钥轮换的负责人。与其展示花哨的一次性演示,不如追求扎实的可靠性。

关于 9f23f8fe623e 的批量说明:不要将提供商密钥放入代码仓库,为每个会话设置令牌使用上限,并将记录与评估用的固定文件放在一起,以便后续更换模型时仍能保持对比性。