Home / Articles / Practical notes: Lab:Langgraph: Integrating Qdrant for Semantic Agentic Memory.

This article is published in English.

Practical notes: Lab:Langgraph: Integrating Qdrant for Semantic Agentic Memory.

Operable walkthrough of Practical notes: Lab:Langgraph: Integrating Qdrant for Semantic Agentic Memory.: contracts, checks, and drop-in code slots for teams shipping this pattern.

2151 words

Use this as an operator-facing rebuild of the ideas in “Lab:Langgraph: Integrating Qdrant for Semantic Agentic Memory.”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Why Semantic Memory is Necessary?

For the Why Semantic Memory is stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

fastapi>=0.110.0
uvicorn>=0.28.0
pydantic>=2.6.0
python-dotenv>=1.0.1
langchain-core>=0.1.30
langchain-openai>=0.1.0
langgraph>=0.0.30
qdrant-client>=1.10.0
#Install venv module (Ubuntu/Debian)
$sudo apt update && sudo apt install python3-venv

#Create a virtual environment
$python3 -m venv myenv

#Activate the environment
$source myenv/bin/activate

#Install saved dependencies
$pip install -r requirements.txt

#Save dependencies
#pip freeze > requirements.txt

#Deactivate when finished:
$deactivate

#Delete the environment:
$rm -rf myenv
#Install uv
$curl -LsSf [https://astral.sh/uv/install.sh](https://astral.sh/uv/install.sh) | sh

#Create a virtual environment
$uv venv myenv

#Activate the environment
$source myenv/bin/activate

#Install saved dependencies:
$uv pip install -r requirements.txt

#Save dependencies:
$uv pip freeze > requirements.txt

#Sync dependencies from lockfile
$uv sync

#Add a new package and auto-update file
$uv add <package_name>
version: '3.8'

services:
  qdrant:
    image: qdrant/qdrant:latest    #image name download from docker.io
    container_name: sbi_semantic_memory  #container name
    ports:
      - "6333:6333"   # REST HTTP API & Web Dashboard
      - "6334:6334"   # High-speed gRPC API
    volumes:
      - ./qdrant_storage:/qdrant/storage #Binds local directory to container for data persistence across restarts
    networks:
      - agent-network  #Attaches container to isolated network for communication
    healthcheck:   #Executes internal HTTP ping against health check endpoint
      test: ["CMD", "curl", "-f", "http://localhost:6333/healthz"]
      interval: 10s  #Run health check every 10 seconds
      timeout: 5s    #Fail if check takes longer than 5 seconds
      retries: 5     #Startup grace period before recording failures

networks:    #Private bridged network for secure container-to-container communication
  agent-network:
    driver: bridge   #Standard single-host bridge driver
image: qdrant/qdrant:latest
$docker compose up -d

$  docker ps
CONTAINER ID   IMAGE                  COMMAND             CREATED        STATUS                    PORTS                                                             NAMES
8922a12f1029   qdrant/qdrant:latest   "./entrypoint.sh"   45 hours ago   Up 45 hours (unhealthy)   0.0.0.0:6333-6334->6333-6334/tcp, [::]:6333-6334->6333-6334/tcp   sbi_semantic_memory

$ curl http://localhost:6333/healthz
healthz check passed

#View recent logs
$docker logs sbi_semantic_memory

#Tail live logs in real time
$docker logs -f sbi_semantic_memory

#View the last 50 lines of logs
$docker logs --tail 50 sbi_semantic_memory

#Inspect health check failure reasons
$docker inspect --format='{{json .State.Health}}' sbi_semantic_memory

#Check container resource consumption (CPU/RAM):
$docker stats sbi_semantic_memory

#Inspect container runtime details and exit codes:
$docker inspect sbi_semantic_memory

#Execute an interactive shell inside the container
$docker exec -it sbi_semantic_memory sh

#Restart the container:
$docker restart sbi_semantic_memory

#Stop, remove, and recreate the container:
$docker compose down && docker compose up -d

#Force-kill a stuck container:
$docker kill sbi_semantic_memory
import os
from dotenv import load_dotenv
from qdrant_client import QdrantClient
from qdrant_client.http import models

# Load environment variables
load_dotenv()

class SemanticMemory:
    def __init__(self):
        host = os.getenv("QDRANT_HOST", "localhost")
        port = int(os.getenv("QDRANT_PORT", 6333))

        self.client = QdrantClient(host=host, port=port)
        self.collection_name = "agent_memories"
        self._ensure_collection()

    def _ensure_collection(self):
        collections = self.client.get_collections().collections
        exists = any(c.name == self.collection_name for c in collections)

        if not exists:
            self.client.create_collection(
                collection_name=self.collection_name,
                vectors_config=models.VectorParams(
                    size=1536,
                    distance=models.Distance.COSINE
                ),
            )

memory_vault = SemanticMemory()
# OpenAI API Key
OPENAI_API_KEY=

# Qdrant Database Configuration
QDRANT_HOST=localhost
QDRANT_PORT=6333
QDRANT_COLLECTION_NAME=agent_memories

# FastAPI Server Setup
HOST=0.0.0.0
PORT=8000
from typing import Dict, Any, List
from dotenv import load_dotenv
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel

# Ensure environment variables are loaded at application start
load_dotenv(override=True)

from graph import app_graph

app = FastAPI(title="Qdrant Semantic Memory Service")

class ChatRequest(BaseModel):
    message: str
    thread_id: str
    metadata: Dict[str, Any] = {}

class ChatResponse(BaseModel):
    thread_id: str
    messages: List[Dict[str, Any]]

@app.post("/chat", response_model=ChatResponse)
async def chat_endpoint(req: ChatRequest):
    try:
        initial_state = {
            "messages": [{"role": "user", "content": req.message}],
            "thread_id": req.thread_id,
            "metadata": req.metadata
        }

        final_state = await app_graph.ainvoke(initial_state)

        return ChatResponse(
            thread_id=req.thread_id,
            messages=final_state["messages"]
        )
    except Exception as e:
        raise HTTPException(status_code=500, detail=str(e))
{
  "message": "My favorite color is Obsidian Blue.",
  "thread_id": "thread-001",
  "metadata": {"source": "mobile_app"}
}
{
  "thread_id": "thread-001",
  "messages": [
    {
      "role": "system",
      "content": "You have access to the following long-term memory facts about the user:\n- My favorite color is Obsidian Blue.\n- My favorite color is Obsidian Blue.\n- Obsidian Blue is a deep, rich shade of blue that often resembles the color of the volcanic glass obsidian. It's a striking and elegant color choice! If you have any questions or need assistance related to colors or anything else, feel free to ask!\n\nUse the facts above to directly answer the user's question."
    },
    {
      "role": "user",
      "content": "My favorite color is Obsidian Blue."
    },
    {
      "role": "assistant",
      "content": "That's a great choice! Obsidian Blue is a deep, rich shade of blue that resembles the color of volcanic glass. It's both striking and elegant. If you have any questions or need assistance related to colors or anything else, feel free to ask!"
    }
  ]
}
{
  "message": "What is my favorite color?",
  "thread_id": "thread-002",
  "metadata": {"source": "web_dashboard" }
}
{
  "thread_id": "thread-002",
  "messages": [
    {
      "role": "system",
      "content": "You have access to the following long-term memory facts about the user:\n- What is my favorite color?\n- My favorite color is Obsidian Blue.\n- My favorite color is Obsidian Blue.\n\nUse the facts above to directly answer the user's question."
    },
    {
      "role": "user",
      "content": "What is my favorite color?"
    },
    {
      "role": "assistant",
      "content": "Your favorite color is Obsidian Blue."
    }
  ]
}
# 1. Generate query vector from user's message
query_vector = embeddings.embed_query(user_query)

# 2. Search Qdrant WITHOUT a payload filter
relevant_docs = memory_vault.client.query_points(
    collection_name="agent_memories",
    query=query_vector,
    limit=2  # Grabs top 2 matches globally across all threads
).points
from qdrant_client.http import models

# Returns matches ONLY if thread_id matches current thread
relevant_docs = memory_vault.client.query_points(
    collection_name="agent_memories",
    query=query_vector,
    query_filter=models.Filter(
        must=[
            models.FieldCondition(
                key="thread_id",
                match=models.MatchValue(value=state["thread_id"])
            )
        ]
    ),
    limit=2
).points

Key Concepts Learned:

For the Key Concepts Learned stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

import os
import uuid
from typing import TypedDict, List, Dict, Any
from dotenv import load_dotenv
from langchain_openai import OpenAIEmbeddings, ChatOpenAI
from qdrant_client.http import models
from langgraph.graph import StateGraph, START, END
from app.memory.vector_store import memory_vault

load_dotenv()

class AgentState(TypedDict):
    messages: List[Dict[str, Any]]
    thread_id: str
    tenant_id: str  # Enforces multi-tenancy boundaries
    user_id: str    # Identifies the end-user
    metadata: Dict[str, Any]

embeddings = OpenAIEmbeddings(model="text-embedding-3-small")
llm = ChatOpenAI(model="gpt-4o", temperature=0)

async def recall_node(state: AgentState) -> dict:
    """Retrieves semantic memories STRICTLY bounded by tenant_id and user_id."""
    user_query = state["messages"][-1]["content"]
    query_vector = embeddings.embed_query(user_query)

    # Multi-Tenant Payload Filter Enforcement
    tenant_filter = models.Filter(
        must=[
            models.FieldCondition(
                key="tenant_id",
                match=models.MatchValue(value=state["tenant_id"])
            ),
            models.FieldCondition(
                key="user_id",
                match=models.MatchValue(value=state["user_id"])
            )
        ]
    )

    # Scoped Query Execution
    relevant_docs = memory_vault.client.query_points(
        collection_name="agent_memories",
        query=query_vector,
        query_filter=tenant_filter,  # Prevents cross-tenant data leaks
        limit=3
    ).points

    if relevant_docs:
        memory_context = "\n".join([f"- {hit.payload['text']}" for hit in relevant_docs])

        system_instruction = {
            "role": "system",
            "content": (
                "You have access to the following long-term memory facts about the user:\n"
                f"{memory_context}\n\n"
                "Use the facts above to directly answer the user's question."
            )
        }
        state["messages"].insert(0, system_instruction)

    return {"messages": state["messages"]}

async def agent_node(state: AgentState) -> dict:
    """LLM reasoning node."""
    response = await llm.ainvoke(state["messages"])
    state["messages"].append({"role": "assistant", "content": response.content})
    return {"messages": state["messages"]}

async def memorize_node(state: AgentState) -> dict:
    """Embeds user facts along with mandatory multi-tenant metadata payloads."""
    user_message = next(
        (m["content"] for m in reversed(state["messages"]) if m.get("role") == "user"),
        None
    )

    if user_message:
        vector = embeddings.embed_query(user_message)

        memory_vault.client.upsert(
            collection_name="agent_memories",
            points=[
                models.PointStruct(
                    id=str(uuid.uuid4()),
                    vector=vector,
                    payload={
                        "text": user_message,
                        "tenant_id": state["tenant_id"],  # Tenant payload tag
                        "user_id": state["user_id"],      # User payload tag
                        "thread_id": state["thread_id"],
                        "metadata": state.get("metadata", {})
                    }
                )
            ]
        )
    return state

# Graph Workflow Wiring
workflow = StateGraph(AgentState)

workflow.add_node("recall", recall_node)
workflow.add_node("agent", agent_node)
workflow.add_node("memorize", memorize_node)

workflow.add_edge(START, "recall")
workflow.add_edge("recall", "agent")
workflow.add_edge("agent", "memorize")
workflow.add_edge("memorize", END)

app_graph = workflow.compile()

Operational checklist

When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.

Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.

Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for e7110284d0c4: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.

For the hardening note 0 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Hardening detail 0/759: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

When working through the hardening note 1 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Hardening detail 1/759: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 2 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Hardening detail 2/759: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.