Home / Articles / Practical notes: Give Your AI Agent a Memory — Then Watch Everything It Does

This article is published in English.

Practical notes: Give Your AI Agent a Memory — Then Watch Everything It Does

Operable walkthrough of Practical notes: Give Your AI Agent a Memory — Then Watch Everything It Does: contracts, checks, and drop-in code slots for teams shipping this pattern.

3610 words

This walkthrough rebuilds the path from raw materials to a working system for: Give Your AI Agent a Memory — Then Watch the author It Does with LangSmith. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

The Architecture

When working through the The Architecture stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

                    ┌─────────────────────┐
                    │       User          │
                    └──────────┬──────────┘
                               │
                               ▼
                    ┌─────────────────────┐
                    │    AI Agent         │
                    │  LangChain Agent    │
                    └──────────┬──────────┘
                               │
                 ┌─────────────┴─────────────┐
                 │                           │
                 ▼                           ▼
        ┌─────────────────┐        ┌──────────────────┐
        │ Conversation    │        │     Tools        │
        │ Memory          │        │                  │
        │ InMemorySaver   │        │ Tavily Web Search│
        └─────────────────┘        └──────────────────┘
                 │                           │
                 └─────────────┬─────────────┘
                               │
                               ▼
                    ┌─────────────────────┐
                    │     Ollama          │
                    │     Llama 3.2       │
                    └─────────────────────┘
                               │
                               ▼
                    ┌─────────────────────┐
                    │     LangSmith       │
                    │  Traces & Debugging │
                    └─────────────────────┘

1. Running an LLM Locally with Ollama

When working through the 1 Running an LLM stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.

from langchain_ollama import ChatOllama
llm = ChatOllama(
    model="llama3.2:latest",
    base_url="http://localhost:11434",
    temperature=0,
)

model

When working through the model stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

model="llama3.2:latest"

When working through the model stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

base_url

The baseurl stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

base_url="http://localhost:11434"

temperature

The temperature stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

temperature=0

2. Creating the Agent

The 2 Creating the Agent stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The 2 Creating the Agent stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

from langchain.agents import create_agent
agent = create_agent(
    model=llm
)
agent = create_agent(
    model=llm,
    tools=[tool1],
    system_prompt="You are an intelligent knowlegable agent"
)

3. Giving the Agent a Web Search Tool

For the 3 Giving the Agent stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

from langchain.tools import tool
@tool("web_search", description="Search the web for information")
def tool1(query: str) -> Dict[str, Any]:
    tavily = TavilyClient()
    return tavily.search(query=query)
tavily = TavilyClient()
return tavily.search(query=query)
User Question
      │
      ▼
     LLM
      │
      │ "I need external information"
      ▼
  Web Search Tool
      │
      ▼
   Tavily API
      │
      ▼
 Search Results
      │
      ▼
     LLM
      │
      ▼
 Final Answer

4. The Problem: LLMs Don’t Automatically Remember the author

For the 4 The Problem LLMs stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

from langgraph.checkpoint.memory import InMemorySaver
agent = create_agent(
    model=llm,
    checkpointer=InMemorySaver()
)
Question → LLM → Answer

5. InMemorySaver — Giving the Agent a Memory

For the 5 InMemorySaver Giving the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the 5 InMemorySaver Giving the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

checkpointer=InMemorySaver()

6. The Most Important Line: thread_id

When working through the 6 The Most Important stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

config = {
    "configurable": {
        "thread_id": "1"
    }
}
Thread 1
────────────
User → My favorite color is Green
AI   → Great!
User → What's my favorite color?
AI   → Green
Thread 2
────────────
User → My favorite color is Blue
AI   → Blue

7. First Conversation

When working through the 7 First Conversation stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

question = HumanMessage(
    content="I am X my favorite color is Green"
)
response = agent.invoke(
    {"messages": [question]},
    config,
)
"thread_id": "1"

8. Asking the Agent Later

When working through the 8 Asking the Agent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the 8 Asking the Agent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

question2 = HumanMessage(
    content="What's my favorite color?"
)
response2 = agent.invoke(
    {"messages": [question2]},
    config,
)
config = {
    "configurable": {
        "thread_id": "1"
    }
}
Your favorite color is Green!

9. Why thread_id Matters So Much

The 9 Why threadid Matters stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Customer A → thread_id = "customer-A"
Customer B → thread_id = "customer-B"
Customer C → thread_id = "customer-C"
AI Agent
                │
       ┌────────┼────────┐
       ▼        ▼        ▼
   Customer A Customer B Customer C
      │           │          │
   Thread A    Thread B   Thread C

10. But Memory Isn’t Enough

The 10 But Memory Isn stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

User Question
     ↓
Agent
     ↓
LLM decides to call web_search
     ↓
Tavily
     ↓
Search results
     ↓
LLM
     ↓
Final answer

11. Enter LangSmith

The 11 Enter LangSmith stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The 11 Enter LangSmith stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

os.environ["LANGSMITH_TRACING"] = "true"
os.environ["LANGSMITH_API_KEY"] = "YYY"
os.environ["LANGSMITH_ENDPOINT"] = "https://langsmith-endpoint"
os.environ["LANGSMITH_PROJECT"] = "local-ollama-agent"
LANGSMITH_TRACING=true

12. Why Tracing Is Better Than print()

For the 12 Why Tracing Is stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

print(response)
Question
   │
   ├── LLM call
   │
   ├── Tool decision
   │
   ├── Web search
   │
   ├── Tool response
   │
   ├── Another LLM call
   │
   └── Final response

13. What Can We Learn From a Trace?

For the 13 What Can We stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Total execution: 8 seconds
LLM call              1.5 sec
Web search             5.2 sec
Final LLM call         1.3 sec

14. Putting Memory and Observability Together

For the 14 Putting Memory and stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the 14 Putting Memory and stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

                         ┌───────────────┐
                         │     User      │
                         └───────┬───────┘
                                 │
                                 ▼
                      ┌────────────────────┐
                      │     AI Agent       │
                      └─────────┬──────────┘
                                │
                 ┌──────────────┼──────────────┐
                 │              │              │
                 ▼              ▼              ▼
             Memory          LLM            Tools
          InMemorySaver     Ollama          Tavily
                 │              │              │
                 └──────────────┼──────────────┘
                                │
                                ▼
                         ┌─────────────┐
                         │ LangSmith   │
                         │   Tracing   │
                         └─────────────┘

Memory

When working through the Memory stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

LLM

When working through the LLM stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

Tools

When working through the Tools stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours. When working through the Tools stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

LangSmith

The LangSmith stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

15. The Complete Conceptual Flow

The 15 The Complete Conceptual stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Step 1 — User provides information

The Step 1 User provides stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

"My name is Amit and my favorite color is Green."

The Step 1 User provides stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

thread_id = 1

Step 2 — User asks another question

For the Step 2 User asks stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

"What's my favorite color?"
"Your favorite color is Green."

Step 3 — User asks a knowledge question

For the Step 3 User asks stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

"Who is the Chief Minister of Tamil Nadu?"
web_search

Step 4 — Tool executes

For the Step 4 Tool executes stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the Step 4 Tool executes stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Step 5 — LLM generates the answer

When working through the Step 5 LLM generates stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.

Step 6 — LangSmith records the execution

16. A Note About Secrets

os.environ["TAVILY_API_KEY"] = "XXX"
os.environ["LANGSMITH_API_KEY"] = "YYY"
export TAVILY_API_KEY="..."
export LANGSMITH_API_KEY="..."
export LANGSMITH_TRACING="true"
export LANGSMITH_PROJECT="local-ollama-agent"

17. Why This Architecture Matters

User → LLM → Response
┌── Memory
                   │
User → Agent → LLM ├── Tools
                   │
                   └── State
                        │
                        ▼
                    Tracing

18. Memory vs. Persistence

InMemorySaver()
Application starts
      ↓
Thread 1 created
      ↓
Conversation stored in memory
      ↓
Application restarts
      ↓
Memory is gone

19. A Simple Mental Model

1⃣ Agent

agent = create_agent(...)

2⃣ Model

llm = ChatOllama(...)

3⃣ Memory

checkpointer=InMemorySaver()

4⃣ Observability

LANGSMITH_TRACING=true
Agent
 ├── Model
 ├── Memory
 ├── Tools
 └── Observability

Conclusion

Ollama
   +
Llama 3.2
   +
LangChain Agent
   +
InMemorySaver
   +
Tavily
   +
LangSmith

What’s Next?

Source Code

Operational checklist