Home / Articles / Practical notes: Self-improving Agentic AI Applications — Feedback loop

This article is published in English.

Practical notes: Self-improving Agentic AI Applications — Feedback loop

Operable walkthrough of Practical notes: Self-improving Agentic AI Applications — Feedback loop: contracts, checks, and drop-in code slots for teams shipping this pattern.

4517 words

Use this as an operator-facing rebuild of the ideas in “Self-improving Agentic AI Applications — Feedback loop integration”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Agent produces answer
        ↓
LLM reflects on answer
        ↓
Agent learns

The Agent Should Not Be the Learning System

For the The Agent Should Not stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Agent made mistake
        ↓
Agent reflects
        ↓
"Always retrieve state policy"
        ↓
Write lesson to memory
        ↓
Future agents use lesson
Production failure
        ↓
Capture evidence
        ↓
Evaluate the run
        ↓
Identify recurring failure
        ↓
Generate lesson candidate
        ↓
Gather supporting and contradicting evidence
        ↓
Validate
        ↓
Canary test
        ↓
Activate

Three Different Feedback Loops

For the Three Different Feedback Loops stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

+------------------------------------------------+
|                RUNTIME PLANE                   |
|                                                |
| plan -> act -> validate -> repair -> respond   |
+-----------------------+------------------------+
                        |
                        v
+------------------------------------------------+
|                LEARNING PLANE                  |
|                                                |
| evaluate -> diagnose -> cluster -> learn       |
+-----------------------+------------------------+
                        |
                        v
+------------------------------------------------+
|                CONTROL PLANE                   |
|                                                |
| test -> approve -> canary -> rollout -> rollback|
+------------------------------------------------+

1. Runtime Self-Healing

For the 1 Runtime Self-Healing stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the 1 Runtime Self-Healing stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Tool failed
   ↓
Retry
   ↓
Retry
   ↓
Retry
User Request
     |
     v
Clarification Gate
     |
     v
Retrieve Context
     |
     v
Plan
     |
     v
Proposed Action
     |
     v
Action Guard
     |
     v
Execute Tool
     |
     v
Sanitize Tool Result
     |
     v
Validate Tool Result
     |
     v
Reason
     |
     v
Validate Answer
     |
     +------ uncertain ------> Critic
     |                           |
     |                           v
     |                      Policy Router
     |                     /    |    |    \
     |                  PASS REPAIR HUMAN FAIL
     |                           |
     +---------------------------+
     |
     v
Response

Validate Before an Action, Not Only After It

When working through the Validate Before an Action stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Tool Outputs Are Also Untrusted Input

When working through the Tool Outputs Are Also stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

External Tool
     |
     v
Tool Result
     |
     v
Sanitizer
     |
     v
Validator
     |
     v
LLM

Separate the Validator from the Critic

When working through the Separate the Validator from stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Separate the Validator from stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Schema correct?
Required fields present?
Allowed value?
Business invariant satisfied?
Evidence exists?
Policy satisfied?
Known contradiction detected?
{
  "correctness": 0.61,
  "groundedness": 0.92,
  "uncertainty": 0.73,
  "defects": [
    "missing_authoritative_evidence"
  ]
}
PASS
REPAIR
HUMAN REVIEW
SAFE FAIL

Repair Should Change the Strategy

The Repair Should Change the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Search policy
    ↓
Generate recommendation
    ↓
Fail validation
Search policy
    ↓
Generate recommendation
failure signature
strategy fingerprint
attempt ID
repair strategy
remaining budget
quality delta
same failure
+
same strategy
+
same evidence
=
do not retry

Measure Whether Repair Actually Helped

The Measure Whether Repair Actually stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

quality score = 0.54
quality score = 0.55
quality score = 0.56
delta = score_after_repair - score_before_repair
small delta
+
small delta
=
human review or safe failure

Human-in-the-Loop Is a Routing Strategy, Not an Exception

The Human-in-the-Loop Is a Routing stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Human-in-the-Loop Is a Routing stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Human
                   |
       +-----------+-----------+
       |           |           |
    Approve       Edit       Reject
       |           |           |
     Continue    Validate    Safe Fail
              Request Repair
                    |
                    v
                  Repair

2. Experience Learning Across Runs

For the 2 Experience Learning Across stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

run_completed
      |
      v
Evaluation worker
      |
      v
Gather feedback
      |
      v
Reflection
      |
      v
Candidate lesson

Feedback Is Not Truth

For the Feedback Is Not Truth stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

thumbs up != correct
thumbs down != incorrectuser correction != authoritative rule
Authoritative business outcome
          >
Expert human label
          >
Deterministic rule
          >
Calibrated evaluator
          >
User feedback
          >
Agent self-confidence

Some of the Best Feedback Arrives Later

For the Some of the Best stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Some of the Best stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Agent recommends payroll code
        |
        v
Payroll system accepts
        |
        v
Two weeks later
        |
        v
Audit rejects transaction

Memory Needs Governance

When working through the Memory Needs Governance stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Profile memory

When working through the Profile memory stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Episodic memory

When working through the Episodic memory stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Episodic memory stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Procedural memory

The Procedural memory stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Observation
     ↓
Candidate lesson
     ↓
Supporting evidence
     +
Counterexamples
     ↓
Validation
     ↓
Active lesson

Learned Memory Must Never Override Authoritative Knowledge

The Learned Memory Must Never stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Authoritative policy
        >
Tenant configuration
        >
Approved procedural lesson
        >
Episodic example
        >
User preference
        >
Unverified claim
        >
LLM reflection

Memory Can Become Stale

The Memory Can Become Stale stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Memory Can Become Stale stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

candidate
   ↓
validated
   ↓
active
   ↓
pending revalidation
   ↓
deprecated
   ↓
retired

3. The Offline Learning Plane

For the 3 The Offline Learning stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Runs
+
User Feedback
+
Human Reviews
+
Downstream Outcomes
+
Evaluation Scores
        |
        v
Failure Classification
        |
        v
Failure Clustering
missing state policy        178
wrong tool selected          63
bad tool argument            52
unsupported inference        41
output schema failure        11

Not Every Failure Is a Prompt Problem

For the Not Every Failure Is stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

Graph rule:
Policy retrieval must occur before this decision.
retrieval
tool schema
validator rule
routing
memory policy
clarification logic
action guard
model configuration

Evaluation Is the Heart of the Feedback Loop

For the Evaluation Is the Heart stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Evaluation Is the Heart stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Correctness
Groundedness
Retrieval relevance
Completeness
Tool selection
Tool arguments
Trajectory efficiency
Repair effectiveness
Safety
Latency
Cost
Business outcome
Agent A
search -> answer
Agent B
search
-> wrong tool
-> retry
-> timeout
-> second search
-> repair
-> answer

Do Not Trust One LLM Judge

When working through the Do Not Trust One stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

Deterministic checks
+
Business outcomes
+
Human labels
+
Multiple evaluator rubrics
+
LLM judges

Observability Is Part of the Learning Architecture

When working through the Observability Is Part of stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Agent Run
 |
 +-- Retrieve Context
 |
 +-- Planner
 |
 +-- Tool Call
 |
 +-- Tool Result Validator
 |
 +-- Repair
 |
 +-- Critic
 |
 +-- Final Answer
LangGraph execution
MongoDB events
Vertex AI calls
Evaluation results
Human feedback
Downstream outcomes

Why Append-Only Agent Events Matter

When working through the Why Append-Only Agent Events stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Why Append-Only Agent Events stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

run_id
agent
release
outcome
latency
tool count
repair count
cost
run.started
retrieval.completed
tool.called
validation.failed
repair.started
human_review.requested
run.completed

Reproducibility Requires More Than a Prompt Version

The Reproducibility Requires More Than stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

model
generation settings
graph version
tool definitions
validator rules
critic rubric
retrieval configuration
memory rules
security rules
agent_release_43
    |
    +-- graph v12
    +-- prompt v43
    +-- Gemini configuration
    +-- tools v17
    +-- validator v11
    +-- retrieval config v9
    +-- memory policy v5
    +-- evaluation suite v8

The Control Plane: Where Improvements Earn Production Access

The The Control Plane Where stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Candidate Improvement
        |
        v
Historical Replay
        |
        v
Regression Evaluation
        |
        v
Shadow Production
        |
        v
Canary
        |
        v
Progressive Rollout
        |
        v
Production

Lessons Need Canary Releases Too

The Lessons Need Canary Releases stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Lessons Need Canary Releases stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Historical replay
     ↓
5% canary
     ↓
Measure outcome
     ↓
25%
     ↓
Measure
     ↓
100%

The Final Architecture

For the The Final Architecture stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

USER
                          |
                          v
+------------------------------------------------------+
|                    RUNTIME                           |
|                                                      |
| clarify -> retrieve -> plan -> action guard          |
|                            |                         |
|                            v                         |
|                           tool                       |
|                            |                         |
|                       sanitize                       |
|                            |                         |
|                      validate                        |
|                            |                         |
|                         reason                       |
|                            |                         |
|                    answer validate                   |
|                            |                         |
|                   critic if needed                   |
|                            |                         |
|                    policy router                     |
|                 /       |       \                    |
|              repair   human     pass                 |
+--------------------------+---------------------------+
                           |
                      agent events
                           |
                           v
+------------------------------------------------------+
|                    LEARNING                          |
|                                                      |
| traces + feedback + outcomes                         |
|            |                                         |
|            v                                         |
|         evaluate                                     |
|            |                                         |
|     classify failures                                |
|            |                                         |
|         cluster                                      |
|        /       \                                     |
|    lessons    improvement candidates                 |
+--------+----------------------+----------------------+
         |                      |
         v                      v
+------------------------------------------------------+
|                    CONTROL                           |
|                                                      |
| validate -> replay -> shadow -> canary -> rollout    |
|                                |                     |
|                              monitor                 |
|                                |                     |
|                             rollback                 |
+------------------------------------------------------+

What This Changes About “Self-Improving AI”

For the What This Changes About stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Observe
   ↓
Measure
   ↓
Diagnose
   ↓
Propose
   ↓
Test
   ↓
Promote
   ↓
Monitor

A Practical Implementation Sequence

For the A Practical Implementation Sequence stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the A Practical Implementation Sequence stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Reliability
   ↓
Observability
   ↓
Evaluation
   ↓
Learning
   ↓
Controlled adaptation

Closing Thought

When working through the Closing Thought stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

production behavior
        ↓
evidence
        ↓
evaluation
        ↓
learning
        ↓
experimentation
        ↓
controlled production change

Operational checklist

When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.

Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.

Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for da99b44a5b86: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.