This article is published in English.
Practical notes: Self-improving Agentic AI Applications — Feedback loop
Operable walkthrough of Practical notes: Self-improving Agentic AI Applications — Feedback loop: contracts, checks, and drop-in code slots for teams shipping this pattern.
Use this as an operator-facing rebuild of the ideas in “Self-improving Agentic AI Applications — Feedback loop integration”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Agent produces answer
↓
LLM reflects on answer
↓
Agent learns
The Agent Should Not Be the Learning System
For the The Agent Should Not stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Agent made mistake
↓
Agent reflects
↓
"Always retrieve state policy"
↓
Write lesson to memory
↓
Future agents use lesson
Production failure
↓
Capture evidence
↓
Evaluate the run
↓
Identify recurring failure
↓
Generate lesson candidate
↓
Gather supporting and contradicting evidence
↓
Validate
↓
Canary test
↓
Activate
Three Different Feedback Loops
For the Three Different Feedback Loops stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
+------------------------------------------------+
| RUNTIME PLANE |
| |
| plan -> act -> validate -> repair -> respond |
+-----------------------+------------------------+
|
v
+------------------------------------------------+
| LEARNING PLANE |
| |
| evaluate -> diagnose -> cluster -> learn |
+-----------------------+------------------------+
|
v
+------------------------------------------------+
| CONTROL PLANE |
| |
| test -> approve -> canary -> rollout -> rollback|
+------------------------------------------------+
1. Runtime Self-Healing
For the 1 Runtime Self-Healing stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the 1 Runtime Self-Healing stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Tool failed
↓
Retry
↓
Retry
↓
Retry
User Request
|
v
Clarification Gate
|
v
Retrieve Context
|
v
Plan
|
v
Proposed Action
|
v
Action Guard
|
v
Execute Tool
|
v
Sanitize Tool Result
|
v
Validate Tool Result
|
v
Reason
|
v
Validate Answer
|
+------ uncertain ------> Critic
| |
| v
| Policy Router
| / | | \
| PASS REPAIR HUMAN FAIL
| |
+---------------------------+
|
v
Response
Validate Before an Action, Not Only After It
When working through the Validate Before an Action stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Tool Outputs Are Also Untrusted Input
When working through the Tool Outputs Are Also stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
External Tool
|
v
Tool Result
|
v
Sanitizer
|
v
Validator
|
v
LLM
Separate the Validator from the Critic
When working through the Separate the Validator from stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Separate the Validator from stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Schema correct?
Required fields present?
Allowed value?
Business invariant satisfied?
Evidence exists?
Policy satisfied?
Known contradiction detected?
{
"correctness": 0.61,
"groundedness": 0.92,
"uncertainty": 0.73,
"defects": [
"missing_authoritative_evidence"
]
}
PASS
REPAIR
HUMAN REVIEW
SAFE FAIL
Repair Should Change the Strategy
The Repair Should Change the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Search policy
↓
Generate recommendation
↓
Fail validation
Search policy
↓
Generate recommendation
failure signature
strategy fingerprint
attempt ID
repair strategy
remaining budget
quality delta
same failure
+
same strategy
+
same evidence
=
do not retry
Measure Whether Repair Actually Helped
The Measure Whether Repair Actually stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
quality score = 0.54
quality score = 0.55
quality score = 0.56
delta = score_after_repair - score_before_repair
small delta
+
small delta
=
human review or safe failure
Human-in-the-Loop Is a Routing Strategy, Not an Exception
The Human-in-the-Loop Is a Routing stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Human-in-the-Loop Is a Routing stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Human
|
+-----------+-----------+
| | |
Approve Edit Reject
| | |
Continue Validate Safe Fail
Request Repair
|
v
Repair
2. Experience Learning Across Runs
For the 2 Experience Learning Across stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
run_completed
|
v
Evaluation worker
|
v
Gather feedback
|
v
Reflection
|
v
Candidate lesson
Feedback Is Not Truth
For the Feedback Is Not Truth stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
thumbs up != correct
thumbs down != incorrectuser correction != authoritative rule
Authoritative business outcome
>
Expert human label
>
Deterministic rule
>
Calibrated evaluator
>
User feedback
>
Agent self-confidence
Some of the Best Feedback Arrives Later
For the Some of the Best stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Some of the Best stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Agent recommends payroll code
|
v
Payroll system accepts
|
v
Two weeks later
|
v
Audit rejects transaction
Memory Needs Governance
When working through the Memory Needs Governance stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Profile memory
When working through the Profile memory stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Episodic memory
When working through the Episodic memory stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Episodic memory stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Procedural memory
The Procedural memory stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Observation
↓
Candidate lesson
↓
Supporting evidence
+
Counterexamples
↓
Validation
↓
Active lesson
Learned Memory Must Never Override Authoritative Knowledge
The Learned Memory Must Never stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Authoritative policy
>
Tenant configuration
>
Approved procedural lesson
>
Episodic example
>
User preference
>
Unverified claim
>
LLM reflection
Memory Can Become Stale
The Memory Can Become Stale stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Memory Can Become Stale stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
candidate
↓
validated
↓
active
↓
pending revalidation
↓
deprecated
↓
retired
3. The Offline Learning Plane
For the 3 The Offline Learning stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Runs
+
User Feedback
+
Human Reviews
+
Downstream Outcomes
+
Evaluation Scores
|
v
Failure Classification
|
v
Failure Clustering
missing state policy 178
wrong tool selected 63
bad tool argument 52
unsupported inference 41
output schema failure 11
Not Every Failure Is a Prompt Problem
For the Not Every Failure Is stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
Graph rule:
Policy retrieval must occur before this decision.
retrieval
tool schema
validator rule
routing
memory policy
clarification logic
action guard
model configuration
Evaluation Is the Heart of the Feedback Loop
For the Evaluation Is the Heart stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Evaluation Is the Heart stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Correctness
Groundedness
Retrieval relevance
Completeness
Tool selection
Tool arguments
Trajectory efficiency
Repair effectiveness
Safety
Latency
Cost
Business outcome
Agent A
search -> answer
Agent B
search
-> wrong tool
-> retry
-> timeout
-> second search
-> repair
-> answer
Do Not Trust One LLM Judge
When working through the Do Not Trust One stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
Deterministic checks
+
Business outcomes
+
Human labels
+
Multiple evaluator rubrics
+
LLM judges
Observability Is Part of the Learning Architecture
When working through the Observability Is Part of stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Agent Run
|
+-- Retrieve Context
|
+-- Planner
|
+-- Tool Call
|
+-- Tool Result Validator
|
+-- Repair
|
+-- Critic
|
+-- Final Answer
LangGraph execution
MongoDB events
Vertex AI calls
Evaluation results
Human feedback
Downstream outcomes
Why Append-Only Agent Events Matter
When working through the Why Append-Only Agent Events stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Why Append-Only Agent Events stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
run_id
agent
release
outcome
latency
tool count
repair count
cost
run.started
retrieval.completed
tool.called
validation.failed
repair.started
human_review.requested
run.completed
Reproducibility Requires More Than a Prompt Version
The Reproducibility Requires More Than stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
model
generation settings
graph version
tool definitions
validator rules
critic rubric
retrieval configuration
memory rules
security rules
agent_release_43
|
+-- graph v12
+-- prompt v43
+-- Gemini configuration
+-- tools v17
+-- validator v11
+-- retrieval config v9
+-- memory policy v5
+-- evaluation suite v8
The Control Plane: Where Improvements Earn Production Access
The The Control Plane Where stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Candidate Improvement
|
v
Historical Replay
|
v
Regression Evaluation
|
v
Shadow Production
|
v
Canary
|
v
Progressive Rollout
|
v
Production
Lessons Need Canary Releases Too
The Lessons Need Canary Releases stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Lessons Need Canary Releases stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Historical replay
↓
5% canary
↓
Measure outcome
↓
25%
↓
Measure
↓
100%
The Final Architecture
For the The Final Architecture stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
USER
|
v
+------------------------------------------------------+
| RUNTIME |
| |
| clarify -> retrieve -> plan -> action guard |
| | |
| v |
| tool |
| | |
| sanitize |
| | |
| validate |
| | |
| reason |
| | |
| answer validate |
| | |
| critic if needed |
| | |
| policy router |
| / | \ |
| repair human pass |
+--------------------------+---------------------------+
|
agent events
|
v
+------------------------------------------------------+
| LEARNING |
| |
| traces + feedback + outcomes |
| | |
| v |
| evaluate |
| | |
| classify failures |
| | |
| cluster |
| / \ |
| lessons improvement candidates |
+--------+----------------------+----------------------+
| |
v v
+------------------------------------------------------+
| CONTROL |
| |
| validate -> replay -> shadow -> canary -> rollout |
| | |
| monitor |
| | |
| rollback |
+------------------------------------------------------+
What This Changes About “Self-Improving AI”
For the What This Changes About stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Observe
↓
Measure
↓
Diagnose
↓
Propose
↓
Test
↓
Promote
↓
Monitor
A Practical Implementation Sequence
For the A Practical Implementation Sequence stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the A Practical Implementation Sequence stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Reliability
↓
Observability
↓
Evaluation
↓
Learning
↓
Controlled adaptation
Closing Thought
When working through the Closing Thought stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
production behavior
↓
evidence
↓
evaluation
↓
learning
↓
experimentation
↓
controlled production change
Operational checklist
When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.
Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.
Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for da99b44a5b86: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.