This article is published in English.
Practical notes: Building Effective AI Agents: Architecture Patterns and
Operable walkthrough of Practical notes: Building Effective AI Agents: Architecture Patterns and: contracts, checks, and drop-in code slots for teams shipping this pattern.
This walkthrough rebuilds the path from raw materials to a working system for: Building Effective AI Agents: Architecture Patterns and Implementation Frameworks. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
1. What Is an AI Agent?
When working through the 1 What Is an stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
2. The Core Agent Architecture
When working through the 2 The Core Agent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
User or System Interface
When working through the User or System Interface stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the User or System Interface stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Agent Runtime
The Agent Runtime stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
3. The Model Layer
The 3 The Model Layer stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
4. Tools Turn Models into Agents
The 4 Tools Turn Models stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices. The 4 Tools Turn Models stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
search_leads()
get_account()
get_customer_history()
get_recent_emails()
create_opportunity()
schedule_meeting()
generate_proposal()
5. Pattern 1: The Tool-Using Agent
For the 5 Pattern 1 The stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.
User
↓
Agent
↓
LLM
↓
Choose Tool
↓
Execute Tool
↓
Tool Result
↓
LLM
↓
Final Response
get_weather("Chicago", "tomorrow")
6. Pattern 2: ReAct
For the 6 Pattern 2 ReAct stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Goal
↓
Reason
↓
Action
↓
Observation
↓
Reason
↓
Action
↓
Observation
↓
Final Answer
User Goal:
Find our highest-value customer with an unresolved support case.
Reason:
I need customer revenue data.Action:
query_customer_database()Observation:
Customer A has the highest revenue.Reason:
Now I need unresolved support cases.Action:
search_support_cases(customer_a)Observation:
Two unresolved cases found.Final Answer:
Customer A is the highest-value customer currently
associated with unresolved support cases.
7. Pattern 3: Plan-and-Execute
For the 7 Pattern 3 Plan-and-Execute stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the 7 Pattern 3 Plan-and-Execute stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Goal:
Prepare me for tomorrow's meeting with Acme Corp.
1. Retrieve account information
2. Review recent opportunities
3. Retrieve previous meeting notes
4. Review recent emails
5. Identify unresolved issues
6. Find relevant company news
7. Generate meeting briefing
User Goal
↓
Planner
↓
Task Plan
↓
Executor
↓
Tools
↓
Results
↓
Evaluator
↓
Final Output
8. Pattern 4: Router Architecture
When working through the 8 Pattern 4 Router stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
┌── HR Agent
│
User → Router ────┼── Finance Agent
│
├── IT Support Agent
│
└── Sales Agent
9. Pattern 5: Supervisor and Worker Agents
When working through the 9 Pattern 5 Supervisor stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Research Agent
↑
│
User → Supervisor → Data Agent
│
↓
Report Agent
10. Pattern 6: Agentic RAG
When working through the 10 Pattern 6 Agentic stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the 10 Pattern 6 Agentic stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Question
↓
Vector Search
↓
Relevant Documents
↓
LLM
↓
Answer
Question
↓
Agent
↓
Do I need retrieval?
↓
Which source?
↓
Search
↓
Evaluate results
↓
Enough information?
↓ ↓
Yes No
↓ ↓
Answer Search Again
11. Pattern 7: Reflection and Self-Evaluation
The 11 Pattern 7 Reflection stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Generate Answer
↓
Evaluate Answer
↓
Is it sufficient?
↓ ↓
Yes No
↓ ↓
Return Improve
12. Memory Architecture
The 12 Memory Architecture stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Working Memory
The Working Memory stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Working Memory stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Conversation Memory
For the Conversation Memory stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Long-Term Memory
For the Long-Term Memory stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Episodic Memory
For the Episodic Memory stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Episodic Memory stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
13. State Is Often More Important Than Memory
When working through the 13 State Is Often stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
state = {
"goal": "",
"user_id": "",
"plan": [],
"completed_tasks": [],
"tool_results": {},
"approval_status": None,
"errors": [],
"final_answer": None
}
14. Graph-Based Agent Architectures
When working through the 14 Graph-Based Agent Architectures stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Start
↓
Classify Request
↓
Retrieve Data
↓
Analyze
↓
Risk Check
↓
Need Approval?
↓ ↓
Yes No
↓ ↓
Human Execute
Approval ↓
↓ Finish
Execute
↓
Finish
15. Human-in-the-Loop Architecture
When working through the 15 Human-in-the-Loop Architecture stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the 15 Human-in-the-Loop Architecture stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Agent Recommendation
↓
Policy Check
↓
High-Risk Action?
↓ ↓
Yes No
↓ ↓
Human Execute
Approval
↓
Execute
16. Agent Guardrails
The 16 Agent Guardrails stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Input Guardrails
The Input Guardrails stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Tool Guardrails
The Tool Guardrails stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve. The Tool Guardrails stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Output Guardrails
For the Output Guardrails stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Execution Guardrails
For the Execution Guardrails stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
17. Implementation Frameworks
For the 17 Implementation Frameworks stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the 17 Implementation Frameworks stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Python
+
LLM API or local model
+
Functions
+
FastAPI
+
Database
Graph-Based Frameworks
When working through the Graph-Based Frameworks stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Multi-Agent Frameworks
When working through the Multi-Agent Frameworks stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Enterprise AI Orchestration Platforms
When working through the Enterprise AI Orchestration Platforms stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Enterprise AI Orchestration Platforms stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
18. A Practical Implementation Framework
The 18 A Practical Implementation stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Step 1: Define the Goal
The Step 1 Define the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Step 2: Define the Agent’s Responsibilities
The Step 2 Define the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Step 2 Define the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Retrieve account information
Retrieve opportunities
Analyze customer communication
Retrieve open support issues
Generate meeting briefing
Recommend discussion topics
Cannot modify CRM records
Cannot send email
Cannot change pricing
Cannot create contracts
19. Step 3: Define the Tools
For the 19 Step 3 Define stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.
get_account(account_id)
get_opportunities(account_id)
get_support_cases(account_id)
get_email_history(account_id)
search_company_news(company_name)
manage_customer()
get_customer_profile()
20. Step 4: Design the Control Flow
For the 20 Step 4 Design stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
21. Step 5: Add State and Memory
For the 21 Step 5 Add stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the 21 Step 5 Add stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Conversation state
Task state
User preferences
Long-term knowledge
Audit history
22. Step 6: Add Observability
When working through the 22 Step 6 Add stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Request
Agent decision
Model used
Prompt version
Tool selected
Tool input
Tool output
Execution time
Token usage
Cost
Errors
Retries
Final response
Human overrides
Request ID: 78425
Step 1:
Intent → Customer Meeting PreparationStep 2:
Tool → get_account()Step 3:
Tool → get_opportunities()Step 4:
Tool → search_support_cases()Step 5:
LLM → Generate briefingTotal execution: 4.8 seconds
Tool calls: 3
Model calls: 2
23. Step 7: Evaluate the Agent
When working through the 23 Step 7 Evaluate stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Task Completion
When working through the Task Completion stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Task Completion stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Tool Selection
The Tool Selection stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Tool Accuracy
Groundedness
Safety
Efficiency
Latency
Cost
24. Failure Recovery
Tool Call
↓
Success?
↓ ↓
Yes No
↓ ↓
Continue Retry
↓
Still Fails?
↓ ↓
Yes No
↓ ↓
Fallback Continue
↓
Escalate
25. Multi-Agent Systems: Use Them Carefully
Supervisor
├── Financial Analysis Agent
├── Legal Analysis Agent
├── Market Research Agent
└── Report Generation Agent
Search Agent
Thinking Agent
Tool Agent
Summary Agent
Response Agent
26. The Enterprise Agent Architecture
User
↓
Agent Gateway
↓
Identity / Access
↓
Router
↓
┌─────────────┼─────────────┐
↓ ↓ ↓
Sales Agent HR Agent Finance Agent
↓ ↓ ↓
Agent Runtime
↓
┌──────────┼──────────┐
↓ ↓ ↓
RAG Tools Memory
↓ ↓ ↓
Knowledge APIs Databases
Base
↓
Policy Engine
↓
Human Approval Layer
↓
Observability
↓
Evaluation
27. Deterministic Software and Probabilistic AI
Probabilistic Intelligence
+
Deterministic Control
=
Reliable Agentic System
28. Start with Workflows, Then Add Autonomy
Stage 1 — Assistant
Stage 2 — Tool-Using Assistant
Stage 3 — Guided Agent
Stage 4 — Autonomous Workflow
Stage 5 — Multi-Agent System
29. What Makes an Effective AI Agent?
30. Final Thoughts
Architecture Summary
AI Agent
│
├── Model
├── Instructions
├── Context
├── Tools
├── Retrieval
├── Memory
├── State
├── Planning
├── Orchestration
├── Guardrails
├── Human Oversight
├── Observability
└── Evaluation