Home / Articles / Practical notes: Building Effective AI Agents: Architecture Patterns and

This article is published in English.

Practical notes: Building Effective AI Agents: Architecture Patterns and

Operable walkthrough of Practical notes: Building Effective AI Agents: Architecture Patterns and: contracts, checks, and drop-in code slots for teams shipping this pattern.

4803 words

This walkthrough rebuilds the path from raw materials to a working system for: Building Effective AI Agents: Architecture Patterns and Implementation Frameworks. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

1. What Is an AI Agent?

When working through the 1 What Is an stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

2. The Core Agent Architecture

When working through the 2 The Core Agent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

User or System Interface

When working through the User or System Interface stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the User or System Interface stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Agent Runtime

The Agent Runtime stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

3. The Model Layer

The 3 The Model Layer stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

4. Tools Turn Models into Agents

The 4 Tools Turn Models stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices. The 4 Tools Turn Models stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

search_leads()
get_account()
get_customer_history()
get_recent_emails()
create_opportunity()
schedule_meeting()
generate_proposal()

5. Pattern 1: The Tool-Using Agent

For the 5 Pattern 1 The stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

User
  ↓
Agent
  ↓
LLM
  ↓
Choose Tool
  ↓
Execute Tool
  ↓
Tool Result
  ↓
LLM
  ↓
Final Response
get_weather("Chicago", "tomorrow")

6. Pattern 2: ReAct

For the 6 Pattern 2 ReAct stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Goal
 ↓
Reason
 ↓
Action
 ↓
Observation
 ↓
Reason
 ↓
Action
 ↓
Observation
 ↓
Final Answer
User Goal:
Find our highest-value customer with an unresolved support case.
Reason:
I need customer revenue data.Action:
query_customer_database()Observation:
Customer A has the highest revenue.Reason:
Now I need unresolved support cases.Action:
search_support_cases(customer_a)Observation:
Two unresolved cases found.Final Answer:
Customer A is the highest-value customer currently
associated with unresolved support cases.

7. Pattern 3: Plan-and-Execute

For the 7 Pattern 3 Plan-and-Execute stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the 7 Pattern 3 Plan-and-Execute stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Goal:
Prepare me for tomorrow's meeting with Acme Corp.
1. Retrieve account information
2. Review recent opportunities
3. Retrieve previous meeting notes
4. Review recent emails
5. Identify unresolved issues
6. Find relevant company news
7. Generate meeting briefing
User Goal
   ↓
Planner
   ↓
Task Plan
   ↓
Executor
   ↓
Tools
   ↓
Results
   ↓
Evaluator
   ↓
Final Output

8. Pattern 4: Router Architecture

When working through the 8 Pattern 4 Router stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

┌── HR Agent
                  │
User → Router ────┼── Finance Agent
                  │
                  ├── IT Support Agent
                  │
                  └── Sales Agent

9. Pattern 5: Supervisor and Worker Agents

When working through the 9 Pattern 5 Supervisor stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Research Agent
                     ↑
                     │
User → Supervisor → Data Agent
                     │
                     ↓
                 Report Agent

10. Pattern 6: Agentic RAG

When working through the 10 Pattern 6 Agentic stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the 10 Pattern 6 Agentic stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Question
 ↓
Vector Search
 ↓
Relevant Documents
 ↓
LLM
 ↓
Answer
Question
 ↓
Agent
 ↓
Do I need retrieval?
 ↓
Which source?
 ↓
Search
 ↓
Evaluate results
 ↓
Enough information?
   ↓        ↓
  Yes       No
   ↓         ↓
Answer    Search Again

11. Pattern 7: Reflection and Self-Evaluation

The 11 Pattern 7 Reflection stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Generate Answer
     ↓
Evaluate Answer
     ↓
Is it sufficient?
  ↓          ↓
 Yes         No
  ↓           ↓
Return       Improve

12. Memory Architecture

The 12 Memory Architecture stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Working Memory

The Working Memory stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Working Memory stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Conversation Memory

For the Conversation Memory stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Long-Term Memory

For the Long-Term Memory stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Episodic Memory

For the Episodic Memory stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Episodic Memory stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

13. State Is Often More Important Than Memory

When working through the 13 State Is Often stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

state = {
    "goal": "",
    "user_id": "",
    "plan": [],
    "completed_tasks": [],
    "tool_results": {},
    "approval_status": None,
    "errors": [],
    "final_answer": None
}

14. Graph-Based Agent Architectures

When working through the 14 Graph-Based Agent Architectures stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Start
  ↓
Classify Request
  ↓
Retrieve Data
  ↓
Analyze
  ↓
Risk Check
  ↓
Need Approval?
 ↓          ↓
Yes         No
 ↓           ↓
Human       Execute
Approval      ↓
 ↓          Finish
Execute
 ↓
Finish

15. Human-in-the-Loop Architecture

When working through the 15 Human-in-the-Loop Architecture stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the 15 Human-in-the-Loop Architecture stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Agent Recommendation
       ↓
Policy Check
       ↓
High-Risk Action?
    ↓        ↓
   Yes       No
    ↓         ↓
Human       Execute
Approval
    ↓
Execute

16. Agent Guardrails

The 16 Agent Guardrails stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Input Guardrails

The Input Guardrails stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Tool Guardrails

The Tool Guardrails stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve. The Tool Guardrails stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Output Guardrails

For the Output Guardrails stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Execution Guardrails

For the Execution Guardrails stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

17. Implementation Frameworks

For the 17 Implementation Frameworks stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the 17 Implementation Frameworks stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Python
+
LLM API or local model
+
Functions
+
FastAPI
+
Database

Graph-Based Frameworks

When working through the Graph-Based Frameworks stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Multi-Agent Frameworks

When working through the Multi-Agent Frameworks stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Enterprise AI Orchestration Platforms

When working through the Enterprise AI Orchestration Platforms stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Enterprise AI Orchestration Platforms stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

18. A Practical Implementation Framework

The 18 A Practical Implementation stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Step 1: Define the Goal

The Step 1 Define the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Step 2: Define the Agent’s Responsibilities

The Step 2 Define the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Step 2 Define the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Retrieve account information
Retrieve opportunities
Analyze customer communication
Retrieve open support issues
Generate meeting briefing
Recommend discussion topics
Cannot modify CRM records
Cannot send email
Cannot change pricing
Cannot create contracts

19. Step 3: Define the Tools

For the 19 Step 3 Define stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

get_account(account_id)
get_opportunities(account_id)
get_support_cases(account_id)
get_email_history(account_id)
search_company_news(company_name)
manage_customer()
get_customer_profile()

20. Step 4: Design the Control Flow

For the 20 Step 4 Design stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

21. Step 5: Add State and Memory

For the 21 Step 5 Add stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the 21 Step 5 Add stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Conversation state
Task state
User preferences
Long-term knowledge
Audit history

22. Step 6: Add Observability

When working through the 22 Step 6 Add stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Request
Agent decision
Model used
Prompt version
Tool selected
Tool input
Tool output
Execution time
Token usage
Cost
Errors
Retries
Final response
Human overrides
Request ID: 78425
Step 1:
Intent → Customer Meeting PreparationStep 2:
Tool → get_account()Step 3:
Tool → get_opportunities()Step 4:
Tool → search_support_cases()Step 5:
LLM → Generate briefingTotal execution: 4.8 seconds
Tool calls: 3
Model calls: 2

23. Step 7: Evaluate the Agent

When working through the 23 Step 7 Evaluate stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Task Completion

When working through the Task Completion stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Task Completion stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Tool Selection

The Tool Selection stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Tool Accuracy

Groundedness

Safety

Efficiency

Latency

Cost

24. Failure Recovery

Tool Call
   ↓
Success?
 ↓      ↓
Yes     No
 ↓       ↓
Continue Retry
          ↓
       Still Fails?
        ↓       ↓
       Yes      No
        ↓        ↓
     Fallback  Continue
        ↓
     Escalate

25. Multi-Agent Systems: Use Them Carefully

Supervisor
 ├── Financial Analysis Agent
 ├── Legal Analysis Agent
 ├── Market Research Agent
 └── Report Generation Agent
Search Agent
Thinking Agent
Tool Agent
Summary Agent
Response Agent

26. The Enterprise Agent Architecture

User
                     ↓
               Agent Gateway
                     ↓
             Identity / Access
                     ↓
                  Router
                     ↓
       ┌─────────────┼─────────────┐
       ↓             ↓             ↓
   Sales Agent    HR Agent    Finance Agent
       ↓             ↓             ↓
             Agent Runtime
                   ↓
        ┌──────────┼──────────┐
        ↓          ↓          ↓
       RAG       Tools      Memory
        ↓          ↓          ↓
    Knowledge    APIs     Databases
      Base
                   ↓
             Policy Engine
                   ↓
          Human Approval Layer
                   ↓
              Observability
                   ↓
              Evaluation

27. Deterministic Software and Probabilistic AI

Probabilistic Intelligence
          +
Deterministic Control
          =
Reliable Agentic System

28. Start with Workflows, Then Add Autonomy

Stage 1 — Assistant

Stage 2 — Tool-Using Assistant

Stage 3 — Guided Agent

Stage 4 — Autonomous Workflow

Stage 5 — Multi-Agent System

29. What Makes an Effective AI Agent?

30. Final Thoughts

Architecture Summary

AI Agent
│
├── Model
├── Instructions
├── Context
├── Tools
├── Retrieval
├── Memory
├── State
├── Planning
├── Orchestration
├── Guardrails
├── Human Oversight
├── Observability
└── Evaluation

Operational checklist