Home / Articles / Practical notes: 21 Agentic Design Patterns Explained Simply

This article is published in English.

Practical notes: 21 Agentic Design Patterns Explained Simply

Operable walkthrough of Practical notes: 21 Agentic Design Patterns Explained Simply: contracts, checks, and drop-in code slots for teams shipping this pattern.

3700 words

This walkthrough rebuilds the path from raw materials to a working system for: 21 Agentic Design Patterns Explained Simply. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Agentic Design Patterns — Architecture Diagrams + Practical Guide (21 Patterns)

When working through the Agentic Design Patterns Architecture stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Table of contents

When working through the Table of contents stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

1) Prompt Chaining (Pipeline)

When working through the 1 Prompt Chaining Pipeline stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the 1 Prompt Chaining Pipeline stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

flowchart TD
A[Input] --> B[Step 1 Prompt e.g. summarize ]
B --> C[Step 2 Prompt e.g. extract structured data ]
C --> D[Step 3 Prompt e.g. format output ]
D --> E[Final Output]

2) Routing

The 2 Routing stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

flowchart TD
I[User Request Input] --> R{Router intent and confidence }
R --> A[Workflow A e.g. Q&A ]
R --> B[Workflow B e.g. coding ]
R --> C[Workflow C e.g. retrieval ]
R --> Q[Ask Clarifying Question]
A --> O[Output]
B --> O
C --> O
Q --> I

3) Parallelization

The 3 Parallelization stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

flowchart TD
I[Input] --> F[Fork]
F --> A[Task A e.g. retrieve source 1 ]
F --> B[Task B e.g. retrieve source 2 ]
F --> C[Task C e.g. retrieve source 3 ]
A --> J[Join Merge]
B --> J
C --> J
J --> O[Output]

4) Reflection (Generate → Critique → Refine)

The 4 Reflection Generate Critique stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The 4 Reflection Generate Critique stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

flowchart TD
D[Draft Output] --> C[Critique Review check requirements errors ]
C --> R[Revise using critique]
R --> D
C --> O[Final Output]

5) Tool Use (Function Calling)

For the 5 Tool Use Function stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

sequenceDiagram
autonumber
participant U as User
participant L as LLM Agent
participant T as Tool API
U->>L: Request
L->>L: Decide tool and arguments
L->>T: Call tool args
T-->>L: Tool result
L-->>U: Answer using result

6) Planning

For the 6 Planning stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

flowchart TD
G[Goal] --> P[Create Plan steps and dependencies and tools ]
P --> S1[Execute Step 1]
S1 --> C1{Step success }
C1 --> S2[Execute Step 2]
C1 --> RP[Revise Plan Recover]
RP --> P
S2 --> C2{Done }
C2 --> S3[Next Steps ]
C2 --> O[Output]
S3 --> C2

7) Multi‑Agent Collaboration

For the 7 Multi Agent Collaboration stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the 7 Multi Agent Collaboration stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

flowchart LR
G[Goal] --> C[Coordinator]
C --> R[Research Agent]
C --> B[Builder Agent]
C --> V[Verifier Reviewer Agent]
R --> S[Synthesis]
B --> S
V --> S
S --> O[Final Output]

8) Memory Management

When working through the 8 Memory Management stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

flowchart TD
E[Events Conversation] --> STM[Short-term Memory session buffer ]
E --> LTM[Long-term Memory Store vector DB ]
Q[Current Query] --> RET[Retrieve relevant memory]
LTM --> RET
STM --> CTX[Assemble Context]
RET --> CTX
CTX --> L[LLM Agent]
L --> O[Output]

9) Learning & Adaptation

When working through the 9 Learning Adaptation stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

flowchart TD
R[Run Agent] --> L[Log outcomes success fail and user edits ]
L --> A[Analyze patterns where it fails ]
A --> U[Update prompts routes retrieval or fine-tune ]
U --> E[Evaluate before rollout]
E --> R
E --> A

10) Model Context Protocol (MCP)

When working through the 10 Model Context Protocol stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the 10 Model Context Protocol stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

flowchart LR
A[Agent LLM] --> C[MCP Client]
C <--> S[MCP Server]
S --> T1[Tool: Documents]
S --> T2[Tool: DB]
S --> T3[Tool: Tickets]
T1 --> S
T2 --> S
T3 --> S
S --> C --> A

11) Goal Setting & Monitoring

The 11 Goal Setting Monitoring stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

flowchart TD
G[Define Goal and Success Criteria] --> X[Execute steps]
X --> M[Monitor state metrics progress budget risk ]
M --> X
M --> A[Adjust plan change route escalate]
A --> X
M --> O[Output]

12) Exception Handling & Recovery

The 12 Exception Handling Recovery stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

flowchart TD
A[Action Tool Call] --> E{Error }
E --> N[Next Step]
E --> R[Retry with backoff]
R --> S{Recovered }
S --> N
S --> F[Fallback route tool]
F --> T{Still failing }
T --> N
T --> H[Escalate to Human Safe Stop]

13) Human‑in‑the‑Loop (HITL)

The 13 Human in the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The 13 Human in the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

sequenceDiagram
autonumber
participant U as Human
participant A as Agent
participant S as System Tools
A->>U: Proposal and rationale
U-->>A: Approve Edit Reject
A->>S: Execute approved action
S-->>A: Result
A-->>U: Confirmation and summary

14) Knowledge Retrieval (RAG)

For the 14 Knowledge Retrieval RAG stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

flowchart TD
Q[Question] --> E[Embed Rewrite Query]
E --> R[Retrieve top-k chunks vector keyword hybrid ]
R --> RR[Rerank Filter optional ]
RR --> C[Compose grounded prompt question and context ]
C --> L[LLM]
L --> O[Answer and citations quotes optional ]

15) Inter‑Agent Communication (A2A)

For the 15 Inter Agent Communication stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

sequenceDiagram
autonumber
participant A as Agent A
participant B as Agent B
participant C as Agent C
A->>B: Task request schema and constraints
B-->>A: Result or stream updates
A->>C: Verification request
C-->>A: Verified flagged findings
A-->>A: Merge and decide next step

16) Resource‑Aware Optimization

For the 16 Resource Aware Optimization stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the 16 Resource Aware Optimization stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

flowchart TD
I[Request] --> S[Score difficulty and risk and SLA]
S --> D{Choose compute level}
D --> L1[Fast Cheap path small model and minimal tools]
D --> L2[Balanced path hybrid retrieval and standard model]
D --> L3[Strong path best model and RAG and Reflection]
L1 --> O[Output]
L2 --> O
L3 --> O

17) Reasoning Techniques

When working through the 17 Reasoning Techniques stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

flowchart TD
P[Problem] --> D[Decompose choose reasoning strategy]
D --> A[Act: tool calls sub-steps optional ]
A --> V[Verify constraints checks tests cross-check]
V --> D
V --> O[Answer]

18) Guardrails / Safety Patterns

When working through the 18 Guardrails Safety Patterns stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

flowchart TD
IN[User Input] --> IV[Input Validation policy risk checks ]
IV --> PC[Policy Constraints system rules and boundaries ]
PC --> TR[Tool Restrictions allowlist and sandbox and rate limits]
TR --> L[LLM Agent]
L --> OV[Output Validation PII leak checks and format checks]
OV --> OUT[Safe Output]
OV --> ESC[Escalate Refuse Human review]

19) Evaluation & Monitoring

When working through the 19 Evaluation Monitoring stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the 19 Evaluation Monitoring stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

flowchart TD
RUN[Agent Runs] --> LOG[Log traces inputs tools outputs latency cost ]
LOG --> EVAL[Evaluate quality golden set and metrics ]
EVAL --> DRIFT[Drift Anomaly detection]
DRIFT --> IMP[Improve prompts routes retrieval model ]
IMP --> DEP[Deploy and A B test]
DEP --> RUN

20) Prioritization

The 20 Prioritization stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

flowchart TD
T[Incoming tasks] --> N[Normalize into task objects]
N --> S[Score tasks urgency impact risk deps cost ]
S --> Q[Queue Scheduler]
Q --> X[Execute next task]
X --> U[Update scores new info failures deadlines ]
U --> S
X --> O[Outputs Results]

21) Exploration & Discovery

The 21 Exploration Discovery stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

flowchart TD
S[Start: unknown space] --> H[Generate hypotheses options]
H --> G[Gather evidence search tools experiments ]
G --> E[Evaluate findings rank eliminate]
E --> R[Refine hypotheses]
R --> G
E --> O[Best answer strategy]

Common pattern “recipes” (what teams actually ship)

The Common pattern recipes what stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Common pattern recipes what stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Operational checklist

The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.

Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.

Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for 30adaa1daab9: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.