Home / Articles / Practical notes: Agentic Context Engineering: A Step-by-Step Guide for Machine

This article is published in English.

Practical notes: Agentic Context Engineering: A Step-by-Step Guide for Machine

Operable walkthrough of Practical notes: Agentic Context Engineering: A Step-by-Step Guide for Machine: contracts, checks, and drop-in code slots for teams shipping this pattern.

5021 words

The following notes reconstruct a practical path around “Agentic Context Engineering: A Step-by-Step Guide for Machine Learning Engineers”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

1. Start With the Fundamental Question: What Is Context?

The 1 Start With the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Context → LLM → Output
User question
      +
Model configuration
      +
Recent evaluation metrics
      +
Training dataset statistics
      +
Recent production images
      +
Deployment history
      +
Data distribution statistics
      +
Recent code changes
      ↓
    LLM
      ↓
Diagnosis

2. Prompt Engineering Is Not Context Engineering

The 2 Prompt Engineering Is stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

Prompt engineering

The Prompt engineering stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices. The Prompt engineering stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

You are an expert ML engineer.
Analyze the following anomaly detection results
and identify the most likely causes of the performance drop.
Consider:
1. Data distribution shift
2. Model degradation
3. Label quality
4. Hardware changes
5. Preprocessing changes

Context engineering

For the Context engineering stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

                    ┌──────────────┐
                    │ User Request │
                    └──────┬───────┘
                           ↓
                    ┌──────────────┐
                    │ Agent State  │
                    └──────┬───────┘
                           ↓
              ┌────────────┴────────────┐
              ↓                         ↓
        Retrieval                    Tools
              ↓                         ↓
       Documentation             Metrics / DB
              └────────────┬────────────┘
                           ↓
                    ┌──────────────┐
                    │   Context    │
                    └──────┬───────┘
                           ↓
                         LLM
                           ↓
                        Action

3. Why Context Becomes Hard for Agents

For the 3 Why Context Becomes stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

User → LLM → Answer
User
 ↓
Agent
 ↓
Search documentation
 ↓
Read results
 ↓
Call database
 ↓
Analyze data
 ↓
Call another tool
 ↓
Observe result
 ↓
Modify hypothesis
 ↓
Search again
 ↓
Take action
 ↓
Final answer

4. The Four Fundamental Context Problems

For the 4 The Four Fundamental stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the 4 The Four Fundamental stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Problem 1: Missing context

When working through the Problem 1 Missing context stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Agent:
"The model may be suffering from data drift."
Reality:
The model was recently changed from ViT-B to ViT-L.

Problem 2: Irrelevant context

When working through the Problem 2 Irrelevant context stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Query
 ↓
Vector database
 ↓
50 documents
 ↓
LLM

Problem 3: Stale context

When working through the Problem 3 Stale context stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Problem 3 Stale context stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Current production model:
anomaly-detector-v4
Old documentation:
anomaly-detector-v2

Problem 4: Poorly structured context

The Problem 4 Poorly structured stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Accuracy: 0.91
Accuracy: 0.87
Accuracy: 0.84
Model: anomaly-detector-v4
Dataset: Production
Metric: Image-level AUROC
Previous week: 0.91
Current week:  0.84
Change:       -7.7 percentage pointsDeployment:
- Version: v4.2
- Date: 2026-08-12
- Preprocessing: resize=336

5. A Useful Mental Model: Context as a Data Pipeline

The 5 A Useful Mental stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

Raw Data
 ↓
Cleaning
 ↓
Feature Extraction
 ↓
Transformation
 ↓
Model
 ↓
Prediction
Raw Information
 ↓
Retrieval
 ↓
Filtering
 ↓
Ranking
 ↓
Transformation
 ↓
Compression
 ↓
Context Assembly
 ↓
LLM
 ↓
Action
 ↓
New Observation
 ↓
Context Update

6. Step 1 — Define the Agent’s Objective

The 6 Step 1 Define stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The 6 Step 1 Define stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Objective:
Diagnose production degradation

7. Step 2 — Identify the Context Sources

For the 7 Step 2 Identify stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

                  ML Debugging Agent
                          │
        ┌─────────────────┼─────────────────┐
        ↓                 ↓                 ↓
   Metrics DB        Git Repository      Experiment DB
        │                 │                 │
        ↓                 ↓                 ↓
   Performance         Code changes       Experiments
        │                 │                 │
        └─────────────────┼─────────────────┘
                          ↓
                       Context

Documents

For the Documents stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Databases

For the Databases stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Databases stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Tools

When working through the Tools stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

Memory

When working through the Memory stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Environment

When working through the Environment stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Environment stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

8. Step 3 — Retrieve Only What Matters

The 8 Step 3 Retrieve stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

Question
 ↓
Embedding
 ↓
Vector DB
 ↓
Top 10 chunks
 ↓
LLM
1. Find current deployment
2. Find previous deployment
3. Compare configurations
4. Retrieve associated code changes
5. Retrieve metric changes

9. Step 4 — Use Structured Context

The 9 Step 4 Use stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

The deployment happened recently and the model seems
to have lower performance and there were some preprocessing
changes and the new version uses 336 resolution...
{
  "model": "anomaly-detector-v4",
  "deployment": "v4.2",
  "deployment_date": "2026-08-12",
  "image_size": 336,
  "previous_image_size": 224,
  "auroc_previous": 0.91,
  "auroc_current": 0.84
}
image_size:
224 → 336
AUROC:
0.91 → 0.84

10. Step 5 — Separate Facts, Hypotheses, and Actions

The 10 Step 5 Separate stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The 10 Step 5 Separate stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Facts

For the Facts stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Current AUROC = 0.84
Previous AUROC = 0.91
Image resolution changed from 224 to 336

Hypotheses

For the Hypotheses stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Hypothesis:
The resolution change may have caused distribution mismatch.

Actions

For the Actions stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Actions stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Action:
Evaluate v4.2 on the previous preprocessing configuration.
{
  "facts": [
    "AUROC dropped from 0.91 to 0.84",
    "Image resolution changed from 224 to 336"
  ],
  "hypotheses": [
    {
      "claim": "Resolution change caused degradation",
      "confidence": 0.65
    }
  ],
  "actions_completed": [
    "Compared deployment configurations"
  ],
  "next_action": "Run controlled preprocessing experiment"
}

11. Step 6 — Compress Context

When working through the 11 Step 6 Compress stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Investigation Summary
Objective:
Diagnose production AUROC degradation.Observed:
- AUROC decreased 0.91 → 0.84.
- Deployment v4.2 introduced 336px preprocessing.
- Model weights unchanged.
- Data volume unchanged.Ruled out:
- Model checkpoint change.
- Infrastructure failure.Current hypothesis:
Preprocessing change may be responsible.Next experiment:
Evaluate v4.2 using 224px preprocessing.

12. Step 7 — Give the Agent a Working Memory

When working through the 12 Step 7 Give stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

                 Agent Memory
                      │
       ┌──────────────┼──────────────┐
       ↓              ↓              ↓
  Working Memory   Long-Term      External
                   Memory          Knowledge
       │              │              │
       ↓              ↓              ↓
 Current task     Past decisions   Documents
 Current facts    User preferences Databases
 Hypotheses       Past results     APIs

Working memory

When working through the Working memory stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Working memory stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Long-term memory

The Long-term memory stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

External knowledge

The External knowledge stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

13. Step 8 — Let the Agent Decide What Context It Needs

The 13 Step 8 Let stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The 13 Step 8 Let stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Question
 ↓
Retrieve
 ↓
LLM
 ↓
Answer
Question
 ↓
LLM
 ↓
"What information am I missing?"
 ↓
Retrieve
 ↓
Observe
 ↓
"What else do I need?"
 ↓
Tool call
 ↓
Observe
 ↓
Update hypothesis
 ↓
Retrieve again
 ↓
Answer
I need:
1. Current metrics
2. Historical metrics
3. Recent deployments
I see a preprocessing change.
I now need:
4. Code/config diff
5. Evaluation by preprocessing version
The degradation occurs only on the new preprocessing path.
Hypothesis strengthened.

14. A Concrete Example: ML Debugging Agent

For the 14 A Concrete Example stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Image-level AUROC
Monday: 0.94
Tuesday: 0.93
Wednesday: 0.92
Thursday: 0.85

Initial context

For the Initial context stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Task:
Diagnose the AUROC degradation.
Current metric:
0.85Previous metric:
0.92

Tool call 1: deployment system

For the Tool call 1 deployment stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

Current model:
v4.2
Previous model:
v4.1

For the Tool call 1 deployment stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Tool call 2: model registry

When working through the Tool call 2 model stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

Weights:
v4.1 == v4.2

Tool call 3: configuration service

When working through the Tool call 3 configuration stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

Resize:
224 → 336
Normalization:
unchanged

Tool call 4: evaluation service

When working through the Tool call 4 evaluation stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.

v4.2 + 224px:
AUROC = 0.93
v4.2 + 336px:
AUROC = 0.85

When working through the Tool call 4 evaluation stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Finding:
The performance degradation is strongly associated with
the preprocessing change from 224px to 336px.
Evidence:
- Model weights unchanged.
- Deployment introduced 336px preprocessing.
- 224px evaluation restores AUROC to 0.93.
- 336px evaluation produces AUROC of 0.85.Recommendation:
Roll back preprocessing to 224px while investigating
why the new preprocessing configuration causes degradation.

15. Context Engineering and RAG

The 15 Context Engineering and stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

Query
 ↓
Retriever
 ↓
Documents
 ↓
LLM
Goal
 ↓
Agent
 ↓
Determine missing context
 ↓
Retrieve / Query / Execute
 ↓
Evaluate results
 ↓
Update state
 ↓
Retrieve again
 ↓
Compress
 ↓
Assemble context
 ↓
LLM
 ↓
Action

16. Context Engineering Is Similar to Feature Engineering

The 16 Context Engineering Is stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Raw data
 ↓
Feature engineering
 ↓
Feature selection
 ↓
Model
 ↓
Prediction
Raw information
 ↓
Context retrieval
 ↓
Context filtering
 ↓
Context transformation
 ↓
Context selection
 ↓
LLM
 ↓
Decision

17. Evaluating Context Quality

Bad retrieval
     ↓
Bad context
     ↓
Bad reasoning
     ↓
Bad action
     ↓
Bad answer

Retrieval quality

Context quality

Reasoning quality

Action quality

Final task success

Retrieval
   ↓
Context
   ↓
Reasoning
   ↓
Action
   ↓
Outcome

18. Common Mistakes

Mistake 1: “Just put everything in the prompt”

Mistake 2: Treating vector search as the entire solution

Mistake 3: Keeping infinite conversation history

Recent details
+
Compressed historical state
+
Relevant retrieved information

Mistake 4: Mixing facts with guesses

Mistake 5: Ignoring temporal context

timestamp
version
deployment
experiment
data snapshot
environment

19. A Practical Architecture for Your First Agent

                 ┌──────────────┐
                 │    User      │
                 └──────┬───────┘
                        ↓
                 ┌──────────────┐
                 │    Agent     │
                 └──────┬───────┘
                        ↓
               ┌─────────────────┐
               │ Context Manager │
               └───────┬─────────┘
                       ↓
          ┌────────────┼────────────┐
          ↓            ↓            ↓
       Search        Database      Tools
          │            │            │
          └────────────┼────────────┘
                       ↓
                 Context Assembly
                       ↓
                      LLM
                       ↓
                     Action
                       ↓
                   Observation
                       ↓
                 Context Update

20. Where Agentic Context Engineering Is Going

LLM + Tools
LLM
+
Memory
+
Retrieval
+
State
+
Tools
+
Environment
+
Context Management

Conclusion

Prompt Engineering
       ↓
RAG
       ↓
Memory
       ↓
Tool Use
       ↓
State Management
       ↓
Agentic Context Engineering

Operational checklist