This article is published in English.
Practical notes: Why No One Talks About LangChain, LangGraph Or AutoAgent
Operable walkthrough of Practical notes: Why No One Talks About LangChain, LangGraph Or AutoAgent: contracts, checks, and drop-in code slots for teams shipping this pattern.
This walkthrough rebuilds the path from raw materials to a working system for: Why No One Talks About LangChain, LangGraph Or AutoAgent. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For Overview, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
┌────────────────────────────────────────┐
│ POST-FRAMEWORK PARADIGM │
└──────────────────┬─────────────────────┘
│
┌─────────────┴─────────────┐
▼ ▼
┌──────────────┐ ┌──────────────┐
│ PIPELINE │ │ HARNESS │
│ MODE │ │ MODE │
├──────────────┤ ├──────────────┤
│ • Bounded │ │ • Unbounded │
│ • Code-Run │ │ • Model-Run │
│ • Compute │ │ • Context │
│ Bound │ │ Bound │
└──────────────┘ └──────────────┘
Table of Contents
When working through Table of Contents, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
1. The Work Classification Nobody Is Doing
When working through 1. The Work Classification Nobody Is Doing, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Is the work decomposable into a finite set of
named states with deterministic transitions?
[ YES ] ──► PIPELINE MODE
• Invoice Extraction
• Document Classification
• Approval Routing
• KYC Verification
• Data Enrichment
[ NO ] ──► HARNESS MODE
• Coding Agents
• Research Agents
• Ephemeral Tool/Shell Use
• Legacy Code Migrations
• Open-Ended Investigation
2. Pipeline Mode: When the Work Is Decomposable
When working through 2. Pipeline Mode: When the Work Is Decomposable, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through 2. Pipeline Mode: When the Work Is Decomposable, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
[Input Data payload]
│
▼
┌──────────────┐
│ State 1 │ ──► [Structured LLM Call]
└──────┬───────┘ │
│ ▼
│ ┌──────────────┐
│ │ Validation │ ──► [Fail] ──► [Error State]
│ └──────┬───────┘
│ │ [Pass]
▼ ▼
┌──────────────┐
│ State 2 │ ──► [Structured LLM Call]
└──────┬───────┘ │
│ ▼
│ ┌──────────────┐
│ │ Validation │ ──► [Fail] ──► [Error State]
│ └──────┬───────┘
│ │ [Pass]
▼ ▼
[Final Success Output]
3. Harness Mode: When the Work Is Not Decomposable
- Harness Mode: When the Work Is Not Decomposable works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
┌──────────────────────────────────────────────┐
│ HARNESS LAYER │
│ Executes container limits & system safety │
├──────────────────────────────────────────────┤
│ [Safety Hooks] [Budgets] [Time Constraints] │
├──────────────────────────────────────────────┤
│ ┌────────────────────────────────────────┐ │
│ │ MODEL OWNS LOOP │ │
│ │ Plan ──► Execute ──► Observe ──► Plan │ │
│ │ │ │
│ │ Native capabilities: │ │
│ │ • File / Sandboxed Shell access │ │
│ │ • Context-isolated subagent spawning │ │
│ │ • Skill discovery & loading │ │
│ └────────────────────────────────────────┘ │
└──────────────────────┬───────────────────────┘
▼
Result, Partial State, or Escalation
The Failure Modes of Harness Mode
The Failure Modes of Harness Mode works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
4. The Mathematical Reality of Both Modes
- The Mathematical Reality of Both Modes works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
- The Mathematical Reality of Both Modes works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Pipeline Mode: The Inference-Time Scaling Law
For Pipeline Mode: The Inference-Time Scaling Law, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Harness Mode: The Context Degradation Curve
For Harness Mode: The Context Degradation Curve, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.
P(Success)
▲
1.0├───────────┐
│ │
│ └───┐
0.5│ │
│ └─────────────► Context Size (Tokens)
0└───────────┬───┬─────────────
50k 100k
5. The Shape of a 2026 Harness (Implementation Blueprint)
For 5. The Shape of a 2026 Harness (Implementation Blueprint), define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For 5. The Shape of a 2026 Harness (Implementation Blueprint), define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
import logging
import os
from typing import Dict, Any, List
from pydantic import BaseModel, Field
# 2026 SDK abstractions (analogous across major provider frameworks)
from modern_agent_sdk import Agent, Skill, SubAgent
from modern_agent_sdk.hooks import PreToolUse, PostToolUse
from modern_agent_sdk.budgets import TokenBudget, StepBudget, WallTimeBudgetlogger = logging.getLogger("EnterpriseHarness")# 1. Skills are dynamically loaded from disk, not declared inline.
# Each skill is an isolated directory containing metadata (SKILL.md),
# runtime scripts, and specialized sandbox requirements.
SKILLS_DIR = "./skills"
all_skills = Skill.load_directory(SKILLS_DIR)# 2. Safety Hooks: Gateway controls running in the host environment.
# These run on the host *before* any action is committed inside the sandbox.
def pre_tool_execution_hook(context: Dict[str, Any], tool_call: BaseModel) -> PreToolUse:
"""
Validates security boundaries and consumption limits before the model executes a tool.
"""
# Hard safety boundary: Prevent destructive shell operations
if tool_call.name == "execute_shell":
command = tool_call.args.get("command", "")
if any(bad_cmd in command for bad_cmd in ["rm -rf", "chmod", "wget"]):
logger.error(f"Execution blocked: Blocked command detected: '{command}'")
return PreToolUse.deny("Destructive shell operations are prohibited in this sandbox.")
# Budget check: Halt network tools if API budget is running low
if tool_call.name == "network_request":
if context["budget"].remaining_tokens < 15_000:
return PreToolUse.deny("Insufficient remaining token budget to execute external network calls.")
logger.info(f"Approved tool call: {tool_call.name}")
return PreToolUse.allow()def post_tool_execution_hook(context: Dict[str, Any], tool_call: BaseModel, result: Any) -> PostToolUse:
"""
Evaluates tool execution outcomes to detect systemic failure loops.
"""
# Detect repeating error patterns to prevent infinite execution loops
if result.is_error and context["consecutive_failures"] >= 3:
logger.warning("System detected a repeating error loop. Forcing escalation.")
return PostToolUse.escalate("Agent is stuck in an execution failure loop.")
return PostToolUse.continue_loop()# 3. Context Isolation via Subagents
# To prevent context poisoning, the parent agent never sees the child's
# scratchpad or intermediate execution steps—only the final verified output.
def spawn_research_subagent(target_query: str) -> str:
"""
Spawns a specialized subagent in a separate context window to perform a task,
keeping the parent's working context completely clean.
"""
logger.info(f"Spawning isolated subagent for query: '{target_query}'")
subagent = SubAgent(
model="claude-sonnet-4-5",
skills=all_skills.filter_by_tag("research"),
budgets=[
TokenBudget(max_input_tokens=40_000, max_output_tokens=8_000),
StepBudget(max_steps=12)
]
)
# Run task and return only the clean, compiled summary
output = subagent.execute(task=target_query)
return output.summary# 4. Assembling the Harness Container
# The harness manages the execution sandbox, enforces constraints, and handles state.
agent_harness = Agent(
model="claude-opus-4-7",
skills=all_skills,
subagent_factories={
"researcher": spawn_research_subagent
},
hooks={
"pre_tool_use": pre_tool_execution_hook,
"post_tool_use": post_tool_execution_hook
},
budgets=[
TokenBudget(max_total_tokens=600_000),
StepBudget(max_steps=100),
WallTimeBudget(max_seconds=1800) # 30-minute hard cap
],
persistence_store="./.agent_state_db", # State is persisted per step to survive restarts
escalation_handler=lambda issue: open_human_review_ticket(issue)
)# 5. Execution
if __name__ == "__main__":
task_input = "Migrate this legacy Django 4.2 codebase to FastAPI with 100% endpoint parity."
try:
# The agent loop is entirely internal to the model.
# There is no manual 'while True' in our orchestration code.
result = agent_harness.run(task=task_input)
logger.info(f"Task completed. Result: {result.status}")
except Exception as e:
logger.critical(f"Harness terminated execution: {e}")
Key Differences from Legacy 2024 Frameworks:
When working through Key Differences from Legacy 2024 Frameworks:, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
6. Architectural Collisions: Putting the Wrong Shape on the Wrong Problem
When working through 6. Architectural Collisions: Putting the Wrong Shape on the Wrong Problem, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
┌─────────────────────────────────────────────────────┐
│ ARCHITECTURAL MATCH │
├──────────────────────────┬──────────────────────────┤
│ PIPELINE WORKLOAD │ HARNESS WORKLOAD │
├──────────────────────────┼──────────────────────────┤
│ • Goal: Exact Extraction │ • Goal: Code Refactoring,│
│ & Routing │ Agent Tasks, Migration │
│ │ │
│ • WRONG: Harness │ • WRONG: Pipeline │
│ (High latency, costly, │ (FSM State explosion, │
│ non-deterministic) │ inflexible schema) │
│ │ │
│ • RIGHT: Pipeline FSM │ • RIGHT: Sandbox Harness │
│ (Predictable transitions)│ (Dynamic loop execution)│
└──────────────────────────┴──────────────────────────┘
Case 1: The Harness-on-Pipeline Error (Over-Engineering)
When working through Case 1: The Harness-on-Pipeline Error (Over-Engineering), write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours. When working through Case 1: The Harness-on-Pipeline Error (Over-Engineering), write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Case 2: The Pipeline-on-Harness Error (Under-Engineering)
Case 2: The Pipeline-on-Harness Error (Under-Engineering) works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Expose tools with narrow schemas and explicit side-effect labels. Hosts need to know which calls mutate state before they auto-approve.
How to Select the Correct Architecture
How to Select the Correct Architecture works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
7. The Skill Layer: The Open Standard of 2026
- The Skill Layer: The Open Standard of 2026 works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
- The Skill Layer: The Open Standard of 2026 works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
skills/
├── document_parser/
│ ├── SKILL.md <-- Human/Model readable metadata & capabilities
│ ├── run.py <-- The execution logic running inside the sandbox
│ └── reference_rules.pdf <-- Domain-specific constraints and edge-cases
8. Conclusion: The New Engineering Job
For 8. Conclusion: The New Engineering Job, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Operational checklist
Operational checklist works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 9f921d978495: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.