This article is published in English.
Beyond Spec-Driven Development: The Agentic Engineering Playbook That’s
Operable walkthrough of Beyond Spec-Driven Development: The Agentic Engineering Playbook That’s: contracts, checks, and drop-in code slots for teams shipping this pattern.
The following notes reconstruct a practical path around “Beyond Spec-Driven Development: The Agentic Engineering Playbook That’s Replacing How We Build Software”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
The Name That Stuck and Why It Matters
The The Name That Stuck stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
The Five Layers:
The The Five Layers stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Layer 1 — Specification: Define What You Want Before You Ask for It
The Layer 1 Specification Define stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Layer 1 Specification Define stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Implement it — the spec template:
For the Implement it the spec stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
# Feature Spec: [Feature Name]
## Last updated: [Date] — this document is the source of truth
## What this must always do
- [ ] [Behaviour 1 — specific, testable]
- [ ] [Behaviour 2 — specific, testable]## What this must never do
- [ ] [Boundary 1 — e.g. "Never process refund >$500 without human approval"]
- [ ] [Boundary 2 — e.g. "Never fabricate a citation"]## Success criteria (measurable)
- Context recall: >0.85
- Faithfulness: >0.90
- Cost per query: <$0.15
- Latency p95: <2s## Failure signals (alert if breached)
- Any metric drops >5% from 7-day baseline
- Cost per task exceeds 3x expected range
- User complaint rate exceeds 2% of sessions## Decomposition (for parallel agents)
- [ ] Task A: [scope] — can be delegated independently
- [ ] Task B: [scope] — depends on Task A output
- [ ] Task C: [scope] — can run parallel with Task A
Layer 2 — Orchestration: Multiple Agents, One Coherent Output
For the Layer 2 Orchestration Multiple stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Implement it — parallel orchestration in practice:
For the Implement it parallel orchestration stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
class AgenticOrchestrator:
"""Plan → Execute (parallel) → Verify"""
def plan(self, feature_spec: dict) -> dict:
"""Decompose feature into independently delegatable tasks."""
tasks = []
for component in feature_spec["decomposition"]:
tasks.append({
"id": component["id"],
"spec": component["scope"],
"constraints": feature_spec["must_never"] + component.get("local_rules", []),
"depends_on": component.get("depends_on", []),
"verification": component.get("success_criteria", []),
})
independent = [t for t in tasks if not t["depends_on"]]
sequential = [t for t in tasks if t["depends_on"]]
return {"parallel": independent, "sequential": sequential} async def execute(self, plan: dict):
"""Run independent tasks in parallel, sequential tasks in order."""
parallel_results = await asyncio.gather(*[
self.delegate_to_agent(task) for task in plan["parallel"]
])
for task in plan["sequential"]:
result = await self.delegate_to_agent(task, prior=parallel_results)
parallel_results.append(result)
return parallel_results def verify(self, results: list, spec: dict) -> dict:
"""Check all outputs against the spec — structural, not line-by-line."""
issues = []
for result in results:
if not self.satisfies_spec(result, spec):
issues.append(f"{result['id']}: does not satisfy spec")
if self.conflicts_with(result, results):
issues.append(f"{result['id']}: conflicts with another module")
return {"passed": len(issues) == 0, "issues": issues}
For the Implement it parallel orchestration stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Layer 3 — Codified Rules: Teaching Agents How Your Team Works
When working through the Layer 3 Codified Rules stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Implement it — the rules hierarchy:
When working through the Implement it the rules stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
your-company/
├── .ai-rules/
│ └── org-rules.md ← Company-wide: security, compliance, style
│
├── team-payments/
│ ├── .ai-rules/
│ │ └── team-rules.md ← Team-level: error handling, testing, deploys
│ │
│ ├── service-checkout/
│ │ ├── CLAUDE.md ← Repo-level: architecture, tech stack, builds
│ │ ├── src/
│ │ │ ├── auth/
│ │ │ │ └── .ai-rules.md ← Module: auth-specific constraints
│ │ │ └── payments/
│ │ │ └── .ai-rules.md ← Module: PCI compliance rules
# org-rules.md (loaded for every agent, every repo)
- Never commit secrets or credentials
- All public APIs require authentication
- Error responses must never expose stack traces
- Log every state-changing operation with user context
# team-rules.md (inherits org, adds team specifics)
- Use Result<T, E> pattern — never throw exceptions
- All database queries go through the repository layer
- Tests must cover the happy path + 2 failure modes minimum# CLAUDE.md (inherits team, adds repo specifics)
- This repo uses Express.js + Prisma + PostgreSQL
- Run `npm test` before suggesting any PR is ready
- Migrations are append-only — never modify applied migrations# module .ai-rules.md (inherits repo, adds module specifics)
- auth/: All token operations must use constant-time comparison
- payments/: PCI DSS requires field-level encryption on card data
Layer 4 — Human Oversight: Delegate, Review, Own
When working through the Layer 4 Human Oversight stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Layer 4 Human Oversight stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Implement it — the agent output review checklist:
The Implement it the agent stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
## Agent Output Review Checklist
### Spec alignment (does it do what was asked?)
- [ ] All "must always do" behaviours are implemented
- [ ] No "must never do" boundaries are violated
- [ ] Success criteria from the spec are met or tested### Architectural coherence (does it fit the system?)
- [ ] No new dependencies introduced without justification
- [ ] Consistent with naming, patterns, and structure of existing code
- [ ] No duplication of logic that exists elsewhere### Safety and edge cases
- [ ] Error handling covers the failure modes the spec anticipated
- [ ] No hardcoded credentials, keys, or environment-specific values
- [ ] Input validation present on all external-facing boundaries### What the agent cannot check for itself
- [ ] Does this make business sense? (not just technical correctness)
Layer 5 — Observable Development: Know What the System Did and Whether It Was Right
The Layer 5 Observable Development stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
In practice, observable development means three things:
The In practice observable development stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The In practice observable development stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Implement it — the minimum observable development setup:
For the Implement it the minimum stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
from dataclasses import dataclass, field
from datetime import datetime
@dataclass
class AgentTrace:
"""Minimum trace for every agent-delegated task."""
task_id: str
agent_id: str
spec_given: str # What was the agent told to do?
tools_used: list[str] # Which tools did it invoke?
files_modified: list[str] # What did it change?
model_version: str # Which model, which version?
tokens_consumed: int # What did it cost?
started_at: datetime = field(default_factory=datetime.utcnow)
completed_at: datetime | None = None
verification_result: str = "" # pass / fail / needs_review
human_reviewer: str = "" # Who signed off?
issues_found: list[str] = field(default_factory=list)
## Weekly Observable Development Review (15 minutes)
1. How many agent tasks were delegated this week? ___
2. How many passed verification on first attempt? ___ (target: >80%)
3. Which task category had the most issues? ___
4. Top 3 issues found during review:
- ___
- ___
- ___
5. Which issues should become codified rules (Layer 3)? ___→ Update CLAUDE.md with any new rules from this week's observations.
What This Means for Your Career?
For the What This Means for stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Here Is The Monday-to-Friday Playbook:
For the Here Is The Monday-to-Friday stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Here Is The Monday-to-Friday stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
What you Got Wrong in the SDD Article?
When working through the What you Got Wrong stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
What Comes Next?
When working through the What Comes Next stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Operational checklist
For the Operational checklist stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.
Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 3031bd6e2b68: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.
For the hardening note 0 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 0/810: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 1 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 1/810: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 2 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 2/810: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 3 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 3/810: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 4 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 4/810: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 5 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 5/810: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 6 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 6/810: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 7 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 7/810: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 8 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 8/810: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 9 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 9/810: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 10 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 10/810: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 11 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 11/810: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 12 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 12/810: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.