This article is published in English.
Practical notes: Stop Watching Your Coding Agent: Build a System You Can Trust
Operable walkthrough of Practical notes: Stop Watching Your Coding Agent: Build a System You Can Trust: contracts, checks, and drop-in code slots for teams shipping this pattern.
Use this as an operator-facing rebuild of the ideas in “Stop Watching Your Coding Agent: Build a System You Can Trust”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
The problem: you are probably still doing half the work
For the The problem you are stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
You: Fix the login bug.
Agent: Done.
You: opens browser
You: It still doesn't work.
Agent: Ah. I found the problem.
You: No, that's not it.
Agent: You're right. I found the REAL problem.
You: sends screenshot
Agent: Ah...
1. Give the agent one command for “done”
For the 1 Give the agent stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
scripts/verify.sh
#!/usr/bin/env bash
set -euo pipefail
echo "== Python lint =="
uv run ruff check backend
echo "== Python types =="
uv run mypy backend
echo "== Python tests =="
uv run pytest -q
echo "== Frontend lint =="
npm --prefix frontend run lint
echo "== Frontend tests =="
npm --prefix frontend test -- --run
echo "Verification passed."
#!/usr/bin/env bash
set -euo pipefail
echo "== Python lint =="
python -m ruff check backend
echo "== Python types =="
python -m mypy backend
echo "== Python tests =="
python -m pytest -q
echo "== Frontend lint =="
npm --prefix frontend run lint
echo "== Frontend tests =="
npm --prefix frontend test -- --run
echo "Verification passed."
chmod +x scripts/verify.sh
verify:
./scripts/verify.sh
make verify
inspect
↓
change code
↓
verify
↓
failure
↓
inspect
↓
change code
↓
verify
2. For bugs, demand proof before the fix
For the 2 For bugs demand stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the 2 For bugs demand stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
def parse_timeout(value: str) -> float:
if value.endswith("s"):
return float(value[:-1])
if value.endswith("m"):
return float(value[:-1]) * 60
return float(value)
250ms is interpreted incorrectly.
def test_parse_timeout_milliseconds():
assert parse_timeout("250ms") == 0.25
uv run pytest tests/test_timeout.py -q
python -m pytest tests/test_timeout.py -q
def parse_timeout(value: str) -> float:
if value.endswith("ms"):
return float(value[:-2]) / 1000
if value.endswith("s"):
return float(value[:-1])
if value.endswith("m"):
return float(value[:-1]) * 60
return float(value)
reported bug
↓
observed failure
↓
code change
↓
observed success
3. Give the agent an onboarding manual
When working through the 3 Give the agent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
# Project
FastAPI backend + React frontend.
Python dependencies are managed with uv.
## Important directories
backend/app/api/ HTTP endpoints
backend/app/services/ business logic
frontend/src/features/ feature code
tests/ backend tests
## Commands
Fast Python tests:
uv run pytest -q tests/unit
Full verification:
make verify
Development:
make dev
## Working rules
Before editing:
1. Reproduce the problem.
2. Inspect the implementation involved.
3. Find similar existing code before creating a new pattern.
4. Identify or add a test.
Before completion:
1. Run relevant tests.
2. Run `make verify`.
3. Inspect `git diff`.
4. Report exactly what was verified.
Fast Python tests:
python -m pytest -q tests/unit
4. Turn recurring lessons into Skills
When working through the 4 Turn recurring lessons stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
skills/debug-with-evidence/SKILL.md
# Debug with evidence
Before modifying production code:
1. Capture the exact symptom.
2. Reproduce it.
3. Find the narrowest failing case.
4. Inspect the code actually executed.
5. Form hypotheses only after gathering evidence.
6. Prefer experiments that distinguish competing explanations.
7. Add a regression test when practical.
8. Make the smallest justified fix.
9. Rerun the reproduction.
10. Run full verification.
For Python projects managed by uv, run Python tools with `uv run`.
Report:
- observed failure
- root cause
- evidence
- files changed
- verification performed
#!/usr/bin/env bash
set -euo pipefail
echo "=== STATUS ==="
git status --short
echo
echo "=== RECENT COMMITS ==="
git log --oneline -10
echo
echo "=== DIFF ==="
git diff --stat
echo
echo "=== TESTS ==="
uv run pytest -q --tb=short
python -m pytest -q --tb=short
if rg 'app\.database' frontend/src
then
echo "Frontend may not import app.database"
exit 1
fi
"Don't import X here."
→ dependency check
"Every endpoint needs authorization."
→ middleware + test
"Don't forget to regenerate the schema."
→ CI check
"Every bug fix needs a regression test."
→ workflow rule
"Don't modify generated files."
→ generated-file check
uv run ruff check .
uv run mypy .
uv run pytest
python -m ruff check .
python -m mypy .
python -m pytest
6. Make the easiest solution the correct solution
When working through the 6 Make the easiest stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the 6 Make the easiest stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
components/
services/
hooks/
types/
validation/
screens/
features/
├── billing/
│ ├── api.ts
│ ├── model.ts
│ ├── BillingPage.tsx
│ └── BillingPage.test.tsx
│
└── login/
├── api.ts
├── model.ts
├── LoginPage.tsx
└── LoginPage.test.tsx
7. Use a fresh agent as reviewer
The 7 Use a fresh stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Agent A
↓
implements
↓
Agent B
↓
reviews from fresh context
Check:
1. Does the change actually satisfy the task?
2. Can you reproduce the original bug?
3. Are edge cases missing?
4. Were tests weakened?
5. Is there unnecessary complexity?
6. Are architectural boundaries violated?
7. Is existing functionality duplicated?
8. Do the tests verify behavior?
For Python changes, run the relevant checks yourself:
uv run ruff check .
uv run mypy .
uv run pytest
python -m ruff check .
python -m mypy .
python -m pytest
confirmed defect
plausible concern
stylistic preference
8. Parallelize with worktrees, not chaos
The 8 Parallelize with worktrees stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
git worktree add ../app-auth -b agent/auth
git worktree add ../app-search -b agent/search
git worktree add ../app-billing -b agent/billing
app-auth/
app-search/
app-billing/
uv sync
uv run pytest
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
python -m pytest
Agent 1: investigate authentication bug
Agent 2: implement CSV export
Agent 3: profile search performance
Agent 1: refactor authentication
Agent 2: refactor authentication differently
Agent 3: rename files both others are editing
9. Treat every human correction as data
The 9 Treat every human stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The 9 Treat every human stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Agent lacked project knowledge?
→ improve AGENTS.md
Agent didn't know the procedure?
→ create a Skill
Bug escaped?
→ regression test
Same architectural mistake again?
→ CI/static rule
Task was ambiguous?
→ improve task template
Agent trusted its own solution too easily?
→ independent reviewer
agent makes mistake
↓
human understands why
↓
lesson becomes process
↓
process becomes Skill/test/CI
↓
future agent avoids whole category of mistake
The setup you would build first
For the The setup you would stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
pyproject.toml
uv.lock
AGENTS.md
Makefile
scripts/verify.sh
skills/debug-with-evidence/SKILL.md
skills/review-change/SKILL.md
uv init
uv sync
uv add --dev pytest ruff mypy
uv run pytest
uv run ruff check .
uv run mypy .
pip install pytest ruff mypy
python -m pytest
python -m ruff check .
python -m mypy .
1. Investigate.
2. Reproduce.
3. Write failing test.
4. Implement smallest fix.
5. Run fast tests.
6. Run full verification.
7. Fresh agent reviews diff.
8. Human corrections become permanent rules.
Own this task end to end.
Before editing:
- inspect the relevant implementation,
- reproduce the problem,
- examine similar existing code.
During implementation:
- make the smallest coherent change,
- add or update tests,
- use `uv run` for Python tools,
- verify while iterating.
Before completion:
- run full verification,
- inspect the final diff,
- independently check the original requirement.
Report what changed, what was verified,
and any remaining uncertainty.
The bigger idea
For the The bigger idea stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
prompt → code
requirement
↓
agent
↓
code
↓
execution
↓
verification
↓
review
↓
feedback
↓
better Skills / tests / architecture
↺
uv run pytest tests/test_bug.py -q
uv run ruff check .
uv run mypy .
make verify
python -m pytest tests/test_bug.py -q
python -m ruff check .
python -m mypy .
make verify
Operational checklist
When working through the Operational checklist stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest.
Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.
Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 780e678b0ae3: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.