This article is published in English.
Practical notes: Why AI Agent Skills Break When We Chain Them and the
Operable walkthrough of Practical notes: Why AI Agent Skills Break When We Chain Them and the: contracts, checks, and drop-in code slots for teams shipping this pattern.
The following notes reconstruct a practical path around “Why AI Agent Skills Break When We Chain Them and the Three-Layer Fix”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Three Levels That Match How Agents Actually Compose
The Three Levels That Match stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Atoms: Single Skills That Do One Thing
The Atoms Single Skills That stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
---
name: verify-email
description: Verify an email address using Hunter.io API.
Use when validating email deliverability before outreach.
allowed-tools: Bash
---
## Verify Email
1. Read the Hunter API key from $HUNTER_API_KEY
2. Call the Hunter email-verifier endpoint
3. Return: status (deliverable/risky/undeliverable), score, smtp_check
4. If the API errors, report the error. Do not guess.
Molecules: Explicit Chains of Atoms
The Molecules Explicit Chains of stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Molecules Explicit Chains of stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
---
name: qualify-lead
description: Research a company, find the right contact, verify
their email, output a qualified lead card.
allowed-tools: Bash Read Write
---
## Qualify Lead
Execute these steps IN ORDER.
### Step 1: Company Research
Use /research-company. Capture: size, industry, funding, tech stack.
### Step 2: Find Contact
Use /find-contact. Target: VP Eng, CTO, Head of Platform.
### Step 3: Find & Verify Email
Use /find-email, then /verify-email.
If undeliverable, return to Step 2 (max 3 attempts).
### Step 4: Output Lead Card
Write structured markdown to leads/{company-slug}.md
Compounds: Subagent Orchestration
For the Compounds Subagent Orchestration stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
---
name: outbound-playbook
description: Run the full outbound playbook for a target segment.
Spawns parallel agents to qualify leads and draft emails.
disable-model-invocation: true
allowed-tools: Bash Read Write Task Teammate
---
## Outbound Playbook
### Phase 1: Build Lead List
Ask the user for: target segment, company size range, geography.
Use /scrape-directory to pull matching companies.
### Phase 2: Parallel Lead Qualification (Task tool)
For each company (batch of 5):
- Spawn a subagent with qualify-lead preloaded
- Each subagent qualifies one company independently
### Phase 3: Draft Emails (Task tool)
For each qualified lead:
- Spawn a subagent with draft-email preloaded
### Phase 4: Human Review Checkpoint
STOP. Present sample drafts. Ask: "Review these. Adjust or proceed?"
Do NOT proceed without explicit user approval.
### Phase 5: Campaign Summary
Compile results to outbound/{segment}/campaign-summary.md
The Folder Structure
For the The Folder Structure stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
.claude/skills/
# ATOMS — single purpose, near-deterministic
verify-email/SKILL.md
find-email/SKILL.md
find-contact/SKILL.md
research-company/SKILL.md
scrape-url/SKILL.md
# MOLECULES - explicit chains of atoms
qualify-lead/SKILL.md
review-and-test/SKILL.md
draft-blog-post/SKILL.md
# COMPOUNDS - subagent orchestration, human-driven
outbound-playbook/SKILL.md
feature-build-ship/SKILL.md
The Same Pattern in Other Frameworks
For the The Same Pattern in stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the The Same Pattern in stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
LangGraph: Tools → Chains → Subgraphs
When working through the LangGraph Tools Chains Subgraphs stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
from langgraph.graph import StateGraph
from langchain_core.tools import tool
# ATOM: a single tool
@tool
def verify_email(email: str) -> dict:
"""Verify email deliverability via Hunter.io."""
response = requests.get(f"https://api.hunter.io/v2/email-verifier?email={email}")
return response.json()
# MOLECULE: explicit sequential graph
workflow = StateGraph(LeadState)
workflow.add_node("research", research_company)
workflow.add_node("find_contact", find_contact)
workflow.add_node("verify", verify_email)
workflow.add_edge("research", "find_contact")
workflow.add_edge("find_contact", "verify")
graph = workflow.compile()
# COMPOUND: subgraph composition
parent = StateGraph(CampaignState)
parent.add_node("qualify", qualify_subgraph) # each is a compiled graph
parent.add_node("draft", email_subgraph) # with its own state
parent.add_node("review", human_review_node)
CrewAI: Tools → Tasks → Crews within Flows
When working through the CrewAI Tools Tasks Crews stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours.
from crewai import Agent, Task, Crew, Flow
from crewai.tools import tool
# ATOM: a tool
@tool
def verify_email(email: str) -> str:
"""Verify email deliverability."""
return requests.get(f"https://api.hunter.io/v2/email-verifier?email={email}").text
# MOLECULE: tasks chained via context
researcher = Agent(role="Researcher", goal="Find company info", tools=[search_tool])
verifier = Agent(role="Verifier", goal="Verify contacts", tools=[verify_email])
research_task = Task(description="Research {company}", agent=researcher)
verify_task = Task(description="Verify the contact", agent=verifier, context=[research_task])
crew = Crew(agents=[researcher, verifier], tasks=[research_task, verify_task])
# COMPOUND: Crews inside a Flow
class OutboundFlow(Flow):
@start()
def qualify_leads(self):
return qualify_crew.kickoff(inputs={"segment": self.state.segment})
@listen(qualify_leads)
def draft_emails(self, qualified):
return email_crew.kickoff(inputs={"leads": qualified})
@listen(draft_emails)
def human_review(self, drafts):
return drafts # pause for human approval
Agno: @tool → Agent → Teams
When working through the Agno tool Agent Teams stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log tool name, args hash, latency, and outcome for every call. Debugging agent loops without that trail wastes hours. When working through the Agno tool Agent Teams stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
from agno.agent import Agent
from agno.models.anthropic import Claude
from agno.tools import tool
# ATOM
@tool
def verify_email(email: str) -> str:
"""Verify email deliverability via Hunter.io."""
return requests.get(f"https://api.hunter.io/v2/email-verifier?email={email}").text
# MOLECULE: agent with ordered tools
qualify_agent = Agent(
model=Claude(id="claude-sonnet-4-6"),
description="Qualify a lead: research company, find contact, verify email. Execute in that order.",
tools=[research_company, find_contact, find_email, verify_email],
)
# COMPOUND: team of agents
from agno.team import Team
outbound_team = Team(
agents=[qualify_agent, email_drafter, campaign_reporter],
description="Run the full outbound playbook for a target segment.",
)
The Pattern Is the Same
The The Pattern Is the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Where It Breaks Today
The Where It Breaks Today stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Atoms that aren’t solid break everything above them:
The Atoms that aren t stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Atoms that aren t stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Molecules beyond 10 atoms get unreliable:
For the Molecules beyond 10 atoms stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Compounds beyond 8–10 molecules hit their own ceiling:
For the Compounds beyond 8 10 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Auto-invocation is less reliable than explicit invocation:
For the Auto-invocation is less reliable stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Auto-invocation is less reliable stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Why This Matters
When working through the Why This Matters stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Operational checklist
For the Operational checklist stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.
Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 373492c8b420: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.