Home / Articles / Practical notes: Your AI Agent Framework Is Probably the Wrong One Here’s How

This article is published in English.

Practical notes: Your AI Agent Framework Is Probably the Wrong One Here’s How

Operable walkthrough of Practical notes: Your AI Agent Framework Is Probably the Wrong One Here’s How: contracts, checks, and drop-in code slots for teams shipping this pattern.

2234 words

This walkthrough rebuilds the path from raw materials to a working system for: Your AI Agent Framework Is Probably the Wrong One Here’s How to Actually Choose. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

The question everyone asks backwards

When working through the The question everyone asks stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Axis 1: How deterministic does your branching need to be?

When working through the Axis 1 How deterministic stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

# A branch where non-determinism is FINE — picking a tone for a summary email.
# If the agent occasionally phrases things slightly differently, nobody's paged.
def draft_summary_tone(context: dict) -> str:
    return llm_call(
        prompt=f"Summarize this incident in a {context['audience']}-appropriate tone.",
        temperature=0.7,  # variability here is a feature, not a bug
    )
# A branch where non-determinism is NOT fine — deciding whether to page a human
# at 4am versus auto-remediating. This must be code, not a prompt.
def route_alert(alert: dict) -> str:
    if alert["severity"] == "critical" and alert["service"] in PAGE_ALWAYS_SERVICES:
        return "page_oncall"
    if alert["auto_remediation_available"] and alert["confidence"] > 0.9:
        return "auto_remediate"
    if alert["severity"] == "critical":
        return "page_oncall"
    return "log_and_monitor"

Axis 2: How long does one unit of work live?

When working through the Axis 2 How long stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

# Short-lived: starts and finishes inside one HTTP request.
# This is the "no framework needed" zone — a framework here is pure overhead.
async def handle_summarize_request(request: SummarizeRequest) -> SummarizeResponse:
    text = await fetch_document(request.doc_id)
    summary = await llm_summarize(text, max_tokens=300)
    return SummarizeResponse(summary=summary)
# Long-lived: this alert might sit in "awaiting human ack" for six hours
# while the on-call engineer is asleep, then resume on a completely
# different process after a deploy rotated the pods underneath it.
class AlertTriageWorkflow:
    async def run(self, alert: dict) -> dict:
        decision = await self.classify_and_route(alert)
        if decision == "page_oncall":
            await self.page(alert)
            await self.wait_for_ack(timeout_hours=1)  # this line is the whole ballgame
        ...

When working through the Axis 2 How long stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Axis 3: What happens if a step runs twice?

The Axis 3 What happens stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

# BEFORE — looks fine in a demo, is a live incident waiting to happen
async def auto_remediate(alert: dict):
    await restart_service(alert["service"])  # what if this activity gets retried?
# AFTER — idempotent by construction
async def auto_remediate(alert: dict, idempotency_key: str):
    if await remediation_ledger.already_applied(idempotency_key):
        logger.info("remediation already applied, skipping", key=idempotency_key)
        return await remediation_ledger.get_result(idempotency_key)
    result = await restart_service(alert["service"])
    await remediation_ledger.record(idempotency_key, result)
    return result

Axis 4: Who needs to read the decision later, and in what form?

The Axis 4 Who needs stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

# A framework-agnostic audit record — this is what actually matters
# in a postmortem, regardless of what orchestrated the steps.
@dataclass
class DecisionRecord:
    alert_id: str
    timestamp: float
    step: str
    reasoning: str        # what the LLM said, verbatim
    decision: str         # the structured outcome, not prose
    confidence: float | None
    human_override: bool

async def log_decision(record: DecisionRecord):
    await audit_store.insert(record)
    # Also emit as a structured log line — cheap insurance for when
    # the audit store itself is the thing that's down during an incident.
    logger.info("agent_decision", **asdict(record))

Axis 5: What’s your actual team velocity constraint?

The Axis 5 What s stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Axis 5 What s stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

# Week-one prototype: prove the concept fast, accept the debt knowingly.
from crewai import Agent, Task, Crew

triage_agent = Agent(role="Alert Triage", goal="Decide how to handle infra alerts")
crew = Crew(agents=[triage_agent], tasks=[Task(description="Triage: {alert}", agent=triage_agent)])
crew.kickoff(inputs={"alert": alert_payload})

Axis 6: What’s your latency and cost budget, per decision?

For the Axis 6 What s stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

# Expensive pattern: every routing decision is its own LLM call,
# multiplied across a multi-agent conversation with several turns.
# At alert volumes (hundreds/day, sometimes bursts of thousands during
# a real incident), this is a real line item, not a rounding error.
async def route_via_llm(alert: dict) -> str:
    return await llm_call(f"How should we handle this alert? {alert}")

# Cheaper, faster, and more auditable: cheap deterministic pre-filtering
# in code, LLM reserved for genuinely ambiguous cases.
async def route_alert_efficiently(alert: dict) -> str:
    if alert["service"] in KNOWN_NOISY_SERVICES and alert["severity"] == "low":
        return "log_and_monitor"          # zero LLM calls for the common case
    if alert["signature"] in KNOWN_REMEDIATION_PLAYBOOK:
        return "auto_remediate"           # deterministic lookup, zero LLM calls
    return await llm_call(f"Novel alert, needs judgment: {alert}")  # LLM only when genuinely needed

Putting it together: a decision path, not a decision tree

For the Putting it together a stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Is this unit of work stateless and finishes in seconds?
  └─ YES → skip the framework entirely. Plain functions + retries. Ship it.
  └─ NO, continue.

Does it need to survive process restarts / wait on humans for hours-to-days?
  └─ YES → you need durable execution (Temporal or equivalent) as the backbone,
           regardless of what else you pick for the reasoning layer.
  └─ NO, continue.

Are the important branches safety- or compliance-critical
(money, infra changes, irreversible external actions)?
  └─ YES → LangGraph-style explicit graphs, keep LLM scoped to narrow nodes.
  └─ NO, mostly exploratory/creative → CrewAI or AutoGen are legitimate defaults.

Is this still a prototype whose findings might get thrown away?
  └─ YES → optimize for speed of iteration over long-term correctness,
           but write down when you'll revisit that tradeoff.
@activity.defn
async def classify_alert_activity(alert: dict) -> dict:
    # LangGraph-style graph runs here — bounded reasoning, deterministic routing —
    # inside an activity Temporal will retry and time-box like any other side effect.
    result = alert_triage_graph.invoke({"alert": alert, "audit_log": []})
    return {"decision": result["decision"], "confidence": result["confidence"]}

@workflow.defn
class AlertTriageWorkflow:
    def __init__(self):
        self._acked = False

    @workflow.signal
    async def acknowledge(self):
        self._acked = True

    @workflow.run
    async def run(self, alert: dict) -> dict:
        classification = await workflow.execute_activity(
            classify_alert_activity, alert,
            start_to_close_timeout=timedelta(seconds=20),
            retry_policy=workflow.RetryPolicy(maximum_attempts=3),
        )
        if classification["decision"] == "page_oncall":
            await workflow.execute_activity(page_oncall, alert, start_to_close_timeout=timedelta(seconds=10))
            await workflow.wait_condition(lambda: self._acked, timeout=timedelta(hours=1))
            if not self._acked:
                await workflow.execute_activity(escalate_to_secondary, alert, start_to_close_timeout=timedelta(seconds=10))
        elif classification["decision"] == "auto_remediate":
            await workflow.execute_activity(
                auto_remediate, alert, f"remediate-{alert['id']}",
                start_to_close_timeout=timedelta(minutes=2),
                retry_policy=workflow.RetryPolicy(maximum_attempts=2),
            )
        return {"alert_id": alert["id"], "decision": classification["decision"]}

Common mistakes you keep seeing

For the Common mistakes you keep stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Common mistakes you keep stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

The actual answer

When working through the The actual answer stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Operational checklist

The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.

Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.

Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for 72c003459fd6: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.