Home / Articles / Practical notes: Your Agent Graph Doesn’t Belong in Python: Compiling a

This article is published in English.

Practical notes: Your Agent Graph Doesn’t Belong in Python: Compiling a

Operable walkthrough of Practical notes: Your Agent Graph Doesn’t Belong in Python: Compiling a: contracts, checks, and drop-in code slots for teams shipping this pattern.

2170 words

The following notes reconstruct a practical path around “Your Agent Graph Doesn’t Belong in Python: Compiling a Multi-Agent Workflow from a Single YAML File”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

The problem nobody warns you about

The The problem nobody warns stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.

What the workflow looks like as data

The What the workflow looks stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.

entry: entry_agent
exit: exit
guardrails:
  - Reject queries that are outside the application's domain.
  - Reject queries about the system, agents, design, or internal workings.state_schema:
  query:
    type: str
    description: "User query or current message."
  chat_history:
    type: list
    annotated_with: add_messages
    description: "Conversation history between user and system."
  result:
    type: dict
    description: "Result from the processing agent."agents:
  - name: agent_one
    kind: function
    impl: your_package.agents.agent_one.agent_one_fn  - name: agent_two
    kind: function
    impl: your_package.agents.agent_two.agent_two_fnworkflow:
  nodes:
    - id: agent_one
      agent: agent_one
      writes: [query, result]
      next: decision_router    - id: decision_router
      kind: router
      router:
        impl: your_package.agents.routers.route_after_agent_one
        reads: [result]
        edges:
          agent_two: agent_two
          human_agent: human_agent

Trick #1: Generating your state class at runtime from a schema

The Trick 1 Generating your stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos. The Trick 1 Generating your stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

# your_package/orchestrator/schema.py
annotations = {}
for key, value in state_schema.items():
    type_str = value.get("type", "str")
    # Convert YAML string to Python type
    py_type = eval(type_str)
    if value.get("annotated_with") == "add_messages":
        py_type = Annotated[list, {}]
    annotations[key] = py_type
spec = Spec(
    ...
    state=TypedDict("State", annotations),   # <- dynamic class, born at boot
    ...
)

Trick #2: Agents referenced by dotted string, resolved by importlib

For the Trick 2 Agents referenced stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.

impl: your_package.agents.agent_one.agent_one_fn
# your_package/orchestrator/schema.py
def _import_from_path(dotted: str) -> Callable[..., Any]:
    """Import a callable from a dotted path like 'package.module.function'."""
    if not dotted or "." not in dotted:
        raise ImportError(f"Invalid impl path: {dotted!r}")
    mod_path, attr = dotted.rsplit(".", 1)
    mod = importlib.import_module(mod_path)
    fn = getattr(mod, attr)
    if not callable(fn):
        raise TypeError(f"Imported object is not callable: {dotted}")
    return fn
def agent_impl_map(spec: Spec) -> Dict[str, Optional[Callable]]:
    """Map agent name -> callable (or None if impl missing)."""
    return {a.name: _import_from_path(a.impl) if a.impl else None
            for a in spec.agents}

Trick #3: The compiler — YAML nodes become graph nodes

For the Trick 3 The compiler stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.

# your_package/orchestrator/runner.py
def build(self):
    graph = StateGraph(state_schema=self.spec.state)   # our generated TypedDict
    def _add_task_node(node):
        async def _node(state: Dict[str, Any]) -> Dict[str, Any]:
            res = await self._call_agent(node.agent, state, node.id)
            if getattr(node, "writes", None):
                if isinstance(res, dict):
                    # Only let the node write the keys it declared in YAML
                    filtered = {k: v for k, v in res.items() if k in node.writes}
                    return filtered or res
                key = node.writes[0]
                return {key: res}
            return res
        graph.add_node(node.id, _node)    # Build every node
    for node in self.spec.workflow.nodes:
        if getattr(node, "router", None):
            _add_router_node(node)
        else:
            _add_task_node(node)    graph.set_entry_point(entry)    # Inline "next:" edges from YAML become static edges
    for node in self.spec.workflow.nodes:
        if getattr(node, "next", None):
            graph.add_edge(node.id, node.next)    # Terminal nodes wire to END
    for node in self.spec.workflow.nodes:
        if getattr(node, "terminal", False):
            graph.add_edge(node.id, END)    self._runnable = graph.compile(checkpointer=self.checkpoint)
    return self

Routers: conditional branching as a lookup table

For the Routers conditional branching as stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine. For the Routers conditional branching as stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

# your_package/orchestrator/runner.py
def _add_router_node(node):
    router = self.router_fns[node.id]
    def _router_fn():
        def _f(state):
            out = router(state)
            # Routers may return either a label, or (state_updates, label)
            if isinstance(out, tuple):
                updates, label = out
                if isinstance(updates, dict):
                    for k, v in updates.items():
                        state[k] = v
            else:
                label = out
            return label
        return _f    graph.add_node(node.id, lambda s: {})
    graph.add_conditional_edges(node.id, _router_fn(), node.router.edges)

Trick #4: The adaptive call — agents write whatever signature they want

When working through the Trick 4 The adaptive stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.

# your_package/orchestrator/runner.py
async def _adapt_and_call(self, fn, state, node_id):
    """
    Adaptively call agent functions so implementations receive what they expect:
    - def agent(**kwargs):        → pass **state (+ inject 'query' if missing)
    - def agent(query, **kwargs): → pass query=..., plus any **extra
    - def agent(state):           → pass state
    - def agent(query):           → pass query
    - def agent():                → call without args
    """
    sig = inspect.signature(fn)
    params = sig.parameters
    has_var_kw = any(p.kind == inspect.Parameter.VAR_KEYWORD
                     for p in params.values())
    kwargs = {}
    if has_var_kw:
        kwargs.update(state)
    if "state" in params:
        kwargs["state"] = state
    if "query" in params or has_var_kw:
        kwargs.setdefault("query", self._fallback_query(state))    # A lone positional 'query' → call it positionally
    if (len(params) == 1
            and next(iter(params.keys())) == "query"):
        return await _maybe_await(fn(self._fallback_query(state)))    res = fn(**kwargs)
    return await res if hasattr(res, "__await__") else res

Trick #5: Hot-swapping a node per session (human-in-the-loop)

When working through the Trick 5 Hot-swapping a stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.

# your_package/services/session_service.py (paraphrased)
if websocket is not None:
    session_handler = SessionHandler(websocket, user_id=user_id, session_id=session_id, ...)
    _runner.agent_fns["human_agent"] = _import_from_function(
        make_input_method(session_handler)
    )

What this architecture actually buys you

When working through the What this architecture actually stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs. When working through the What this architecture actually stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

The takeaway

The The takeaway stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.

Operational checklist

For the Operational checklist stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.

Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.

Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.

Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for 2822ea5988ca: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.