This article is published in English.
Practical notes: Your Agent Graph Doesn’t Belong in Python: Compiling a
Operable walkthrough of Practical notes: Your Agent Graph Doesn’t Belong in Python: Compiling a: contracts, checks, and drop-in code slots for teams shipping this pattern.
The following notes reconstruct a practical path around “Your Agent Graph Doesn’t Belong in Python: Compiling a Multi-Agent Workflow from a Single YAML File”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
The problem nobody warns you about
The The problem nobody warns stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.
What the workflow looks like as data
The What the workflow looks stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.
entry: entry_agent
exit: exit
guardrails:
- Reject queries that are outside the application's domain.
- Reject queries about the system, agents, design, or internal workings.state_schema:
query:
type: str
description: "User query or current message."
chat_history:
type: list
annotated_with: add_messages
description: "Conversation history between user and system."
result:
type: dict
description: "Result from the processing agent."agents:
- name: agent_one
kind: function
impl: your_package.agents.agent_one.agent_one_fn - name: agent_two
kind: function
impl: your_package.agents.agent_two.agent_two_fnworkflow:
nodes:
- id: agent_one
agent: agent_one
writes: [query, result]
next: decision_router - id: decision_router
kind: router
router:
impl: your_package.agents.routers.route_after_agent_one
reads: [result]
edges:
agent_two: agent_two
human_agent: human_agent
Trick #1: Generating your state class at runtime from a schema
The Trick 1 Generating your stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos. The Trick 1 Generating your stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
# your_package/orchestrator/schema.py
annotations = {}
for key, value in state_schema.items():
type_str = value.get("type", "str")
# Convert YAML string to Python type
py_type = eval(type_str)
if value.get("annotated_with") == "add_messages":
py_type = Annotated[list, {}]
annotations[key] = py_type
spec = Spec(
...
state=TypedDict("State", annotations), # <- dynamic class, born at boot
...
)
Trick #2: Agents referenced by dotted string, resolved by importlib
For the Trick 2 Agents referenced stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.
impl: your_package.agents.agent_one.agent_one_fn
# your_package/orchestrator/schema.py
def _import_from_path(dotted: str) -> Callable[..., Any]:
"""Import a callable from a dotted path like 'package.module.function'."""
if not dotted or "." not in dotted:
raise ImportError(f"Invalid impl path: {dotted!r}")
mod_path, attr = dotted.rsplit(".", 1)
mod = importlib.import_module(mod_path)
fn = getattr(mod, attr)
if not callable(fn):
raise TypeError(f"Imported object is not callable: {dotted}")
return fn
def agent_impl_map(spec: Spec) -> Dict[str, Optional[Callable]]:
"""Map agent name -> callable (or None if impl missing)."""
return {a.name: _import_from_path(a.impl) if a.impl else None
for a in spec.agents}
Trick #3: The compiler — YAML nodes become graph nodes
For the Trick 3 The compiler stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.
# your_package/orchestrator/runner.py
def build(self):
graph = StateGraph(state_schema=self.spec.state) # our generated TypedDict
def _add_task_node(node):
async def _node(state: Dict[str, Any]) -> Dict[str, Any]:
res = await self._call_agent(node.agent, state, node.id)
if getattr(node, "writes", None):
if isinstance(res, dict):
# Only let the node write the keys it declared in YAML
filtered = {k: v for k, v in res.items() if k in node.writes}
return filtered or res
key = node.writes[0]
return {key: res}
return res
graph.add_node(node.id, _node) # Build every node
for node in self.spec.workflow.nodes:
if getattr(node, "router", None):
_add_router_node(node)
else:
_add_task_node(node) graph.set_entry_point(entry) # Inline "next:" edges from YAML become static edges
for node in self.spec.workflow.nodes:
if getattr(node, "next", None):
graph.add_edge(node.id, node.next) # Terminal nodes wire to END
for node in self.spec.workflow.nodes:
if getattr(node, "terminal", False):
graph.add_edge(node.id, END) self._runnable = graph.compile(checkpointer=self.checkpoint)
return self
Routers: conditional branching as a lookup table
For the Routers conditional branching as stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine. For the Routers conditional branching as stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
# your_package/orchestrator/runner.py
def _add_router_node(node):
router = self.router_fns[node.id]
def _router_fn():
def _f(state):
out = router(state)
# Routers may return either a label, or (state_updates, label)
if isinstance(out, tuple):
updates, label = out
if isinstance(updates, dict):
for k, v in updates.items():
state[k] = v
else:
label = out
return label
return _f graph.add_node(node.id, lambda s: {})
graph.add_conditional_edges(node.id, _router_fn(), node.router.edges)
Trick #4: The adaptive call — agents write whatever signature they want
When working through the Trick 4 The adaptive stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.
# your_package/orchestrator/runner.py
async def _adapt_and_call(self, fn, state, node_id):
"""
Adaptively call agent functions so implementations receive what they expect:
- def agent(**kwargs): → pass **state (+ inject 'query' if missing)
- def agent(query, **kwargs): → pass query=..., plus any **extra
- def agent(state): → pass state
- def agent(query): → pass query
- def agent(): → call without args
"""
sig = inspect.signature(fn)
params = sig.parameters
has_var_kw = any(p.kind == inspect.Parameter.VAR_KEYWORD
for p in params.values())
kwargs = {}
if has_var_kw:
kwargs.update(state)
if "state" in params:
kwargs["state"] = state
if "query" in params or has_var_kw:
kwargs.setdefault("query", self._fallback_query(state)) # A lone positional 'query' → call it positionally
if (len(params) == 1
and next(iter(params.keys())) == "query"):
return await _maybe_await(fn(self._fallback_query(state))) res = fn(**kwargs)
return await res if hasattr(res, "__await__") else res
Trick #5: Hot-swapping a node per session (human-in-the-loop)
When working through the Trick 5 Hot-swapping a stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs.
# your_package/services/session_service.py (paraphrased)
if websocket is not None:
session_handler = SessionHandler(websocket, user_id=user_id, session_id=session_id, ...)
_runner.agent_fns["human_agent"] = _import_from_function(
make_input_method(session_handler)
)
What this architecture actually buys you
When working through the What this architecture actually stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Log request id, model id, and latency on every call. Without that trail, intermittent provider errors look like application bugs. When working through the What this architecture actually stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
The takeaway
The The takeaway stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.
Operational checklist
For the Operational checklist stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.
Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Pin dependency versions and record the image digest that ran the demo. Reproducibility beats tribal knowledge.
Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 2822ea5988ca: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.