Home / Articles / Practical notes: Fixing Non-Deterministic Replay Bugs in AI Agent Loops

This article is published in English.

Practical notes: Fixing Non-Deterministic Replay Bugs in AI Agent Loops

Operable walkthrough of Practical notes: Fixing Non-Deterministic Replay Bugs in AI Agent Loops: contracts, checks, and drop-in code slots for teams shipping this pattern.

3563 words

The following notes reconstruct a practical path around “Fixing Non-Deterministic Replay Bugs in AI Agent Loops”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

The bug, before the names

The The bug before the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

# ❌ What I almost wrote — LLM call INSIDE the workflow.
@workflow.defn
class SourcingWorkflow:
    @workflow.run
    async def run(self, brief: SourcingBriefInput) -> SourcingResult:
        suppliers = await workflow.execute_activity(research_activity, brief)
        scored = await workflow.execute_activity(score_activity, suppliers)

        # The LLM call — right here, in the workflow. THIS IS THE BUG.
        # If worker dies here, replay will re-call the LLM and mutate state/history.
        response = await llm_client.complete(
            messages=build_decide_prompt(scored),
            temperature=0.0,
        )
        selected = parse_llm_decision(response.content, scored)

        approval = await workflow.wait_condition(...)
        # ...create PO, initiate payment
# ✅ The fix — LLM call moved into an activity.
@activity.defn
async def decide_activity(scored: list[dict], brief: SourcingBriefInput) -> dict:
    """The LLM call lives here — in the activity, not the workflow."""
    from app.agentmesh.llm import get_llm_client

    client = get_llm_client()
    response = await client.complete(
        messages=build_decide_prompt(scored),
        temperature=0.0,
    )
    selected, rationale = parse_llm_decision(response.content, scored)
    return {"selected_supplier": selected, "decision_reason": rationale}


@workflow.defn
class SourcingWorkflow:
    @workflow.run
    async def run(self, brief: SourcingBriefInput) -> SourcingResult:
        suppliers = await workflow.execute_activity(research_activity, brief)
        scored = await workflow.execute_activity(score_activity, suppliers)

        # The LLM call is NOWHERE in the workflow.
        # The activity result is recorded. On replay, it's injected.
        decision = await workflow.execute_activity(
            decide_activity,
            args=(scored, brief),
        )

The law: replay must produce the same history

The The law replay must stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

The violation: what happens when the LLM is in the workflow

The The violation what happens stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices. The The violation what happens stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Why “temperature=0.0” doesn’t save you

For the Why temperature 0 0 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

The boundary: the activity is the non-determinism membrane

For the The boundary the activity stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

The code: how this actually looks in AgentMesh

For the The code how this stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the The code how this stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Temporal Workflow (deterministic — replayed)
    └── Activity: run_graph_until_interrupt (non-deterministic — recorded once)
            └── LangGraph StateGraph
                    └── decide_node (async function)
                            └── get_llm_client().complete()  ← the LLM call

The workflow — pure orchestration (workflow.py)

When working through the The workflow pure orchestration stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

@workflow.defn
class SourcingWorkflow:
    @workflow.run
    async def run(self, brief: SourcingBriefInput) -> SourcingResult:
        # ↓ This is the boundary. The activity runs once. Result is recorded.
        graph_result = await workflow.execute_activity(run_graph_until_interrupt, brief, ...)

        # Wait for human signal — a Temporal primitive, not an LLM call.
        # On replay, the signal is injected from history.
        await workflow.wait_condition(lambda: self._approval_received, timeout=timedelta(hours=24))

        # ↓ Another boundary. Resume activity runs once. Result is recorded.
        resume_result = await workflow.execute_activity(resume_graph, self._approval_data, ...)

        # ↓ Side effects — each is its own activity with its own retry policy.
        po_result = await workflow.execute_activity(create_po_activity, args=(...), ...)

The activity — where the non-determinism lives (activity.py)

When working through the The activity where the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

@activity.defn
async def run_graph_until_interrupt(brief: SourcingBriefInput) -> dict:
    graph = build_sourcing_graph(checkpointer=await get_checkpointer())
    config = {"configurable": {"thread_id": activity.info().workflow_id}}

    # ↓ Everything inside this call is non-deterministic. It runs ONCE.
    #   The result dict is recorded in the event history. On replay, it's injected.
    final_state = await graph.ainvoke({"brief": brief, "suppliers": [], "attempts": 0, ...}, config)

    return {"paused": True, "selected_supplier": selected, "suppliers": final_state.get("suppliers", [])}

The graph node — where the LLM actually fires (graph.py)

When working through the The graph node where stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the The graph node where stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

async def decide_node(state: AgentState) -> dict:
    scored = state.get("scored_suppliers", [])
    messages = build_decide_prompt(brief.item, brief.quantity, brief.budget, scored, past_decisions)

    # ↓ THE LLM CALL. This is the non-determinism that must never be in the workflow.
    response = await get_llm_client().complete(messages=messages, temperature=0.0, max_tokens=1000)

    selected, rationale = parse_llm_decision(response.content, scored)
    return {"selected_supplier": selected, "decision_reason": rationale, "cost_incurred": response.cost_usd}    # ↓ THE LLM CALL. This is the non-determinism that must never be in the workflow.
    response = await get_llm_client().complete(messages=messages, temperature=0.0, max_tokens=1000)

When the boundary gets blurry

The When the boundary gets stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Temporal event history             LangGraph checkpoint (Postgres)
  └── ActivityCompleted(result)      └── graph state at interrupt()
        paused: True                       node: "approve"
        selected: SupplierB                 selected: SupplierB
        suppliers: [A, B, C]                suppliers: [A, B, C]

The Isolation Pattern — same constraint everywhere

The The Isolation Pattern same stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

The same constraint in other agent stacks

The The same constraint in stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The The same constraint in stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

The cost: what you give up for determinism

For the The cost what you stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Edge case 1: Long-running activities and the heartbeat problem

For the Edge case 1 Long-running stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

@activity.defn
async def run_graph_until_interrupt(brief: SourcingBriefInput) -> dict:
    # Long-running: graph may run 5+ minutes with multiple LLM calls
    for node in graph.stream(initial_state, config):
        activity.heartbeat()  # ← "I'm alive, don't timeout me"
        # ... process node output

Edge case 2: Streaming LLM tokens back through Temporal

For the Edge case 2 Streaming stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call. For the Edge case 2 Streaming stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Edge case 3: Activity retry policies — not all activities should retry the same way

When working through the Edge case 3 Activity stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

# Side effect — strict, non-retryable. Idempotency key handles safety.
po_result = await workflow.execute_activity(
    create_po_activity,
    args=(supplier, item, quantity, price),
    retry_policy=workflow.RetryPolicy(
        initial_interval=timedelta(seconds=1),
        maximum_attempts=1,  # ← don't retry. Idempotency key prevents duplicates.
        non_retryable_error_types=["DuplicatePOError"],
    ),
)

# Verification — aggressive retry, but with backoff for eventual consistency
po_verification = await workflow.execute_activity(
    verify_po_exists,
    args=(po_result["po_id"],),
    retry_policy=workflow.RetryPolicy(
        initial_interval=timedelta(seconds=2),  # ← give the DB time to sync
        maximum_attempts=5,
    ),
)

The mental model

When working through the The mental model stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

What’s coming in the next post

When working through the What s coming in stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the What s coming in stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

References

The References stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Operational checklist

For the Operational checklist stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.

Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.

Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for 11b06b04feeb: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.

When working through the hardening note 0 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Hardening detail 0/898: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 1 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Hardening detail 1/898: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 2 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Hardening detail 2/898: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

When working through the hardening note 3 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.

Hardening detail 3/898: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 4 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Hardening detail 4/898: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 5 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Hardening detail 5/898: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

The hardening note 0 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Hardening detail 0/917: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.

For the hardening note 1 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Hardening detail 1/917: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.