Home / Articles / Resumable LLM streaming with Redis Streams as a per-turn rendezvous

This article is published in English.

Resumable LLM streaming with Redis Streams as a per-turn rendezvous

Survive disconnects and multi-minute tool pauses by publishing agent events to a turn-scoped Redis Stream clients can resume.

2049 words

What we wanted

Agent answers can take tens of seconds: reason, load a skill, call tools, wait, stream tokens. Demos keep one open connection. Production drops mid-answer, redeploys gateways, and pauses for tools that return minutes later. The goal: resumable real-time streaming so a client can reconnect and continue the same turn.

The journey

Iteration 1: client talks to the agent directly

Simple and fragile. Any network blip ends the stream. Horizontal scale means sticky sessions or lost events.

Iteration 2: gRPC streaming between services

Better internal contracts, still awkward for browser clients and still weak at multi-minute pauses across pod restarts.

Why not Kafka?

Great for durable logs; heavier than needed for a per-turn rendezvous with short retention and consumer groups that map poorly to “one browser tab.”

Iteration 3 (what worked): a named rendezvous per turn

// Agent output
{
  "type": "tool_result",
  "tool": "product_search",
  "data": {
    "items": [...]
  }
}

// Gateway -> TV
{
  "type": "product_carousel",
  "items": [...]
}

// Gateway -> Mobile
{
  "type": "product_list",
  "items": [...]
}
cursor = last_event_id or "0-0"

while True:
    entries = xread({key: cursor}, block=30_000)

    if not entries:          # the only timeout check point
        check_timeouts()
        continue

    for entry_id, event in entries:
        # writes to the socket; not an ack that the client received it
        sse.send(id=entry_id, data=event.payload)
        cursor = entry_id
        if event.type in TERMINAL:
            return
id: 1755600000123-0
data: {"type":"tool_selected","tool":"search"}

id: 1755600000871-0
data: {"type":"response_block","block":{...}}
GET /sessions/{sid}/turns/{tid}/stream
Last-Event-ID: 1755600000871-0
# turn starts: one atomic step (MULTI/EXEC, or a Lua script)
xadd(key, first_event)
expire(key, GENEROUS_TTL)

# producer finishes: bring it in
expire(key, RECONNECT_TTL)

Each user turn gets a Redis Stream (or stream+consumer group pattern) keyed by turn_id. The agent publishes token/tool events; the gateway tails from the client’s last id. Reconnect resumes from that cursor. Pods can die; the stream retains enough history for the turn.

Details that mattered

  • Cursor discipline — clients ACK the last seen stream ID; never restart from 0 after first connect.
  • Heartbeat events — keep intermediaries from closing idle connections during tool waits.
  • Turn lifecycle — explicit turn_started / turn_paused / turn_completed / turn_failed.
  • TTL — expire streams after the turn settles so Redis does not become an infinite archive.
  • Authz — turn_id is not authorization; bind turns to the authenticated session.

The case that settled it: a multi-minute pause

A tool handed work to another pipeline that replied minutes later. Direct HTTP streaming died. With Redis Streams, the agent published a pause event, the client showed “still working,” and later tokens resumed on reconnect without re-running the whole plan.

What we send to the client

Typed events: tokens, tool_start, tool_result summaries (never secrets), errors, and completion. Keep payloads small; stash bulky artifacts in object storage and send references.

Costs, limits, watch-outs

Watch Redis memory, max stream length, and fan-out if many gateways tail one turn. Cap concurrent turns per user. Load-test reconnect storms after a gateway deploy.

The end state

Gateway is a resume-aware reader; the agent is a writer; Redis Streams is the rendezvous. Real-time UX survives the boring failures that kill demo architectures.

Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.

Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.

Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.

Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.

Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.

Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.

Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.

Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.

Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.

Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.

Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.

Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.

Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.

Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.

Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.

Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.

Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.

Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.

Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.

Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.

Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.

Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.