This article is published in English.
Resumable LLM streaming with Redis Streams as a per-turn rendezvous
Survive disconnects and multi-minute tool pauses by publishing agent events to a turn-scoped Redis Stream clients can resume.
What we wanted
Agent answers can take tens of seconds: reason, load a skill, call tools, wait, stream tokens. Demos keep one open connection. Production drops mid-answer, redeploys gateways, and pauses for tools that return minutes later. The goal: resumable real-time streaming so a client can reconnect and continue the same turn.
The journey
Iteration 1: client talks to the agent directly
Simple and fragile. Any network blip ends the stream. Horizontal scale means sticky sessions or lost events.
Iteration 2: gRPC streaming between services
Better internal contracts, still awkward for browser clients and still weak at multi-minute pauses across pod restarts.
Why not Kafka?
Great for durable logs; heavier than needed for a per-turn rendezvous with short retention and consumer groups that map poorly to “one browser tab.”
Iteration 3 (what worked): a named rendezvous per turn
// Agent output
{
"type": "tool_result",
"tool": "product_search",
"data": {
"items": [...]
}
}
// Gateway -> TV
{
"type": "product_carousel",
"items": [...]
}
// Gateway -> Mobile
{
"type": "product_list",
"items": [...]
}
cursor = last_event_id or "0-0"
while True:
entries = xread({key: cursor}, block=30_000)
if not entries: # the only timeout check point
check_timeouts()
continue
for entry_id, event in entries:
# writes to the socket; not an ack that the client received it
sse.send(id=entry_id, data=event.payload)
cursor = entry_id
if event.type in TERMINAL:
return
id: 1755600000123-0
data: {"type":"tool_selected","tool":"search"}
id: 1755600000871-0
data: {"type":"response_block","block":{...}}
GET /sessions/{sid}/turns/{tid}/stream
Last-Event-ID: 1755600000871-0
# turn starts: one atomic step (MULTI/EXEC, or a Lua script)
xadd(key, first_event)
expire(key, GENEROUS_TTL)
# producer finishes: bring it in
expire(key, RECONNECT_TTL)
Each user turn gets a Redis Stream (or stream+consumer group pattern) keyed by turn_id. The agent publishes token/tool events; the gateway tails from the client’s last id. Reconnect resumes from that cursor. Pods can die; the stream retains enough history for the turn.
Details that mattered
- Cursor discipline — clients ACK the last seen stream ID; never restart from
0after first connect. - Heartbeat events — keep intermediaries from closing idle connections during tool waits.
- Turn lifecycle — explicit
turn_started/turn_paused/turn_completed/turn_failed. - TTL — expire streams after the turn settles so Redis does not become an infinite archive.
- Authz —
turn_idis not authorization; bind turns to the authenticated session.
The case that settled it: a multi-minute pause
A tool handed work to another pipeline that replied minutes later. Direct HTTP streaming died. With Redis Streams, the agent published a pause event, the client showed “still working,” and later tokens resumed on reconnect without re-running the whole plan.
What we send to the client
Typed events: tokens, tool_start, tool_result summaries (never secrets), errors, and completion. Keep payloads small; stash bulky artifacts in object storage and send references.
Costs, limits, watch-outs
Watch Redis memory, max stream length, and fan-out if many gateways tail one turn. Cap concurrent turns per user. Load-test reconnect storms after a gateway deploy.
The end state
Gateway is a resume-aware reader; the agent is a writer; Redis Streams is the rendezvous. Real-time UX survives the boring failures that kill demo architectures.
Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.
Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.
Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.
Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.
Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.
Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.
Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.
Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.
Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.
Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.
Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.
Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.
Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.
Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.
Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.
Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.
Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.
Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.
Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.
Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.
Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.
Operational tip: store the last stream ID in the session cookie or client memory and also server-side for support replay. When a user says “it froze,” support should reconstruct the turn from the stream without asking them to reproduce a ten-minute tool wait. Add dashboards for resume rate, abandoned turns, and average pause duration so product can see whether agents are slow or networks are flaky.