Server-Sent Events promise something WebSocket does not: when the connection drops, the browser reconnects by itself and tells the server the last event it saw, in a Last-Event-ID request header. That promise is only as good as the server behind it. A handler that ignores the header, numbers events per connection, or answers a reconnect with the wrong status code turns automatic recovery into silent data loss or a client that never comes back.
This article covers that server side, one connection at a time: what the WHATWG HTML specification makes the browser do, event IDs as cursors, the four cases a handler decides on every reconnect, the race between replay and live events, and which status codes keep clients alive. For SSE basics read Server-Sent Events in depth; for resuming across a cluster of nodes, read scaling SSE horizontally.
The contract the browser holds you to
The EventSource behaviour is defined precisely in the HTML specification, and every server design decision follows from it. The rules that matter for reconnection:
| Server does | Browser does |
|---|---|
Sends id: x | Stores x in a buffer; when the event is dispatched, x becomes the last event ID. This happens even if the event has no data, in which case no message event fires |
Sends an id containing a NULL character | Ignores the field |
Sends id: with an empty value | Resets the last event ID to empty, so the next reconnect sends no header |
Sends retry: n (ASCII digits only) | Sets the reconnection time to n milliseconds; anything else is ignored |
| Ends the response normally, or the network fails | Sets readyState to CONNECTING, fires error, waits the reconnection time, reconnects |
Replies with a status other than 200, or a Content-Type other than text/event-stream | Fails the connection: readyState becomes CLOSED and it never retries |
| Replies 204 No Content | Stops reconnecting; the spec names this as the way to tell a client to stop |
| Redirects with 301 or 307 | Follows the redirect |
On reconnect the browser adds Last-Event-ID only when its last event ID is not empty. The initial reconnection time is implementation-defined, described by the spec as probably a few seconds, and browsers may add exponential backoff after failures but are not required to. Two consequences surprise people. First, a stream that ends cleanly is not finished as far as the browser is concerned: it reconnects. Second, an error status is not a temporary failure: it is final, and the page has to create a new EventSource to recover.
The flow of one resume
Each event ends with a blank line, comments start with a colon, and an id line on its own moves the client's cursor forward without delivering a message. That last trick is useful for filtered streams: if a client only receives some events from a busy topic, sending a bare id now and then keeps its cursor near the head, so a later reconnect does not replay a long stretch of events it would discard anyway.
id: e7:1040
retry: 3000
id: e7:1041
event: order
data: {"order": 88, "status": "packed"}
: keep-alive comment, ignored by the parser, resets proxy idle timers
id: e7:1050
event: reset
data: {"reason": "cursor_expired", "snapshot": "/orders/88"}
Designing the event ID as a cursor
The id is not a label for the event; it is a cursor into a log the server can read from again. That leads to four rules.
- It must name a position in durable, ordered storage. A per-connection counter restarts at 1 on every connection and makes
Last-Event-ID: 57meaningless. Use the offset or sequence number of the log that events come from: a database sequence, a Kafka offset for a single partition, a Redis stream ID. - It must be ordered within the stream the client reads. The handler needs to answer "what comes after this?", which requires a total order per stream. If a stream merges several partitions, the cursor has to carry a position for each, or you must restrict the stream to one partition.
- It should carry an epoch. If the log is ever rebuilt, truncated or migrated, old offsets point at the wrong events. A prefix such as
e7:lets the handler recognise a cursor from a previous generation and reset instead of replaying the wrong data. - Treat it as opaque and untrusted. Clients echo whatever you sent, but anyone can send any header. Parse strictly, bound its length, and never use it to address data the caller is not authorised to read.
Timestamps make poor cursors: two events can share one, clocks move backwards, and "after 12:00:01.250" is ambiguous. Ordering guarantees in general are covered in message ordering for bidirectional systems.
The resume decision: four cases
Every new connection lands in one of four cases, and each needs a deliberate answer.
| Case | How you know | Correct response |
|---|---|---|
| No cursor | No header, no query parameter | Start at the live tail, or at a position the page supplies, such as the version of a snapshot it has just loaded |
| Valid and retained | Epoch matches, sequence at or after the oldest retained one | Replay everything after it, then continue live |
| Valid but expired | Epoch matches, but the events after it were already trimmed | Send an explicit reset event so the client reloads state, then continue live |
| Unknown or malformed | Wrong epoch, unparsable, or impossibly far ahead | Same as expired: reset, never guess |
The dangerous answer is the silent one: treating an expired cursor as "no cursor" and jumping to live. The client believes it is up to date and is not. A reset event with a pointer to a snapshot URL lets the page refetch the current state once and carry on. Size your retention window from data: the longest gap you want to heal by replay, such as a laptop closed over lunch.
Because EventSource cannot send custom headers on the first connection, accept the starting cursor as a query parameter too, and prefer the header when both are present: the header reflects what the browser actually received, while the URL holds where the page started.
The replay race, and a handler that avoids it
A naive handler reads the log from the cursor, sends those events, then subscribes to live updates. Any event published between the end of the read and the start of the subscription is lost,. The fix is to reverse the order: subscribe first, buffering live events in memory; then replay from the log; then drain the buffer, skipping anything with a sequence number already sent. Because IDs are ordered, deduplication is a single comparison.
import asyncio, json
from starlette.requests import Request
from starlette.responses import Response, StreamingResponse
EPOCH = "e7" # changes whenever the log is rebuilt and old offsets lose meaning
def parse_cursor(raw):
"""Return ("ok", seq), ("foreign", None) or ("bad", None) for a Last-Event-ID value."""
if not raw:
return ("none", None)
epoch, _, seq = raw.partition(":")
if epoch != EPOCH:
return ("foreign", None)
return ("ok", int(seq)) if seq.isdigit() else ("bad", None)
def frame(seq, event, payload):
return f"id: {EPOCH}:{seq}\nevent: {event}\ndata: {json.dumps(payload)}\n\n"
async def stream(request: Request):
if not await authorised(request):
return Response(status_code=401) # ends EventSource for good: see below
topic = request.path_params["topic"]
raw = request.headers.get("last-event-id") or request.query_params.get("after")
kind, after = parse_cursor(raw)
async def body():
nonlocal kind
live = log.subscribe(topic) # 1. subscribe first, buffer in memory
try:
head = log.head_seq(topic)
if kind == "ok" and after > head:
kind = "bad" # from the future: never trust it
replay = kind == "ok" and log.oldest_seq(topic) <= after + 1
last = after if replay else head
yield f"id: {EPOCH}:{last}\nretry: 3000\n\n" # 2. never a bare retry frame
if replay:
async for seq, ev, data in log.read_after(topic, after): # 3. replay
yield frame(seq, ev, data); last = seq
elif kind != "none": # 3. explicit reset, then live
yield frame(last, "reset", {"reason": "cursor_expired" if kind == "ok" else kind})
async for seq, ev, data in live: # 4. live, skipping what replay sent
if seq <= last:
continue
yield frame(seq, ev, data); last = seq
finally:
live.close()
return StreamingResponse(body(), media_type="text/event-stream",
headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"})The opening frame repeats the client's cursor for a reason. The spec starts every new stream with an empty ID buffer, and a blank line dispatches it, so a bare retry frame would reset the client's last event ID and its next reconnect would arrive with no header. This uses Starlette's StreamingResponse; log and authorised stand in for your event store and auth check. The subscription buffer must be bounded: if replay takes long enough for it to overflow, close the stream and let the client reconnect from the last ID it received, which is safe precisely because the cursor is durable. The X-Accel-Buffering header stops nginx buffering the response.
Status codes: which ones keep clients alive
Because any non-200 status ends an EventSource for good, choose them deliberately.
- Overload: do not answer 503 or 429. Answer 200 with the event stream content type, send a large
retrysuch as 30000 plus a random offset, and end the response. The browser treats it as a normal close and comes back later, spread out over time. - Deploys: before shutting a node down, send each connection a
retrywith a randomised value and close it. Clients then reconnect over a window, not all at once. - Expired credentials: a 401 ends the stream, so the page must listen for
error, check thatreadyStateis CLOSED, refresh its token and create a newEventSource. - Stream genuinely finished: for example, a job's progress feed after completion. Send a final
doneevent and have the page callclose(), and answer any later reconnect with 204 so stray clients stop. - Resource gone or forbidden: 404 and 403 are correct precisely because they stop retries.
Client-side visibility handling, such as closing streams in background tabs, is covered in SSE client lifecycle and visibility.
fetch-based clients: you implement the contract
Many applications, notably LLM token streaming, read SSE with fetch and a stream parser instead of EventSource, usually to send POST bodies or an Authorization header. That loses everything in the table above: no automatic reconnect, no Last-Event-ID, no retry. If you need resumption, implement it, with jittered backoff because nothing else will add it.
// fetch-based clients get none of EventSource's reconnection logic for free
async function follow(url, onEvent) {
let lastId = "", delay = 1000;
for (;;) {
try {
const res = await fetch(url, { headers: lastId ? { "Last-Event-ID": lastId } : {} });
if (res.status === 204) return; // server said: stop
if (!res.ok) throw new Error(`status ${res.status}`);
delay = 1000;
for await (const ev of parseSSE(res.body)) { // your parser: id, event, data, retry
if (ev.id !== undefined) lastId = ev.id;
if (ev.retry) delay = ev.retry;
if (ev.data !== undefined) onEvent(ev);
}
} catch (e) { /* network error or non-OK status: fall through and retry */ }
await new Promise(r => setTimeout(r, delay * (0.5 + Math.random())));
delay = Math.min(delay * 2, 30000);
}
}The same backoff reasoning as for WebSockets applies; see WebSocket reconnection strategies.
Worked example: order tracking for 40,000 open tabs
An e-commerce site streams order status. Events come from an outbox table with a sequence column; the stream ID is e3:<sequence>, per customer topic. Peak rate is 300 events per second across all customers, and the log retains 6 hours, about 6.5 million rows. A customer's laptop sleeps for 50 minutes; on wake, the browser reconnects with Last-Event-ID: e3:918220, the handler subscribes, replays the 4 events for that customer since 918220, drains 1 duplicate from the live buffer, and continues. A tab left open over a weekend reconnects with a cursor older than 6 hours, receives reset, refetches the order list once, and continues live. After a table migration the epoch becomes e4 and old cursors reset.
Failure modes
- Per-connection IDs: every reconnect replays from the wrong place or not at all.
- Read-then-subscribe: events published during the gap vanish, invisibly and only under load.
- Silent jump to live: expired cursors look like fresh clients and stale state persists.
- Error status for transient problems: a 503 during a deploy closes every affected tab's stream until reload.
- No retry jitter: thousands of clients reconnect on the same tick after a node restarts.
- Proxy buffering: a proxy holds the response until a buffer fills, so replayed events and heartbeats reach the client late or in clumps.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Long retention | More reconnects healed by replay | Storage, and slower replays for long gaps |
| Short retention plus reset | Simple, bounded storage | More snapshot reloads |
| Per-topic sequences | Small replays per client | More sequences to manage |
| Global sequence with filtering | One ordered log | Replays scan events the client discards; use bare id checkpoints |
| EventSource | Reconnect and resume built in | GET only, no custom headers |
| fetch streaming | POST, headers, binary-safe parsing | You write the reconnect logic |
What to do next
- Check what your stream's
idvalues are; if they are per-connection counters, replace them with log positions. - Add an epoch prefix and a strict cursor parser that classifies none, valid, expired and unknown.
- Reorder your handler to subscribe before replaying, and deduplicate by sequence.
- Define a
resetevent and a snapshot endpoint, and make the page handle it. - Audit every non-200 status your stream route can return; replace transient ones with 200 plus a jittered
retryand close. - Handle
readyStateCLOSED in the page for 401 and recreate the source after refreshing credentials. - Test it: kill the connection mid-stream, sleep past retention, and restart a node under load, checking for gaps and duplicates each time.