The event stream is the API

ADK is event-sourced from the ground up. A run is a sequence of immutable Event objects: the user message, each model response, each function call the model emits, each function response your tools return, each state mutation, each control signal. The runner yields them as they happen and appends the durable ones to the session. Nothing about that shape is streaming-specific — it is how a plain turn-based agent works too.

What streaming adds is subdivision. In the default non-streaming mode, a model response arrives as one event whose content is the complete message: you wait for the last token before you see the first. Flip the run into a streaming mode and the same logical response arrives as many events, each carrying a fragment of text, with a final aggregated event closing the sequence. Your handler code barely changes — it is still async for event in ... — but the loop body now runs dozens of times per response instead of once, and each iteration has to decide whether it is looking at a draft or a fact. Almost every streaming bug in ADK is a failure to make that decision correctly.

Advertisement

Anatomy of an event — the fields that matter on the wire

An ADK Event is deliberately fat: it is a log record, not a chat message. A handful of its fields do all the work when you are relaying a stream.

FieldWhat it tells you
authorWho produced it — user, or the agent’s name (essential once sub-agents are involved)
content.partsThe payload: text, inline media, a function call, or a function response
partialTrue for an incremental chunk — a draft, not the record
turn_completeThe agent has yielded the floor (live mode)
interruptedGeneration was cut off, typically by barge-in
invocation_id / idWhich run this belongs to, and a unique id for this event
actionsSide effects: state deltas, artifact saves, transfer and escalation signals

Two conveniences save you from parsing parts by hand: get_function_calls() and get_function_responses() pull tool traffic out of the parts list, and is_final_response() answers ‘is this the user-visible answer?’ — which is emphatically not the same question as ‘is this the last event?’, because state-only and tool events keep arriving around it.

Advertisement

Turning it on — RunConfig, StreamingMode, and the two modes

Streaming is a property of the run, not of the agent. You pass a RunConfig whose streaming_mode selects the granularity, and the same agent definition serves a batch job, a typing chat box, and a voice call unchanged.

from google.adk.agents.run_config import RunConfig, StreamingMode

async for event in runner.run_async(
        user_id=user_id, session_id=session.id,
        new_message=message,
        run_config=RunConfig(streaming_mode=StreamingMode.SSE)):
    if event.partial and event.content:
        yield delta_frame(event)      # forward the token chunk
    elif event.is_final_response():
        yield done_frame(event)       # close the message

NONE is the default: whole messages, one event each. SSE is server-sent-events granularity — the model streams tokens down and you relay them, but the request direction is still one message in, one response out. BIDI is the live mode reached through run_live, where the upstream channel stays open too.

Partials, aggregates, and the double-render bug

Here is the trap that catches everyone exactly once. In streaming mode the model emits a run of events with partial=True, each carrying a fragment — "The", " order", " shipped" — and then a final, non-partial event carrying the whole assembled message. A client that naively appends event.content.parts[0].text for every event renders the sentence twice: once letter by letter, then again in full.

The rule is simple and worth writing on the wall: partial events are deltas to append; the non-partial event is a replacement, not an addition. Treat the aggregated event as the authoritative version of the message and overwrite your accumulated buffer with it, or ignore its text entirely and use it purely as an end-of-message marker. Either discipline works; mixing them does not. The same rule governs the server side — partial events exist for the transport, and it is the aggregated, non-partial events that the runner commits to the session as the durable record, so a crash mid-stream loses a half-typed sentence rather than corrupting the history with forty fragmentary turns.

run_live and the two pumps — bidirectional streaming

run_async streams in one direction: you hand it a message, it hands you events. run_live opens both directions at once. Alongside the downstream event iterator it takes a LiveRequestQueue, an upstream pipe you push into for as long as the session lasts — realtime media frames on one path, text and control signals on another. The model consumes your input continuously instead of waiting for a completed turn.

The structural consequence is two concurrent, never-blocking pumps. One task drains the client transport and pushes into the queue; a second iterates the event stream and forwards frames back to the client; asyncio.gather runs both. Neither may await the other, because the entire point is that the user can speak or type while the agent is still producing. That symmetry is what makes interruption possible at all. Everything else in a live session — resumption handles that restore context after a dropped socket, streaming tools that feed the model a continuous signal — hangs off this loop.

ADK streaming - live bidi sessions: audio in, audio out, tools mid-streamfrom request/response to conversationClient (mic/cam)WebSocket / WebRTCrun_live loopbidi runner modeLiveRequestQueueupstream framesLive model sessionGemini Live APIStreaming eventspartial text + audioInterruptionbarge-in handlingTool calls mid-streamasync function execSession resumptionreconnect + contextTranscriptioninput + output textStreaming toolsvideo feeds, live dataOps - latency budget + connection lifecycle + cost per minuteframeseventsforwardstreamtextcancelfeedoperateoperate
Client frames flow up a LiveRequestQueue; model events stream down; interruption cuts across.

The wire — SSE, WebSocket, and your own envelope

ADK gives you events; it does not choose your transport. Server-sent events is the natural fit for StreamingMode.SSE: it is one-way, text-framed, survives proxies that mangle WebSockets, and reconnects on its own. WebSocket is required for BIDI, because SSE has no upstream channel at all. Plain HTTP with a chunked body works but gives you no framing, so you end up reinventing SSE badly.

Whichever you pick, do not serialize the raw event. Events are internal records carrying state deltas, agent names, and tool arguments you may not want in a browser. Define a small envelope and project events onto it:

{"type": "delta", "msg": "e_7f2", "text": " shipped"}
{"type": "tool",  "msg": "e_7f2", "name": "lookup_order", "state": "running"}
{"type": "done",  "msg": "e_7f2", "text": "The order shipped Tuesday."}
{"type": "flush", "msg": "e_7f2"}

Four frame types cover the whole protocol. The msg id is what lets a client with several concurrent agent messages in flight route each delta to the right bubble instead of concatenating them into nonsense.

Client-side reassembly without lying to the user

The client is a small state machine over that envelope: a map from msg id to an accumulating buffer. A delta appends; a done replaces and seals; a flush truncates. Three details separate a reassembler that works from one that mostly works.

Order is per-connection, not global. Deltas arrive in order on one socket, but if you reconnect and replay, or run two agents concurrently, you can receive a delta for a message you already sealed. Ignore deltas for sealed ids rather than resurrecting them. Rendering must be idempotent. Key your UI nodes by msg id so a duplicated frame updates a node instead of creating a second one — retries and reconnects will duplicate frames eventually. What you display must match what you log. If a message was truncated by an interruption, the transcript you keep should be the truncated text, not the full generation; otherwise every later turn reasons over a conversation that never happened.

Tool calls interleaved with streamed text

Tool traffic does not pause the stream; it flows through it. A model producing a reply can emit a function call part mid-response, and what you observe is a sequence: some partial text (‘Let me check that…’), an event whose get_function_calls() is non-empty, a gap while the runner executes your tool, an event carrying the function_response, then more partial text as the model resumes with the result folded in. Parallel calls appear as several call parts in one event.

Two consequences for the client. First, the gap is the user experience problem: a slow tool is dead air in the middle of a sentence, so surface tool events explicitly — a ‘checking orders…’ chip beats a frozen cursor, and it doubles as free observability. Second, tool events are not the final response; is_final_response() is false for them, which is exactly why you should not treat ‘last event I received’ as ‘the answer’. Long-running tools are flagged so a client can show a pending state and before_tool callbacks gate sensitive actions identically in streaming and non-streaming runs — policy code is transport-agnostic.

Backpressure — where pull semantics stop

The downstream half is better behaved than people expect. run_async and run_live return asynchronous generators, which are pull-based: the runner produces the next event only when your loop asks for it. If your await ws.send(...) blocks because the client’s TCP window is full, the loop stops asking and the generator simply waits. Backpressure propagates for free — right up to the point where you break the chain.

You break it the moment you decouple. Push events into an unbounded asyncio.Queue for a separate sender task and that queue is now your buffer, growing without limit against a slow consumer until the process dies of memory. Bound it, and decide what happens when it fills: block, drop the oldest deltas, or hang up. The upstream direction is a different mechanism entirely — a LiveRequestQueue you push into is not the same pull-based contract, so a client shipping frames faster than the model consumes them is your problem to police — shed the heavy, low-value input first.