Every A2A client makes the same decision on every call: send a message and wait for the answer in one response, or open a stream and watch the task progress. Early drafts of the protocol named the two methods tasks/send and tasks/sendSubscribe, and many tutorials, SDK samples and blog posts still use those names. The names have changed twice since then, but the decision has not, and getting it wrong shows up as gateway timeouts, duplicate work, silent hangs and users staring at a spinner for a minute.
This article explains the choice as it exists in A2A 1.0, where there are really three ways to wait rather than two: what each returns, a client that survives a dropped stream, what the server must guarantee, and the failure modes that decide the question in practice. For the full list of operations and objects, start with the A2A protocol overview.
The names, then and now
The method names have moved, so the first job is to translate whatever you are reading into the version you speak. The behaviour behind each row is the same idea: one method returns a single result, one returns a stream of events, and one reattaches to the stream of a task that already exists.
| Purpose | Early drafts | 0.3 specification | 1.0 specification |
|---|---|---|---|
| Send, single response | tasks/send | message/send | SendMessage |
| Send, event stream | tasks/sendSubscribe | message/stream | SendStreamingMessage |
| Reattach to a task's stream | not confirmed | tasks/resubscribe | SubscribeToTask |
In 1.0 the JSON-RPC method strings are the PascalCase names, and every request carries an A2A-Version header so the server knows which vocabulary you mean. Two 1.0 changes matter for code written against older samples. The send result now holds exactly one of a message or a task, so the agent may answer a trivial question with a plain Message and no task at all. And the final flag that older status-update events carried was removed in 1.0.0 as redundant: the stream ends when the task does, so the client reads the state instead of a flag. Code that waits for final: true never sees it from a 1.0 server and tends to mistake the normal close for a dropped connection.
Three ways to wait, not two
The old two-way framing hides a third option that is often the right one. A2A 1.0 gives the client three shapes for the same work.
Blocking SendMessage. By default the server must hold the request open until the task reaches a terminal state (completed, failed, canceled or rejected) or an interrupted state (input-required or auth-required), and then return the task. One request, one response, no progress in between. It is the simplest code you can write.
SendMessage with returnImmediately. Setting returnImmediately: true in the request configuration makes the server return the task as soon as it exists, typically in the submitted state. The client then learns the outcome by polling GetTask or by registering a webhook, either in the same request's taskPushNotificationConfig or separately; see configuring push notifications. This is the shape for work measured in minutes or hours, and for callers that should not hold a connection at all.
SendStreamingMessage. The same parameters, but the response is a Server-Sent Events stream. It follows one of two patterns: either exactly one Message and then the stream closes, or a Task first, followed by status-update and artifact-update events. The specification requires the stream to close when the task reaches a terminal state. It does not say the stream must close on input-required or auth-required, so a client should treat those states as the end of its turn whether or not the connection stays open.
What travels on the wire
Both send methods take identical parameters, which is what makes switching cheap. In the JSON-RPC binding each streamed event is a full JSON-RPC response carrying the original request id, whose result is a StreamResponse with exactly one of task, message, statusUpdate or artifactUpdate. Annotated end-to-end transcripts are in A2A examples; the essentials look like this.
# Blocking: one JSON-RPC request, one response
POST /a2a/jsonrpc
A2A-Version: 1.0
{"jsonrpc":"2.0","id":7,"method":"SendMessage","params":{
"message":{"messageId":"m-7f2a","role":"ROLE_USER",
"parts":[{"text":"Summarise the Q3 incident reports"}]}}}
-> {"jsonrpc":"2.0","id":7,"result":{"task":{"id":"task-91","contextId":"ctx-4",
"status":{"state":"TASK_STATE_COMPLETED"},"artifacts":[ ... ]}}}
# Streaming: same params, different method, response is text/event-stream
{"jsonrpc":"2.0","id":8,"method":"SendStreamingMessage","params":{ ...same... }}
data: {"jsonrpc":"2.0","id":8,"result":{"task":{"id":"task-92","contextId":"ctx-4","status":{"state":"TASK_STATE_SUBMITTED"}}}}
data: {"jsonrpc":"2.0","id":8,"result":{"statusUpdate":{"taskId":"task-92","contextId":"ctx-4","status":{"state":"TASK_STATE_WORKING"}}}}
data: {"jsonrpc":"2.0","id":8,"result":{"artifactUpdate":{"taskId":"task-92","contextId":"ctx-4","artifact":{"artifactId":"summary","parts":[{"text":"Three incidents ..."}]}}}}
data: {"jsonrpc":"2.0","id":8,"result":{"statusUpdate":{"taskId":"task-92","contextId":"ctx-4","status":{"state":"TASK_STATE_COMPLETED"}}}}Artifact updates can arrive in pieces. An update with append: true adds parts to an artifact already sent under the same artifactId, and lastChunk: true says that artifact is complete. A client that renders text as it arrives needs this; a client that only wants the final answer can ignore chunks and read the finished task.
Streaming is optional for the agent. If its card does not declare capabilities.streaming, calls to SendStreamingMessage or SubscribeToTask must fail with UnsupportedOperationError. Read the card first and pick the method from it, rather than catching the error on every call.
Choosing per call
There is no global answer, because the right shape depends on how long the work takes, whether a human is watching and what sits on the network path. The table is a starting rule; the failure modes section explains the exceptions.
| Situation | Use | Why |
|---|---|---|
| Answer in a few seconds, machine caller | Blocking SendMessage | Least code; one round trip; easy retries |
| Human watching, output useful before it is complete | SendStreamingMessage | Partial artifacts and status as they happen |
| Tens of seconds, gateway with a short timeout in the path | Streaming or returnImmediately | A silent blocking request gets cut by idle timeouts |
| Minutes to hours, or caller may go away | returnImmediately plus push or polling | No connection held; survives client restarts |
| Agent card has no streaming capability | returnImmediately plus polling | Streaming methods are rejected |
| Orchestrator fanning out to many agents | returnImmediately, then GetTask or push | Thousands of open streams cost memory and sockets |
A useful default: stream when a person is waiting, return immediately otherwise, and keep blocking for short calls that are cheap to retry.
A client that survives a dropped stream
The hard part of streaming is not reading events; it is what happens when the connection breaks. A stream that dies after the Task event has told you the task id, and the task keeps running on the server regardless of your connection. The client should reattach with SubscribeToTask, whose first event is the current Task, so you rebuild your view from it instead of replaying missed events. If the task finished while you were disconnected, subscribing fails with UnsupportedOperationError because the task is terminal, and the right fallback is GetTask. If the stream dies before the Task event, you do not know whether a task was created, so treat it as a failed send.
import json, time, uuid
import httpx
URL = "https://agent.example.com/a2a/jsonrpc"
HDRS = {"A2A-Version": "1.0", "Authorization": "Bearer <token>"}
TERMINAL = {"TASK_STATE_COMPLETED", "TASK_STATE_FAILED",
"TASK_STATE_CANCELED", "TASK_STATE_REJECTED"}
INTERRUPTED = {"TASK_STATE_INPUT_REQUIRED", "TASK_STATE_AUTH_REQUIRED"}
def rpc(method, params, rid):
return {"jsonrpc": "2.0", "id": rid, "method": method, "params": params}
def sse_events(resp):
"""Yield the JSON-RPC result of each SSE data line."""
for line in resp.iter_lines():
if line.startswith("data:"):
msg = json.loads(line[5:])
if "error" in msg:
raise RuntimeError(msg["error"])
yield msg["result"]
def apply(task, ev):
"""Fold one StreamResponse event into a local task view."""
if "task" in ev:
return ev["task"]
if "statusUpdate" in ev:
task["status"] = ev["statusUpdate"]["status"]
elif "artifactUpdate" in ev:
u = ev["artifactUpdate"]; art = u["artifact"]
arts = task.setdefault("artifacts", [])
old = next((a for a in arts if a["artifactId"] == art["artifactId"]), None)
if old and u.get("append"):
old["parts"].extend(art["parts"])
elif old:
old["parts"] = art["parts"]
else:
arts.append(art)
return task
def stream_task(client, text, max_reconnects=5):
msg = {"messageId": str(uuid.uuid4()), "role": "ROLE_USER",
"parts": [{"text": text}]}
req = rpc("SendStreamingMessage", {"message": msg}, 1)
task, attempt = None, 0
while True:
try:
with client.stream("POST", URL, json=req, headers=HDRS,
timeout=httpx.Timeout(10, read=90)) as resp:
for ev in sse_events(resp):
if "message" in ev: # message-only stream
return ev["message"]
task = apply(task or {}, ev)
state = task["status"]["state"]
if state in TERMINAL or state in INTERRUPTED:
return task
# stream ended without a final state: fall through and reattach
except httpx.TransportError:
pass
if task is None: # never saw the Task: we cannot resume, only retry
raise RuntimeError("stream failed before the task was created")
attempt += 1
if attempt > max_reconnects:
break
time.sleep(min(2 ** attempt, 30))
req = rpc("SubscribeToTask", {"id": task["id"]}, 1 + attempt)
# last resort: read the task; it may have finished while we were away
r = client.post(URL, json=rpc("GetTask", {"id": task["id"]}, 99), headers=HDRS)
return r.json()["result"]The read timeout applies per read, so a long stream stays open while events keep coming, and reconnects back off. The local view is replaced whenever a full Task arrives, which is what makes reconnecting safe: the server's task store is the truth and the stream only a view of it. Background on recovering from crashes on both sides is in A2A error recovery.
What the server must get right
- The task store is the truth. Persist each state change before emitting its event. A reconnecting client calls
SubscribeToTaskand must receive a Task that reflects everything already streamed. - Task first. For a task-lifecycle stream the first event must be the Task. Clients depend on it to learn the id they need for reconnecting.
- Broadcast in order. The specification allows several concurrent streams for one task and requires every stream to receive the same events in the same order. Fan out from one ordered event log per task, not from separate code paths.
- Never let a slow reader block the work. Give each stream a bounded queue. If a client cannot keep up, close its stream and let it resubscribe; the work itself must not wait for a socket.
- Close on terminal states. Leaving a stream open after completion leaks connections and leaves clients waiting for something that will never come.
Worked example: a ninety-second report
An orchestrator asks a research agent for an incident summary. The agent spends 5 seconds planning, then 80 seconds reading documents, and writes the summary in the last 5 seconds. Between them sits a load balancer with a 60-second idle timeout, a common default for managed HTTP load balancers.
With blocking SendMessage, the request is silent for 90 seconds. At second 60 the load balancer closes the connection and the client sees a gateway error. The task, however, is still running, and it completes at second 90 with nobody listening. A naive retry starts a second, identical task.
With SendStreamingMessage, the client gets the Task at about 100 milliseconds, a working status at second 5, and progress updates while documents are read, so the connection is never idle for long enough to be cut if the agent emits an update at least every few tens of seconds. The summary arrives in chunks between seconds 85 and 90 and the stream closes on the completed state. If a deploy restarts the proxy at second 40, the client reattaches with the task id and loses nothing.
With returnImmediately, the client has the task id within 100 milliseconds and polls GetTask every 10 seconds, making about nine extra small requests, or receives one webhook at second 90. Nothing is held open, so this scales to thousands of concurrent reports, at the cost of up to one polling interval of extra latency.
Failure modes
- Buffering proxies. A reverse proxy that buffers responses holds every event until the stream ends, so streaming silently degrades to slow blocking. Disable response buffering for the A2A route (in nginx,
proxy_buffering offor anX-Accel-Buffering: noresponse header) and test through the real path, not localhost. - Idle timeouts. Long silent periods inside a stream are cut just like silent blocking requests. Emit status updates during long steps; some servers also send SSE comment lines as keep-alives, which conforming SSE parsers ignore.
- Retrying a blocking send that timed out. The task usually survived. Before retrying, look for it with
ListTasksin the context, or reuse the samemessageIdso a server that deduplicates can recognise the retry. - Waiting for a removed flag. Clients ported from older samples loop until
finalis true. In 1.0 decide from the task state. - Hanging on interrupted states. A client that only stops on terminal states will wait forever when the agent asks for input. Treat input-required and auth-required as the end of your turn.
- Resubscribing to a finished task.
SubscribeToTaskfails on terminal tasks; fall back toGetTaskinstead of retrying the subscription.
Trade-offs
| Dimension | Blocking | returnImmediately | Streaming |
|---|---|---|---|
| Client complexity | Lowest | Polling loop or webhook endpoint | SSE parsing, chunk assembly, reconnects |
| Time to first feedback | End of task | Task id at once, then polling interval | Task at once, then live |
| Connection held | Whole task, silent | None | Whole task, active |
| Survives client restart | No | Yes | Yes, with the task id |
| Agent requirement | None | None (push needs capability) | capabilities.streaming |
| Scales to many tasks | Poorly | Best | Moderately |
What to do next
- Search your code and SDK samples for
tasks/send,tasks/sendSubscribe,message/streamand loops onfinal, and port them to the 1.0 names and state checks. - Read each peer's Agent Card at startup and record whether it declares streaming and push notifications; choose methods from that, not from try-and-catch.
- Classify your calls by expected duration and audience using the table above, and set blocking only for short machine-to-machine calls.
- Put a stream reconnect path in your client: remember the task id, resubscribe with backoff, fall back to GetTask.
- Run a streaming call through your real load balancers and proxies and confirm events arrive as they are sent, not all at the end.
- On the server, persist before emitting, emit the Task first, close streams on terminal states and cap per-stream buffers. For lifecycle rules see A2A task lifecycle states.