A notification, not a request — and why that is deliberate

Cancellation travels as a JSON-RPC notification: a message with a method and params but no id of its own, and therefore no response and no error. The sender writes notifications/cancelled with the requestId it wants abandoned, optionally a reason, and moves on. Nothing comes back.

That shape is a choice, not an oversight. If cancelling were itself a request it would have an id, which means it could hang, time out, fail — and need cancelling — a supervisory channel on top of the channel you were trying to interrupt. Fire-and-forget also matches the semantics honestly: the sender is not asking permission, and the receiver cannot truthfully answer “yes, it stopped” the instant the message lands. An unacknowledged hint is a smaller lie than a confirmation you cannot back up. One request is conventionally off-limits: initialize has no mid-flight state to unwind.

Advertisement

The race you cannot close

Cancellation and completion travel in opposite directions over the same wire, and neither side sees the other’s in-flight bytes. A client that cancels at T may find the server finished at T minus one millisecond and already flushed a perfectly good response. No protocol trick removes this window: closing it needs a synchronous round trip, exactly what a notification refuses to be.

The race is therefore not a bug to fix but a condition to tolerate, symmetrically. The requester must accept a full response for a request it already cancelled and drop it. The receiver must accept a cancellation for an id it no longer recognizes — already completed, already cancelled, never seen — and ignore it silently. Neither is an error. Treating either as one is the classic bug: a server replying “unknown request id” turns a timing artifact into noise, and a client that surfaces a late result shows output the user explicitly asked to stop.

Advertisement

What a server should actually do on receipt

The correct posture is best-effort abort. On receipt, a server looks the id up in its table of in-flight requests. No entry, no action. If there is one, it trips that request’s cancellation primitive — an AbortController, a cancelled asyncio task, a Go context — and returns immediately, without waiting to see whether the handler noticed.

Then it stays quiet. A cancelled request gets no response: not a result, and not a JSON-RPC error either. Abandoned is not failed, and inventing an error reply recreates the ambiguity the design was avoiding. The one real obligation is to make handlers cancellable at all — a token nobody checks is decoration — by threading it into every slow call, so an HTTP read or model stream aborts instead of running to completion inside a task the runtime believes it cancelled.

MCP cancellation — notifications/cancelled races the in-flight request to a safe stopadvisory, best-effort, race-awareClient calleruser cancels / timeoutRequest trackerrequestId + AbortControllerTransportstdio / streamable HTTPServer dispatcherper-request contextcancelled notifrequestId + reasonIn-flight handlertool / LLM / IO callCancellation tokenchecked at await pointsResource cleanupsockets, temp, spendResponse suppressiondrop late result, no errorRace windowresult already sent = ignoreOps — orphan detection + cancel latency + leaked-work metrics + idempotency keysemittrackroutedispatchsuppressunwindreleaseoperateoperate
MCP cancellation: the client sends notifications/cancelled by requestId; the server races it against the in-flight handler, unwinding work and suppressing any late response.

Making handlers cooperative

Cancellation on every mainstream runtime is cooperative: nothing forcibly kills a running function, so the handler must reach a point where it observes the token and unwinds. That yields a checklist. Check the token before each expensive step and inside any loop over a large input, so a thousand-item batch does not grind through item 998 after the caller left. Pass the token, not just a URL, to your HTTP client so the socket read is genuinely abortable. Break out of streaming loops when it trips instead of draining the stream.

The trap is blocking code. A synchronous CPU-bound loop, or a driver with no cancellation support, holds its thread until it finishes, token or not; the remedies are running it somewhere you can abandon (a worker thread, a subprocess) or chunking it so checkpoints exist. Measure the delay from cancel received to work actually stopped — that number, not the protocol, is your real guarantee.

Cleanup and partial side effects — the genuinely hard part

Stopping is easy. Stopping safely, when the operation already changed something, is the real work. An abandoned handler may hold an open transaction, a half-written temp file, a lock, a leased sandbox, a pooled connection. Structured unwinding — try/finally, context managers, defer — is non-negotiable, because cleanup sprinkled through a happy path never runs on the abandonment path.

The harder class is effects you cannot take back. If the tool already POSTed a payment or triggered a deployment, cancellation stops the polling, not the consequence. Three habits keep that survivable: order the handler so irreversible external steps happen as late as possible, shrinking the mid-commit window; carry an idempotency key so a cancelled-then-retried operation reuses the effect rather than duplicating it; and record the abandonment, because a request cancelled after committing an external effect is a reconciliation item, not a no-op.

Propagating downward — subprocesses, fan-out, and gateways

A tool call is rarely a leaf. It may have spawned a subprocess, issued three parallel HTTP calls, asked the host for an LLM completion via sampling, or forwarded work to another MCP server behind a gateway. Cancellation is only as good as its reach, so the per-request token should be the root of a tree: every child operation derives from it and dies with it.

Each edge needs its own mechanism. Child tasks inherit the context. Subprocesses need a signal, usually a process group signal — killing a shell that spawned a compiler leaves the compiler running — with a grace period before a hard kill. A gateway or proxy must translate an inbound cancellation into an outbound notifications/cancelled for the id it minted downstream, which means keeping that id mapping alive for exactly this purpose. Anything you forget to wire becomes an orphan: work that continues, and bills, with nobody waiting for it.