Why architecture matters here
Rolling your own RPC is a mistake most teams make once: framing, schema evolution, streaming, cancellation, load balancing and retry-under-deadline are each subtly hard, and gRPC is a widely-deployed implementation of all of them behind a generated, language-native API.
The wire savings are real but not a fixed ratio — protobuf writes field numbers rather than names, varint-encodes integers and needs no quoting or delimiters, so the gap against JSON is wide for many small numeric fields and narrow for a payload that is mostly free text. HTTP/2 multiplexing collapses N connections into one, which is the source of the efficiency and, as the load balancing section shows, of the most common way gRPC deployments go wrong. The rest — propagating deadlines, one place for auth and tracing, retry and balancing policy as configuration — is why you take the framework instead of putting protobuf on plain HTTP yourself.
The HTTP/2 foundation, and what gRPC adds
gRPC is not a transport. It is a calling convention layered on HTTP/2, and most of its operational surprises are an HTTP/2 mechanism showing through.
HTTP/2 supplies four things. Binary framing: HEADERS, DATA, SETTINGS, PING, WINDOW_UPDATE, RST_STREAM and GOAWAY frames, each tagged with a stream ID. Streams: an independently-terminated lane, which is what one RPC actually is — cancelling a call is a RST_STREAM on its stream, and MAX_CONCURRENT_STREAMS caps how many run at once per connection. HPACK: header compression against a shared dynamic table, so the :path, content-type and auth headers repeated on every call cost a few bytes after the first. Flow control: a credit window per stream and per connection, initially 65,535 bytes, replenished by WINDOW_UPDATE, so a slow reader stops the writer instead of buffering without bound — mature implementations grow that window from a bandwidth-delay estimate rather than leaving it at the spec default.
What gRPC adds is what HTTP never had: a method naming scheme, message framing inside the byte stream, a machine-readable status independent of the HTTP status, a deadline that travels with the request, and a schema.
The wire format: length-prefixed frames and trailing status
A unary call is one HTTP/2 stream carrying a POST. The request HEADERS frame holds :method POST and :path /pkg.Service/Method — the fully-qualified proto name, which is why gRPC has no URL design debate — plus content-type: application/grpc+proto, te: trailers, and any metadata.
DATA frames carry length-prefixed messages. Each message is preceded by exactly five bytes: one compressed-flag byte (0 = identity, 1 = compressed with the algorithm named in grpc-encoding), then a 4-byte big-endian length. Because of that prefix a message boundary never depends on a frame boundary — a 6 MB message spans many DATA frames, several tiny messages share one — and the common 4 MiB receive limit is enforceable before any payload is parsed.
Then the part that trips people up: the status arrives in trailers. The response ends with a second HEADERS frame carrying grpc-status and optionally grpc-message. It has to be trailers, because on a streaming call the server cannot know whether the RPC succeeded until after the last message is written, and even a unary handler can fail after headers are flushed. The HTTP status is 200 on a failed RPC: the transport worked, the call did not. A server failing immediately sends a Trailers-Only response — one HEADERS frame with END_STREAM carrying both statuses, no DATA at all.
The architecture: every piece explained
Read the diagram top to bottom. The top row is one RPC in flight: a stub generated from your .proto, a stream multiplexed onto a shared HTTP/2 connection, a generated dispatcher on the far side. The middle rows are what rides on it — the schema, the interceptor chain, and the four call shapes the streaming model allows. The bottom rows are the operational surface: finding backends and spreading load across them, bounding the whole call tree with a deadline, and the bridges browsers and REST clients need. Each gets its own section below.
Four call types and what they really guarantee
The four shapes differ only in which side may send more than one message. Unary: one request, one response. Server-streaming: the client sends one message and half-closes, the server writes many. Client-streaming: the reverse. Bidirectional: both write independently, and the interleaving is entirely the application's business — a bidi stream is not request/response unless you make it so.
What you get: messages on one stream are ordered and never lost without the stream erroring, since it is all one TCP byte stream. What you do not get is any guarantee once the stream breaks. A server stream that dies at message 400 of 1000 is indistinguishable, to the client, from a server that meant to send 400 — unless your protocol carries a completion marker or a resumption token. Streams are not queues either: no persistence, no replay, no consumer groups. The only real backpressure is the HTTP/2 window, and it helps only if your binding exposes manual flow control instead of draining into an unbounded application buffer.
Protobuf as the IDL, and the evolution rules that hold
The .proto file is the contract; codegen is convenience on top. On the wire a message is key-value pairs where the key is a varint packing the field number and wire type, not the name. Field 1 with varint type is tag byte 0x08. That one choice explains both why protobuf is small and why the compatibility rules are what they are.
Safe: adding a field with a fresh number — old readers keep it as an unknown field and most runtimes preserve it through a parse/serialize round trip, which is what stops a service in the middle from silently stripping data; renaming a field; adding an enum value if readers tolerate unknown ones.
Unsafe: reusing a number after deleting a field, which deserializes old data into the wrong field — hence reserved 4, 7; and reserved "email";; changing a type across wire-type families; changing repeated to singular. Mind the defaults too: proto3 scalars have no presence unless declared optional, so an unset int32 and a deliberate zero are identical bytes, and enum value 0 is what every old reader sees for a value it does not know.