Prose results make everybody guess

A tool result is, at its simplest, an array of content blocks — usually one block of text. That is enough when the only consumer is a language model reading an answer in a chat window. It stops being enough the moment anything deterministic wants that answer. A host that needs to draw a chart, cache a row, or compare a value against a threshold has to recover fields out of a sentence, and sentences vary: today Temperature: 27.4C, tomorrow it is currently 27.4 degrees and humid. Regexes accumulate; none of them are testable against a phrasing that has not happened yet.

The model gets drafted into the same job. It reads the prose, extracts a number, and retypes it into the next call’s arguments — tokens spent on clerical work, plus a transcription step that can silently go wrong. Both sides are reconstructing a structure the server already had and discarded the moment it formatted the string. Structured output exists to stop the server discarding it.

Advertisement

outputSchema — declaring the shape a tool returns

A tool definition already carries an inputSchema: a JSON Schema describing the arguments a caller may send. Current revisions of the protocol let a tool also carry an outputSchema — conventionally an object type — describing the fields a successful result will contain. It is the mirror image of the input contract, published the same way and in the same place: in the tool listing, at discovery time, before any call happens.

Publishing the shape early is what makes it useful. A client can compile a validator once instead of per call. A host can decide in advance which renderer or route fits this tool. And a model told what a call returns can judge whether the tool serves its goal at all, rather than calling it to find out. Declaring the schema is also a promise: the server commits that what it sends back will conform to it.

Advertisement

structuredContent — the typed payload

The other half of the contract lives in the result. Alongside the familiar content array, a tool result may carry a structuredContent field holding exactly one JSON object that matches the declared outputSchema. It is a single object rather than a list of blocks precisely because it is not presentational: it is the deserialized return value, the thing a program will index into.

The distinction is worth holding onto, because the two fields answer different questions. content is a channel of things to display or read — text, images, embedded resources, ordered for legibility. structuredContent is a value. A field named total there is the total, at a known type, at a known path, whatever the tool chose to say in its summary. Once a result carries a value channel, the host stops doing natural-language processing on its own tool output — which was never a reasonable thing to ask of it.

Why both channels ship together

The obvious simplification — send the structured object and drop the prose — is the one the design deliberately refuses. A tool that returns structuredContent is expected to also include a text representation of the same data in content, typically the JSON serialized into a text block.

There are two consumers and they want different things. The model reasons over language; handing it a bare object with no readable rendering makes it worse at exactly the summarizing and explaining you invoked it for. The application wants a typed value and no parsing at all. Sending both costs a modest duplication and serves each consumer in its native currency. One rule keeps the arrangement honest: the two views must agree. structuredContent is authoritative for programs and content is a rendering of it, not a second, subtly different answer travelling in the same envelope.

MCP structured tool output — outputSchema declares the shape, structuredContent carries typed datacontent[] stays human-readable; structuredContent is machine-consumable and validatedTool definitionname, inputSchema, outputSchema (JSON Schema)Server tool handlerruns logic, returns CallToolResultClient / hostvalidates result vs outputSchemacontent[]text/image blocks for the model + humansstructuredContentone JSON object matching outputSchemaBackward-compatcontent mirrors JSON as text tooLLM consumes typed fields; host can render/route without re-parsing proseVersioning — outputSchema evolves additively; clients ignore unknown fields, validate known onesdeclaresresultcontentstructuredtypedcontext
A tool declares an outputSchema; the handler returns both a readable content[] and a structuredContent object; the client validates the object and feeds typed fields onward without re-parsing prose.

Validation, and what happens when a tool breaks its own contract

Validate on both ends. A server should check its own payload against its declared outputSchema before returning, because a handler is ordinary code and ordinary code drifts from its documentation. A client should check on receipt, because a contract nobody enforces is only a comment. Wiring that same check into a test suite catches most drift before a user ever sees it.

When a payload fails, treat it as a broken tool rather than a confused model. A temperature arriving as the string "27.4C" where the schema promised a number is a server defect, and catching it at the boundary is enormously better than letting the value travel into a gauge, a comparison, or the next call’s arguments, where it fails somewhere unrelated and unrecognizable. What the client does next — reject outright, or fall back to the readable content with the structure marked untrusted — is a policy choice, and a separate matter from a tool that ran correctly and legitimately reported failure.

Backward compatibility with clients that ignore structure

Not every client understands structured output, and a server generally cannot know which ones do. That is precisely why the mirrored text block is not optional politeness. An older client reading only content receives a serialized JSON string and displays it: it loses the typing, but it does not lose the answer. Nothing breaks, no negotiation flag is required, and a server can start emitting structure without waiting for the ecosystem to catch up.

The obligation runs the other way too. A client that does understand structured output must not assume it is present. A tool that declares no outputSchema, an older server, or an error result will send content alone, and the client’s fallback has to be to use the text. Treat structured output as an enhancement layered on a channel that still works without it, and its absence becomes an ordinary case rather than a crash.