Everything the model sees is a Part

ADK has no separate image API bolted onto a text API. It has one content model, borrowed from the underlying google-genai types, and every modality lives inside it. A message is a Content with a role and a list of Part objects, and a Part is a tagged union: it carries either text, or binary data, or a reference to data stored elsewhere, or a function call, or a function response. Nothing else. Once that lands, ‘sending an image’ is just appending a differently-shaped part to a list you were already building.

Three shapes matter for media. part.text is the ordinary string. part.inline_data is a Blob — a mime_type plus raw data bytes — carried inside the request. part.file_data is a FileData — a mime_type plus a file_uri — carried as a pointer the model service dereferences on its own side. One turn can mix all of them:

from google.genai import types

message = types.Content(
    role="user",
    parts=[
        types.Part(text="Which line item is the delivery fee?"),
        types.Part(inline_data=types.Blob(
            mime_type="image/jpeg", data=jpeg_bytes)),
        types.Part(file_data=types.FileData(
            mime_type="application/pdf",
            file_uri="gs://receipts/invoice-8812.pdf")),
    ],
)

Order is not cosmetic: the model reads parts in sequence, so a question placed before several images is answered with all of them in view.

Advertisement

Three ways to hand over bytes

Given the same photo you have three genuinely different transport options, and picking wrong causes both cost surprises and hard errors. Inline puts the bytes in the request: no upload step, no cleanup, bounded by the total request size the endpoint accepts — tens of megabytes including base64 expansion, smaller than it sounds once you attach a second file. Use it for screenshots and camera captures, anything a user just typed a message about.

File URI uploads the bytes once to a store the model can reach — the provider’s file service, or a Cloud Storage bucket on Vertex — and references them by URI. That clears the request ceiling and is far cheaper when one document is referenced across many turns: you upload the fifty-page contract once, not on every request. The cost is lifecycle — uploaded files expire, and a stale URI fails at request time, not attach time. Artifacts are ADK’s own layer over all this, answering a different question: not ‘how does the model reach the bytes’ but ‘where do they live between turns, and who owns them.’

ADK multimodal — agents that see, hear, and readimages, audio, documents as first-class inputsMultimodal inputimage, audio, PDF, videoContent partstyped message partsMultimodal modelGemini native vision/audioGrounded reasoninganswer about the mediaArtifactslarge media out of contextTool outputs as mediacharts, screenshotsMultimodal outputgenerated images / audioModality routingwhich model, which pathDocument understandingPDFs, forms, tablesCost + latencymedia tokens are expensiveOps — media handling + token budgets + modality-specific evalsstoreproducegeneraterouteparsebudgetevaluateoperateoperate
The whole surface at a glance: media arrives as typed content parts — inline bytes, a file URI, or an artifact loaded just in time — is reasoned over by a natively multimodal model, and can come back out as generated media.
Advertisement

Artifacts — where media lives between turns

An ADK artifact is a named, versioned binary belonging to a session or a user, held by an ArtifactService you configure on the Runner just as you configure a SessionService — in-memory for development, Cloud Storage—backed in production. The unit it stores is a Part, the same type the model consumes, so an artifact is not a side-channel file; it is model content you have parked.

Tool and callback contexts expose the save and load operations, and two properties shape how you use them. Saves are versioned: writing the same filename appends a version rather than overwriting, so chart.png can be regenerated across a conversation with every earlier version still addressable. Filenames are also namespaced: a plain name is scoped to the session and dies with it, while the conventional user: prefix promotes the artifact to the user and makes it visible from every future session — the difference between ‘the image they just uploaded’ and ‘their ID document on file.’

The discipline that matters is that an artifact is not in the context. Saving a PDF costs zero tokens and keeps costing zero until something loads it into a request. ADK ships a built-in tool that lets the model pull artifacts in by name when it decides it needs to look; the alternative is loading them deterministically in a callback. Either way, media enters the context just in time.

MIME types are the contract

The mime_type on a blob or file reference is not metadata — it is the instruction that decides how the bytes are decoded. Get it wrong and you do not get a helpful message about image formats; you get a decode error, or a confident answer about nothing. Declare what the bytes are, never what the filename claims.

ModalityTypically acceptedWhat to watch
Imageimage/png, image/jpeg, image/webp, HEIC/HEIFNo SVG — it is markup, not raster; rasterise first
AudioWAV, MP3, FLAC, AAC, OGG, AIFFChannels are downmixed; long files are the cost risk
VideoMP4, MOV, WebM, MPEG, AVI, 3GPPSampled to frames, not decoded whole
Documentapplication/pdf, text/plainOffice formats are usually not native — convert to PDF
Code / datatext/plain, CSV, source filesPlain text is the cheapest path

Two traps recur. DOCX, XLSX and PPTX arrive constantly, are ZIP containers rather than documents, and passed through as opaque blobs produce gibberish — convert to PDF at the door. And double encoding: the blob field wants raw bytes, so code that base64-encodes first hands the model an ASCII string it tries to read as a JPEG.

What media actually costs

Media is the biggest source of context bloat in an agent, and the arithmetic is worth memorising because it is unintuitive. A small image costs roughly what a page of prose costs. A minute of audio costs a couple of thousand tokens. A minute of video costs an order of magnitude more than a minute of audio, because you pay for sampled frames and the soundtrack. A PDF page costs an image plus the text on it.

The published figures for the current Gemini generation sharpen that: an image small enough to fit one tile bills as a flat couple of hundred tokens, and larger images are cut into tiles that each bill the same amount — cost scales with area, not with file size. Audio runs at a few tens of tokens per second; video at a few hundred per second at default resolution, with a low-resolution mode cutting that by roughly four. Treat the constants as a shape rather than a promise and check them for the model you deploy.

Two consequences follow. Uploading a 12-megapixel phone photo is waste — past the resolution the model tiles at, extra pixels buy tokens rather than detail; crop to the region if you need fine print. And media is sticky: an image attached at turn three sits in the history at turn twenty, billed again on every request in between.

Documents — a PDF is read as pages, not as text

Native document understanding deletes the most fragile pipeline you own. The model does not receive your PDF as an extracted text dump; it receives the pages and reads them the way a person does — columns as columns, a table as a grid, handwriting in the margin as handwriting. Everything a text extractor silently destroys about a form is exactly what the model needs to fill it in.

The envelope is generous but real: current models accept documents in the high hundreds of pages, and each page bills as an image plus its text. Large limit, per-page cost — that combination drives the design. If a question concerns one clause of a two-hundred-page contract, sending all two hundred pages is economically foolish; keep the document as an artifact, maintain a cheap text index or per-section summary beside it, and load only the relevant page range.

Two habits sharpen accuracy. Ask for structure, not prose — pair the document with a response schema or explicit field list, because ‘extract the invoice details’ returns a paragraph you must then parse, reintroducing the brittleness you removed. And ask for provenance: a required page number beside every extracted value turns an unverifiable claim into a checkable one.

Video — sampled, not watched

Video is the modality whose mental model most needs correcting. The model does not watch your file; the service samples it into stills at a low fixed rate — on the order of one frame per second — extracts the audio track, and reasons over that sequence. Motion faster than the sampling rate is therefore invisible: a ball crossing the frame in 300 ms may not appear at all, and ‘how many times did the light blink’ asks a question the representation cannot answer. The audio track, by contrast, is continuous, so a lecture is understood far better than a magic trick.

Because frames are billed, duration is the cost variable, and both levers are about sending less. Most services accept a clipping interval — a start and end offset — so a question about one moment in an hour-long recording sends ninety seconds. A low-resolution setting is the right default whenever you need to know what is happening rather than to read text on screen. One property comes free: frames are timestamped, so the model answers ‘when’ questions with offsets you can seek to.

Audio files, and where the voice agent takes over

Audio enters an ADK agent in two entirely different ways, and conflating them causes real confusion. This article covers the first: an audio file — a voicemail, a recorded meeting — attached to a normal turn as a blob or a file reference. The second is live audio, where PCM frames stream continuously into a LiveRequestQueue under run_live — a real-time system with its own format constraints, latency budget and barge-in problem, and the voice companion’s territory.

For files, the thing worth understanding is that native audio comprehension is not transcription with extra steps. Transcription is one thing you can ask for; the model also hears what a transcript deletes — who was speaking, whether they sounded frustrated, that two people talked over each other, that eight seconds of silence preceded the answer. ‘Summarise the complaint and flag whether the caller became angry’ is one request, not an ASR pass plus sentiment analysis over its output. Design around duration: chunk multi-hour recordings and carry per-segment summaries forward rather than the audio.