A confidential computing deployment runs a model inside hardware-isolated memory and can prove, through remote attestation, which software is running. Most write-ups stop at the server: build a measured image, verify the GPU, release the weight key. That protects the model owner. It does not, by itself, protect the person typing the prompt, because that person usually has no way to check anything. Their prompt goes to a hostname, a load balancer terminates TLS, and the guarantee they were promised depends on infrastructure they cannot see.
This page is about closing that gap: designing confidential inference so that the end user's client can verify it, and so that the operator could not quietly break it. It covers who verifies what, why the request must be encrypted to the attested node rather than to the front door, how an oblivious relay and node selection stop an attacker from steering one user to a compromised machine, how a transparency log turns a private allow-list into a public commitment, what stateless serving means inside the node, and what a TEE can never hide. Server bring-up, image measurement, GPU confidential mode and multi-GPU limits are covered in confidential compute for LLMs, in depth; the comparison with homomorphic encryption and MPC is in secure inference for LLMs.
Who verifies, and against what
The IETF's Remote ATtestation procedureS architecture, RFC 9334, gives the parties names that are worth using because they force precision. The attester is the inference node: it produces evidence, signed claims about its hardware and the software it booted. The verifier appraises that evidence against endorsements from the hardware vendors (certificates that say a signing key belongs to a genuine chip) and reference values from whoever builds the software (the measurements a good build should produce). The verifier's output is an attestation result. The relying party is whoever acts on that result.
In a server-only design the relying party is a key broker the operator runs. In a client-verifiable design the user's device is a relying party too. In RFC 9334's passport model the node shows the client an attestation result from a verifier; in the background-check model the client receives raw evidence and verifies it, or sends it to a verifier it chooses. The first is lighter for phones; the second removes trust in the operator's choice of verifier. Either way, ask: who supplies the reference values, and can the operator change them unnoticed?
The request path, end to end
The architecture that answers that question has five parts, and the diagram shows what each one is allowed to see. The client fetches node evidence, verifies it and encrypts the prompt. A relay, run by a different organisation, forwards the encrypted request so the operator never learns the client's IP address. A gateway run by the operator balances load but holds no key that can decrypt anything. The attested node decrypts, runs the model and encrypts the reply. A transparency log publishes every software measurement the operator has ever deployed, and a key directory serves each node's current public key alongside its evidence.
Encrypt to the node, not to the front door
The commonest way to undermine confidential inference is ordinary: TLS terminates at a load balancer, which forwards plaintext to the enclave. Every prompt now exists in plaintext on an unattested machine. Binding the node's TLS key into its report data only helps if the TLS session really runs from client to node, which rules out most managed load balancers.
The cleaner pattern is to encrypt the request body to a key only the attested node holds. Hybrid Public Key Encryption, RFC 9180, is the standard tool: the client encapsulates a fresh key to the node's public key, seals the request, and derives a response key from the same context. The node generates its HPKE key pair inside the TEE at boot and commits its hash into the report data, so evidence and key cannot be separated. Gateways and middleware then just route ciphertext.
# Client-side request path. Function names marked "sdk" are placeholders for your
# attestation SDK and HPKE library; the order of checks and the refusals are the point.
def send_private_prompt(prompt: bytes, directory, log, relay):
candidates = directory.fetch_nodes(count=3) # small subset, not the whole fleet
usable = []
for node in candidates:
ev = sdk.verify_evidence(node.evidence, # signatures, cert chains, revocation
nonce=node.challenge) # freshness from the directory
if ev.debug or ev.tcb < POLICY.min_tcb:
continue
if ev.measurement not in POLICY.allowed or not log.included(ev.measurement):
continue # unknown or unlogged build
if ev.report_data != sha256(node.hpke_public_key + node.challenge):
continue # key not bound to this evidence
usable.append(node)
if not usable:
raise RefuseToSend("no node passed verification") # fail closed, never fall back
node = secrets.choice(usable)
sealed, ctx = sdk.hpke_seal(node.hpke_public_key, info=b"llm-request-v1",
plaintext=pad_to_bucket(prompt))
reply = relay.post(node.id, sealed) # OHTTP: relay sees IP, not content
return unpad(ctx.open_response(reply))Four refusals carry the security: debug or out-of-date hardware is skipped, the measurement must be allow-listed and logged, the key must be bound to the evidence, and when nothing passes the client refuses to send rather than falling back to a plain endpoint.
Transparency logs: making the allow-list public
An operator-held allow-list has an obvious weakness: the operator can add a build that logs prompts, deploy it to a few nodes, and remove it later. The fix is the one certificate transparency applied to TLS certificates. Every deployable measurement is appended to an append-only Merkle tree log, and clients accept a node only if its measurement is provably in the log. A malicious build can still be deployed, but not secretly.
The client-side check is small. The log signs a tree head (size and root hash); the client asks for an inclusion proof for its leaf and recomputes the root. The function below implements the inclusion check from RFC 9162, the certificate transparency 2.0 specification. It was tested against trees of 1 to 39 leaves built with the same specification's hashing rules, accepting every genuine proof and rejecting every tampered leaf.
import hashlib
def leaf_hash(data: bytes) -> bytes:
return hashlib.sha256(b"\x00" + data).digest()
def node_hash(left: bytes, right: bytes) -> bytes:
return hashlib.sha256(b"\x01" + left + right).digest()
def verify_inclusion(leaf: bytes, index: int, tree_size: int, proof: list, root: bytes) -> bool:
"""Inclusion proof check following RFC 9162, section 2.1.3.2."""
if index >= tree_size:
return False
fn, sn, r = index, tree_size - 1, leaf_hash(leaf)
for p in proof:
if sn == 0:
return False
if fn & 1 or fn == sn:
r = node_hash(p, r)
if not fn & 1:
while fn and not fn & 1:
fn >>= 1
sn >>= 1
else:
r = node_hash(r, p)
fn >>= 1
sn >>= 1
return sn == 0 and r == rootInclusion does not prove the log shows everyone the same tree. Close that with consistency proofs between tree heads a client has seen, and independent monitors comparing tree heads across clients.
Non-targetability: the operator should not be able to pick its victim
Attestation proves what a node runs. It does not stop an operator, or an attacker who has compromised one node, from routing a particular user's requests to that node. Non-targetability is the property that a limited compromise cannot be aimed. Three mechanisms provide it. An oblivious relay following Oblivious HTTP, RFC 9458, splits knowledge: the relay sees the client's IP address but cannot read the encapsulated request, and the service decrypts the request but never sees the client's IP. Anonymous authorisation, such as tokens issued under RSA blind signatures (RFC 9474), lets the service check that a request is entitled to capacity without linking it to an account. Client-side node choice from a small verified subset means a compromised node only ever sees a random fraction of traffic, not a chosen user's.
Apple's Private Cloud Compute is the most detailed public example. Its June 2024 security blog post lists stateless computation, enforceable guarantees, no privileged runtime access, non-targetability and verifiable transparency as requirements, and describes a third-party OHTTP relay, requests encrypted to the public keys of specific nodes the device has validated, RSA blind signatures for authorisation, and an append-only transparency log of all production software measurements. Read the post before copying any detail; the point is that every part of this pattern has been built at scale.
Stateless serving inside the node
Inside the TEE the threat changes from the operator reading memory to the software remembering things. Stateless serving means a prompt is used to produce its answer and then ceases to exist. In practice that requires the image to contain no persistent writable storage for request data, an encrypted scratch volume whose key is generated at boot and never leaves memory (so a reboot destroys anything written), no remote shell or debugger, and logging that is structured, allow-listed and free of content.
LLM serving adds one complication: caches. Prefix caching and KV-cache reuse are major performance features, and they keep fragments of earlier prompts in memory by design. A shared prefix cache is also a side channel, because a faster first token reveals that someone else sent the same prefix. Scope caches to one request, or to one authenticated tenant when the tenant is the data owner, as described in tenant isolation for LLM services, and evict on a timer you can state. Measure the prefill cost before promising it.
What the TEE cannot hide: metadata
Everything outside the green box still sees traffic shape: request sizes, arrival times, the chosen node, response sizes and, with streaming, each token's arrival. Published research on encrypted LLM traffic has used streamed token lengths to infer response content, so treat this as a real channel.
The defences are blunt but effective. Pad requests and responses to size buckets, for example powers of two up to the context limit. Batch streamed tokens into fixed-size chunks sent on a fixed cadence, or do not stream at all for sensitive tiers. Keep timing coarse: a reply that always takes at least a bucketed minimum reveals less than one that returns as fast as possible. Each costs latency or bandwidth, so set them per product tier. For outbound paths from the node, see egress filtering for LLM systems.
Worked example: a clinic&#x27;s note summariser
A clinic group wants a hosted model to summarise visit notes without the provider being able to read them. The clinician's app pins a policy at install: acceptable hardware vendors, minimum firmware levels, the log's public key, and the rule that a measurement must be allow-listed and logged. For each request it verifies three candidate nodes, drops one with old firmware, picks one of the other two at random, pads the note to the 8 KB bucket, seals it with HPKE and sends it through a relay run by another company. The gateway forwards ciphertext it cannot read. The node summarises with request-scoped caching and logs only a request ID, bucket, latency and status.
Now ask how the provider could read notes. A new server build needs a new measurement, which must appear in the public log the clinic's auditors watch. A debug node fails verification. Terminating TLS at the gateway yields only ciphertext. Compromising one node exposes a random share of requests, without knowing whose. What remains is trust in the hardware vendors, the reviewed code and the metadata budget the clinic accepted.
Failure modes
- Plaintext at the front door. TLS ends at a gateway and the node receives plaintext. Attestation is then decorative for users.
- Silent fallback. When verification fails the client retries against a non-confidential endpoint. Make refusal visible and fail closed.
- Evidence not bound to the key. The report data does not commit to the HPKE public key and a fresh challenge, so genuine evidence can be replayed in front of an attacker's key.
- Stale policy on devices. Clients pinned to old firmware minimums keep accepting vulnerable hardware. Ship policy updates like security patches.
- Caches and logs inside the node. A genuine, attested image still leaks if it logs prompts or shares prefix caches across users.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Client verifies evidence itself | No trust in the operator's verifier | More code and certificate handling on every client |
| Passport model (verifier result) | Light clients, simple caching | The verifier is now trusted; choose it independently |
| HPKE to the node | Gateways can be ordinary infrastructure | Custom protocol in front of standard serving APIs |
| OHTTP relay | Operator cannot link requests to IP addresses | A second company in the path; extra latency and cost |
| Public transparency log | Malicious builds cannot be deployed secretly | Release process must publish binaries and measurements |
| Padding and chunked streaming | Less metadata leakage | Bandwidth and time to first visible token |
What to do next
- Draw your request path and mark every hop that holds plaintext. If any hop outside the TEE does, fix that before anything else.
- Name the RATS roles for your system: who verifies, who supplies reference values, and whether clients use the passport or background-check model.
- Generate an HPKE key pair inside the node at boot, commit its hash and a challenge into the report data, and encrypt request bodies to it.
- Publish every production measurement to an append-only log and make clients verify inclusion and consistency before sending.
- Put an independent oblivious relay in front of the service and let clients choose among a small verified subset of nodes.
- Audit the image for persistent storage, content logging and shared caches; make prompt state request-scoped.
- Choose padding buckets and a streaming policy per product tier, and write down the metadata you accept leaking.
- Test failure paths: a debug node, an unlogged measurement, stale firmware and a mismatched key must each cause a visible refusal.