Encryption at rest and in transit leave one gap: data in use. While a model serves a request, the prompt, the weights and the KV cache sit in plaintext in memory, readable by anyone who controls the host, from a cloud operator to a compromised hypervisor. Confidential computing closes that gap with hardware. The CPU runs the workload in a confidential virtual machine whose memory the host cannot read, the GPU runs in a confidential mode that blocks host access to its protected memory, and remote attestation lets someone far away check exactly what code is running before trusting it with a secret.

The overview of the options, comparing TEEs with homomorphic encryption and multi-party computation, is in secure inference for LLMs, in depth, including a basic key-release check. This page goes one level down, for the engineer who has to build and run such a system. It covers the full bring-up sequence, what each measurement actually covers, how to produce images whose measurements you can predict, the data path and where the overhead comes from, the hard limits of multi-GPU deployments, and the day-two problems that sink most projects: firmware updates, policy churn and debugging a machine you are not allowed to look inside.

Advertisement

Who is protected from whom

Name the parties first. The data owner sends prompts and documents and wants them unreadable to everyone else. The model owner supplies weights and wants them unextractable. The operator runs the hardware and wants neither liability. Confidential computing removes the operator, the host software stack and other tenants from the trusted set. It does not remove the hardware vendors, whose keys sign the attestation evidence, and it does not remove whoever wrote the code inside the CVM.

Three shapes come up: an enterprise running its own fine-tuned model on rented GPUs, where attestation gates the weight key; a provider serving an open model to users who want proof prompts are not logged, where clients verify a published measurement before sending anything; and a vendor deploying proprietary weights into a customer's environment, where the broker releases weights only to measured images. Weight theft more broadly is covered in model extraction and broader node hardening in LLM infrastructure security.

The bring-up sequence, step by step

Bringing up a confidential LLM server: every step either measures something or refuses to continue1. CPU launches the CVMSEV-SNP / TDX launch digest2. Measured bootfirmware, kernel, cmdline, initrd3. Verified root filesystemdm-verity root hash in cmdline4. Driver opens SPDM sessionGPU identity + session keys5. GPU attestationdevice-signed report, CC mode on6. GPU set readyonly after evidence verifies7. Composite evidenceCPU report + GPU report + nonce8. Verifier + key brokerpolicy: image, TCB, fw, debug off9. Weight key releasedwrapped to an in-CVM key10. Serve: TLS terminates inside the CVM, report_data binds the TLS keyclients verify the same evidence before sending promptsCPU to GPU data pathencrypted bounce buffers in shared memory, AES-GCM over PCIeOutside the boundaryhost OS, hypervisor, BMC, operators: ciphertext onlyA failed check at any step means no key, so a mis-measured server cannot decrypt weights or receive prompts.
Ten steps from launch to serving. Steps 1 to 3 build the CPU-side measurement, 4 to 6 establish and verify the GPU, 7 to 9 turn evidence into a key, and 10 lets clients check the same evidence.

CPU launch. With AMD SEV-SNP, the processor computes a launch measurement over the initial guest memory, typically the virtual firmware, and signs attestation reports with a chip-specific key chained to AMD's certificates; the report also carries TCB version numbers and 64 bytes of report data chosen by the guest. With Intel TDX, the initial contents are measured into MRTD, and four runtime measurement registers (RTMRs) can be extended later in boot, with the quote likewise carrying 64 bytes of report data. Many cloud confidential VMs add a virtual TPM and measure boot into it, so check which evidence your platform actually exposes.

Measured boot and root filesystem. The launch measurement covers the firmware, so the firmware must measure what it loads next: kernel, command line and initrd. The usual way to cover gigabytes of root filesystem is dm-verity: put the verity root hash on the measured kernel command line, and every block read afterwards is checked against it. A writable or unverified root filesystem makes the whole measurement meaningless, because an attacker with disk access changes the server after it is measured.

GPU session and attestation. In confidential mode the NVIDIA driver inside the CVM uses the SPDM protocol to authenticate the GPU and set up session keys, then requests a GPU attestation report signed by a device key whose certificate chains to NVIDIA. The report covers firmware measurements and confidential-mode state. NVIDIA offers a remote service (NRAS) and a local verifier in its open-source nvtrust tooling to check these reports, including revocation. The GPU does not accept work until software marks it ready, and deployments should do that only after the GPU evidence verifies.

Composite evidence and key release. A relying party needs both halves. A genuine CPU report with an unverified GPU means the weights may be copied into a GPU the host controls, and a genuine GPU with an unmeasured CVM means unknown code drives it. The key broker checks both reports, a fresh nonce, and the binding of the report data to a public key generated inside the CVM, then wraps the weight key to that public key. Finally the server publishes the same evidence so clients can verify before sending prompts; binding the TLS key into the report data makes the TLS session itself attested.

Advertisement

Building images you can measure

Attestation compares a measurement against an expected value, so you need to know that value before the server boots. That turns image building into a security control. The build must be reproducible: pinned base images by digest, pinned package versions, fixed timestamps and no secrets baked in. CI builds the image, computes the launch measurement with the vendor's measurement tooling for your exact virtual firmware and vCPU configuration, and publishes both to the policy repository in one signed change.

Keep the policy as data, reviewed like code.

{
  "policy_version": "2026-10-02.1",
  "cpu": {
    "tee": "sev-snp",
    "allowed_launch_measurements": [
      "9f1c...e2  # image llm-serve 4.12.0, built by CI run 18832",
      "41aa...07  # image llm-serve 4.11.3, kept for rollback until 2026-10-16"
    ],
    "min_tcb": "from the current AMD security bulletin; raised after the fleet is patched",
    "debug_allowed": false,
    "migration_agent_allowed": false
  },
  "gpu": {
    "vendor": "nvidia",
    "cc_mode_required": true,
    "min_driver": "pinned-by-ci",
    "min_vbios": "pinned-by-ci",
    "allowed_topologies": ["single", "ppcie-8"],
    "verifier": "remote-or-local, must check revocation"
  },
  "binding": {"report_data": "sha256(tls_public_key || nonce)"},
  "release": {"key_id": "weights/llama-ft-v7", "wrap_to": "tls_public_key"}
}

Note the second allowed measurement: the previous image remains valid for a fixed rollback window, then expires. Note also what is absent. There is no wildcard, debug is refused, and minimum TCB versions are explicit, so a host running unpatched firmware with a known vulnerability is refused even when the image is correct. The key broker enforces this policy, not the server: the server's own boot gate exists to fail fast and log clearly, but it is inside the thing being verified and cannot be its own judge.

# Boot-time gate run as the first service inside the CVM. Function names are placeholders
# for your attestation SDK and key broker client; the order and the refusals are the point.
def bring_up():
    tls_key = generate_keypair_in_memory()             # never written to disk
    nonce = broker.get_nonce()                         # freshness from the relying party

    gpu_reports = []
    for gpu in list_gpus():
        if not gpu.cc_mode_enabled():
            fail("GPU %s not in confidential mode" % gpu.id)
        gpu_reports.append(gpu.attestation_report(nonce))   # signed by the device key

    if not local_gpu_verifier.verify(gpu_reports, policy.gpu):    # chain, revocation, fw
        fail("GPU evidence rejected")
    for gpu in list_gpus():
        gpu.set_ready()                                # accept work only now

    cpu_report = cpu_tee.report(report_data=sha256(tls_key.public + nonce))
    wrapped = broker.release(policy.release.key_id,
                             evidence={"cpu": cpu_report, "gpu": gpu_reports, "nonce": nonce},
                             wrap_to=tls_key.public)   # broker re-verifies everything itself
    weights_key = tls_key.unwrap(wrapped)
    start_server(model=decrypt_stream("weights.enc", weights_key), tls_key=tls_key,
                 evidence_endpoint=(cpu_report, gpu_reports))

def fail(reason):
    log_to_serial_console(reason)                      # no secrets in this message
    halt()                                             # no key means nothing to serve

The data path and what it costs

On Hopper, the GPU cannot read the CVM's encrypted private memory. Every host-to-device copy therefore goes through bounce buffers: the driver encrypts data with an authenticated cipher (AES-GCM) into shared memory, the GPU copies it across PCIe and decrypts it internally, and results return the same way. Computation on the GPU runs at full speed; the cost lands on transfers and on CPU cycles spent encrypting. On H100 the GPU's memory is not encrypted; the GPU blocks host access to its protected region instead. The H100 architecture page covers the rest of that chip.

That predicts where overhead shows up. The H100 benchmark study (arXiv 2409.03992), which ran vLLM on a single H100 with confidential mode on and off, reported throughput overheads of 6.85 percent for Llama-3.1-8B, 4.58 percent for Phi-3-14B-128k and essentially none for Llama-3.1-70B, but time-to-first-token overheads of about 19 and 18 percent for the two smaller models. Prefill moves the prompt across PCIe before computation starts, so latency-sensitive small-model workloads feel it most, while large models spend their time computing on the GPU.

  • Keep data on the GPU. Offloading the KV cache to host memory, which is common for long contexts, now costs encryption on every swap. Size the GPU memory so swapping is rare.
  • Sample on the GPU. Copying logits to the host each step adds encrypted transfers per token.
  • Batch. Larger batches amortise fixed transfer costs; measure TTFT and tokens per second separately.
  • Budget CPU. Bounce-buffer encryption runs on CVM vCPUs, so an undersized CVM throttles the GPU.

Multi-GPU: the hard limit

Large models need several GPUs, and that is where Hopper confidential computing is weakest. In single-GPU passthrough each GPU is its own trust domain. For eight-GPU HGX Hopper systems NVIDIA added a Protected PCIe mode that passes all eight GPUs of the board to one CVM. CPU-to-GPU traffic is encrypted through bounce buffers as before, but GPU-to-GPU traffic over NVLink and NVSwitch is not encrypted on Hopper. Tensor-parallel activations therefore cross the NVLink fabric in plaintext, and the security argument relies on the eight GPUs and switches being attested together as one physical unit.

NVIDIA's Blackwell and Hopper security whitepaper states that Blackwell's multi-GPU pass-through mode also encrypts the NVLink path, for up to eight GPUs per CVM, and that compatible CPUs and firmware can replace bounce buffers with inline TDISP/IDE encryption. Across nodes, traffic needs its own encryption and mutual attestation at the application layer, which many serving stacks lack. If your model fits on one 8-GPU board, stay there.

Confidential containers and Kubernetes

Teams already running inference on Kubernetes rarely want to manage one-off VMs. The Confidential Containers project runs each pod inside its own lightweight confidential VM through Kata Containers, with an attestation agent that fetches secrets only after the pod's environment verifies. What gets measured shifts from a disk image to a small guest plus pulled container images, so image digests belong in policy, and the pod specification, including environment variables and mounts, is untrusted input from the control plane. GPU passthrough support here depends on GPU, driver and runtime versions; validate it yourself.

Failure modes

  • Telemetry leaks. The CVM is protected, but logs, traces, crash dumps and metrics labelled with prompt fragments leave it in plaintext. Treat every egress path as a disclosure channel and allow-list what leaves.
  • Debug or unmeasured modes accepted. A policy that skips the debug flag, accepts any measurement or skips GPU evidence turns attestation into theatre.
  • Evidence not bound to the channel. Without report data tied to the TLS key and a nonce, a genuine report can be replayed in front of an impostor.
  • Measurement drift. A firmware update, a new vCPU count or a different virtual firmware build changes the launch measurement and every server fails to get keys at once. This is the most common outage.
  • Stale TCB. Hosts that miss microcode or firmware patches are correctly refused, which looks like a fleet-wide incident if nobody tracks TCB levels.
  • Side channels and metadata. Outsiders still see how large requests are, when they arrive and how many tokens return, and shared prefix caches can reveal whether another tenant sent the same text.

Day-two operations

Run policy changes as deployments. Add the new measurement to the allow-list before rolling the new image, roll gradually, then remove the old measurement once the rollback window closes. Track firmware, VBIOS, driver and TCB versions per host in inventory, and stage vendor updates the same way: canary hosts first, policy minimums raised only after the fleet is patched. Alert on key-release denials by reason code: a spike is an attack or a broken rollout.

Debugging is harder because you cannot attach to a production CVM. Build a debug variant of the image that is identical except for its debug flag and is never in the allow-list, reproduce problems with synthetic data, and expose only health and structured, prompt-free diagnostics from production. Publish image measurements and their source so clients can audit them.

Trade-offs

Confidential computing gives strong protection against host and operator access with single-digit throughput cost for large models, but it narrows your hardware and region choices, pins your driver and firmware stack, slows boot by the weight-decryption time and adds an attestation service on the startup path. A prompt-logging bug inside the measured image is still a leak. For most LLM workloads that need data-in-use protection today, a well-built TEE deployment is the practical choice, provided the image and policy pipeline is treated as security-critical code.

What to do next

  1. Write down the parties and which of them the deployment must exclude, then check that confidential computing actually excludes them.
  2. Make your inference image reproducible, add dm-verity for the root filesystem and compute its launch measurement in CI.
  3. Store the attestation policy as reviewed data with explicit measurements, TCB minimums and debug refused.
  4. Implement the boot gate: GPU evidence verified before ready, composite evidence bound to an in-CVM key, and a key broker that re-verifies everything.
  5. Benchmark TTFT and throughput with confidential mode on and off for your model, batch sizes and context lengths.
  6. If you need more than one GPU, document the NVLink assumption for Hopper or plan for hardware with NVLink encryption.
  7. Allow-list every egress path, then rehearse a firmware update and an image rollout through the policy pipeline before going live.
Key takeaway: Confidential computing for LLMs combines a measured confidential VM, a GPU in confidential mode verified over SPDM, and a key broker that releases weights only to composite evidence that matches policy. The cryptography is the easy part. The real work is reproducible images, measurements computed in CI, explicit TCB and debug policy, and evidence bound to the TLS key. Expect modest overhead concentrated in PCIe transfers, plaintext NVLink on Hopper multi-GPU boards, and outages caused by measurement drift unless policy changes ship like deployments.