Security teams describe data in three states. At rest it sits on disk or in object storage, and AES with a managed key is routine. In transit it crosses a network, and TLS is routine. In use it is being processed, which means it is sitting in RAM, in CPU caches, in GPU memory, in a KV cache, in a log line or in a crash dump, and for most of computing history it had to be plaintext there because processors cannot add encrypted numbers. Encryption in use is the set of techniques that narrow that gap: hardware that encrypts memory so the host cannot read it, keys that are only released to a measured workload, and software designs that keep the most sensitive values out of the processing path entirely.

LLM systems make this state unusually important. A prompt can contain an entire medical record, a contract or a customer's account history, and the model must see the content to be useful. This article maps where that plaintext actually lives in an LLM pipeline, explains what each in-use control does and does not cover, and works through a design you can build. The confidential-serving architecture itself, along with homomorphic encryption and multi-party computation, is covered in secure inference for LLMs; here the focus is the data lifecycle around it.

Advertisement

Why in use is the hard state

Encryption at rest and in transit share a property: the data is not being computed on, so it can stay ciphertext until the moment it is needed. In use, the computation needs the values. There are three ways out. You can compute on ciphertext directly (homomorphic encryption, practical today for narrow operations and far too slow for a full transformer forward pass). You can split the computation between parties who each see only shares (MPC, with heavy communication costs). Or you can let the computation see plaintext but put it inside a boundary the attacker cannot reach, which is what trusted execution environments do. For production LLM serving, the third option plus aggressive minimisation of what enters the boundary is what actually ships.

That framing gives you the right question. Not "is the data encrypted" but "in which places does plaintext exist, for how long, and who can read each place". The answer is always a list, and every item on it is either protected by a boundary, minimised, or an accepted risk you have written down.

Where plaintext lives in an LLM pipeline

Follow one request through a typical hosted deployment and record every place its content exists unencrypted:

LocationWhat is thereWho can normally read itPrimary control
Gateway or load balancerFull request after TLS terminationGateway operators, its memory and logsTerminate TLS inside the boundary, or log nothing
Application server RAMPrompt, retrieved documents, outputHost root, hypervisor, memory dumpsConfidential VM; minimise content
Swap and core dumpsPages of process memoryAnyone with disk accessDisable or encrypt; LimitCORE=0
Logs and tracesPrompts copied for debuggingEveryone with log access, often for monthsStructured logging without content fields
GPU memoryWeights, activations, KV cacheHost drivers, other tenants on shared GPUsGPU confidential mode; no unsafe sharing
PCIe transfersInputs and outputs moving CPU to GPUAnyone who can observe the bus or DMAEncrypted bounce buffers in CC mode
Prompt or prefix cachesReusable prefixes, sometimes across usersOther requests via timing or reusePer-tenant cache keys; TTLs
Vector store and RAGChunks and embeddingsDatabase operators and backupsEncrypt at rest, filter by ACL at query time
Where a prompt exists as plaintext on its way through an LLM serviceClientTLS to the edgeGatewayTLS ends: plaintextPseudonymiserPII to tokensToken vaultkeys in KMS or HSMConfidential VM boundary: CPU memory encrypted per VM, measured at bootInference serverprompts, logs, dumpsBounce bufferAES-GCM over PCIeGPU in CC modeweights, KV cachetokens onlyAttestation agentevidence of the VMKMS key releasepolicy on measurementdata keyHost and operatoroutside the boundaryPlaintext still exists at the gateway, inside the VM and inside the GPU package; the design questionis who can reach each of those places, and how little sensitive content you let arrive there.
The boundary covers the VM and the GPU; the pseudonymiser decides how much sensitive content ever enters it.
Advertisement

What memory encryption actually does

Confidential VMs on current server CPUs encrypt guest memory with keys the hypervisor never sees. On AMD SEV-SNP, the AMD Secure Processor generates a per-VM key and loads it into an AES-128 engine in the memory controller, so DRAM contents are ciphertext to the host; SNP adds the Reverse Map Table, which records which guest owns each page and stops the hypervisor from remapping or replaying guest memory. Intel TDX encrypts trust-domain memory with AES-XTS through the multi-key memory encryption engine and adds integrity protection against software and, in its cryptographic mode, against tampered memory. Both measure the initial guest image and sign that measurement in an attestation report.

For GPUs, NVIDIA's confidential computing mode on H100 and later puts the GPU inside the trust boundary: the driver in the confidential VM establishes a session with the GPU, and data moving across PCIe goes through bounce buffers in shared memory, encrypted with AES-GCM. The GPU produces its own attestation report. Treat the GPU package as the boundary and do not build your threat model on finer claims about how on-package memory is protected; descriptions differ, so read the current documentation for your generation.

What none of this covers is equally important. The code inside the boundary sees plaintext and can log it, cache it or send it anywhere. A bug that writes prompts to a log inside a confidential VM ships the log out as plaintext. Side channels remain a research area. And the boundary only helps if something checks the attestation before handing over secrets, which is the next section.

Keys that only an attested workload can get

Memory encryption protects data that is already inside the VM. The data usually arrives encrypted, under a data key from a key management service, and the whole design rests on the KMS releasing that key only to the right workload. The pattern is envelope encryption with an attestation condition: the workload produces a signed report of its measurement, the KMS or an attestation service verifies it, and the key policy permits decryption only when the measurement matches an approved build.

AWS expresses this directly for Nitro Enclaves. The enclave passes a signed attestation document as the Recipient of a KMS call, and the key policy can require that the enclave image digest, which corresponds to PCR0, matches an approved value:

{
  "Sid": "DecryptOnlyFromMeasuredEnclave",
  "Effect": "Allow",
  "Principal": {"AWS": "arn:aws:iam::111122223333:role/inference-enclave-parent"},
  "Action": ["kms:Decrypt", "kms:GenerateDataKey"],
  "Resource": "*",
  "Condition": {
    "StringEqualsIgnoreCase": {
      "kms:RecipientAttestation:ImageSha384": "<sha384 of the approved enclave image (PCR0)>"
    }
  }
}

With the condition in place, KMS returns the plaintext key encrypted to the enclave's public key, so even the parent instance that relays the call cannot read it. Azure and Google offer comparable flows for confidential VMs through their attestation services and key release policies; the names differ but the shape is the same. Three rules make it hold up: pin measurements, not instance identities; keep the list of approved measurements in version control with the build that produced each one; and alert when a decrypt is refused, because that is either an attack or a deploy that forgot to register its new measurement.

Shrinking plaintext in software

The strongest in-use control is not to have the value in use at all. Most LLM tasks do not need a customer's real account number, email address or national ID; they need to know that one exists and to refer to it consistently. Replace identifiers with stable tokens before the text reaches the model, keep the mapping outside the model's reach, and restore the real values in the response only where the caller is entitled to see them.

import hmac, hashlib, re

EMAIL = re.compile(r"[\w.+-]+@[\w-]+\.[\w.-]+")
ACCT = re.compile(r"\b\d{10,12}\b")

class Pseudonymiser:
    """Replace identifiers with stable tokens before text reaches the model.

    The HMAC key comes from the KMS at startup and never leaves this process.
    The reverse map lives only for the request and is never logged."""

    def __init__(self, hmac_key: bytes):
        self.key = hmac_key

    def _token(self, kind: str, value: str) -> str:
        digest = hmac.new(self.key, f"{kind}:{value}".encode(), hashlib.sha256)
        return f"[{kind}_{digest.hexdigest()[:10]}]"

    def protect(self, text: str):
        reverse = {}
        def sub(kind):
            def repl(m):
                tok = self._token(kind, m.group(0))
                reverse[tok] = m.group(0)
                return tok
            return repl
        text = EMAIL.sub(sub("EMAIL"), text)
        text = ACCT.sub(sub("ACCT"), text)
        return text, reverse

    @staticmethod
    def restore(text: str, reverse: dict) -> str:
        for tok, value in reverse.items():
            text = text.replace(tok, value)
        return text

# p = Pseudonymiser(key_from_kms)
# safe, rev = p.protect("Refund 4711002233 for ana@example.com")
# -> "Refund [ACCT_...] for [EMAIL_...]"; the model never sees the raw values

Keyed HMAC tokens are deterministic, so the same email gets the same token across a conversation and the model can reason about "the same customer", but nobody without the key can reverse or dictionary-attack them. Regexes catch structured identifiers; names and free-text addresses need a named-entity detector, and PII leakage in LLM systems covers how to evaluate detectors and what they miss. Where analytics need equality joins on protected fields, deterministic encryption or tokenisation in the data store gives you that without exposing values, at the cost of revealing which records share a value.

Pseudonymisation also shrinks every other row of the plaintext table at once: the logs, the KV cache, the prompt cache and the vendor's systems now hold tokens, not identifiers.

Process hygiene inside the boundary

Memory encryption stops the host from reading RAM. It does nothing about the copies your own process makes. Core dumps write the whole address space to disk, swap writes pages, debug logging copies prompts into a pipeline with long retention, and exception trackers capture local variables including the request body. Turn these off deliberately:

# systemd unit for the inference service: no core files, no swap-backed secrets
[Service]
LimitCORE=0
MemorySwapMax=0
Environment=PYTHONFAULTHANDLER=1
# kernel side, on the host image
#   kernel.core_pattern=|/bin/false     (or route to an encrypted, access-controlled store)
#   vm.swappiness=0 and encrypted swap if swap exists at all

Two further habits matter. First, define the log schema so content fields do not exist: request ID, tenant, token counts, latency, model version and a hash of the prompt for deduplication, but never the prompt. Second, accept that managed runtimes cannot reliably zeroise memory. Python strings are immutable and copied freely, so "wipe the secret after use" is not achievable; the realistic goal is short-lived processes, no persistence paths, and secrets fetched from a proper secrets manager rather than baked into images or environment dumps.

On GPUs, avoid sharing a device between tenants unless the sharing mode gives memory isolation, and remember that caches keyed only on prefix text can leak across users through timing. Key prefix caches by tenant.

Worked example: an insurance claims assistant

An insurer wants an LLM to summarise claim files that include policyholder names, policy numbers, medical notes and bank details, served from a GPU cluster the insurer does not physically control. Walking the plaintext table produces this design.

  1. Claim documents are stored encrypted with per-tenant data keys. Their key policy releases them only to inference VMs whose measurement matches the approved image.
  2. A pseudonymiser runs inside the confidential VM before prompt assembly. It replaces policy numbers, bank details and names with HMAC tokens. Medical content stays, because the summary needs it.
  3. Inference runs on GPUs in confidential mode, attached to the same VM. The VM checks the GPU's attestation before loading any data.
  4. Logs carry claim ID, token counts and latency. Prompts and outputs are never logged, and the prompt cache is keyed by tenant.
  5. The summary leaves the VM with tokens in place. The claims application, which is authorised to see identities, restores them for the adjuster.

What remains is a short, explicit list. Medical text is in plaintext inside an attested boundary, and the insurer accepts that. Identifiers never enter the model. The risk of a compromised inference image is handled by build provenance and measurement pinning. Each line on that list has an owner.

Failure modes

SymptomCauseFix
Prompts found in a log indexDebug logging or an exception tracker captured request bodiesContent-free log schema; scrub tracker payloads; test with canary strings
Key release works for any VMPolicy checks the role, not the measurementAdd the attestation condition; test that an unapproved image is refused
Deploy fails with decrypt deniedNew image measurement not registeredRegister measurements in the release pipeline before rollout
Identifiers still reach the vendorDetector missed free-text names or new formatsMeasure recall on labelled samples; add NER; alert on regex hits in outbound text
Crash leaks memory to diskCore dumps enabled on the host or containerLimitCORE=0, encrypted dump store if dumps are needed
Confidential mode silently offInstance type or driver fell back to normal modeVerify attestation at startup and refuse to serve without it

Trade-offs

Confidential VMs and GPUs cost some throughput, mostly from encrypting CPU-GPU transfers, and narrow your choice of instance types and drivers. Pseudonymisation costs some answer quality when the model would have used the real value, and needs a detector you maintain. Attestation-gated keys add a release step to every deploy. Against these, the alternative is trusting every operator, hypervisor, log pipeline and backup that touches the data, and that list is long. Start with the cheap controls (content-free logs, pseudonymisation, no dumps), then add hardware boundaries where the data or the hosting arrangement demands them. LLM infrastructure security covers the rest of the node and control-plane hardening that these controls depend on.

Key takeaway: <p><strong>What to do next.</strong> Encryption in use is a plaintext inventory plus three tools: hardware boundaries, attestation-gated keys and minimisation. Use all three, and write down what remains.</p><ol><li>Trace one real request and list every place its content exists unencrypted, with who can read each place.</li><li>Make the log schema content-free and test it with a canary string sent through the system.</li><li>Disable core dumps and swap on inference hosts, and key any prompt cache by tenant.</li><li>Pseudonymise identifiers before prompt assembly with a keyed HMAC whose key comes from the KMS.</li><li>Where hosting is untrusted, run in confidential VMs and GPUs and verify attestation at startup.</li><li>Bind data-key release to approved measurements and alert on every refused decrypt.</li></ol>