An audit log is a record you expect to be read by someone who does not trust you: an investigator after an incident, an auditor checking a control, a regulator, or a court. For an LLM application, that reader will ask questions ordinary logs cannot answer. What exactly did the model see? Which documents did retrieval put in front of it? Which tool did it call, with what arguments, and who allowed that? Which version of the system prompt and which safety policy were live at the time?

The site's LLM audit logging architecture article covers the pipeline as a whole: capture, redaction, hash chains and WORM storage. This article goes one level down, to the records themselves: the event schema, a lean audit stream beside an encrypted content vault, erasure without breaking immutability, proving nothing is missing, what to query, and how long to keep it.

Advertisement

Audit log, telemetry and traces are different contracts

Your observability stack is not your audit log; the three make different promises.

PropertyTelemetry and metricsDistributed tracesAudit log
Question answeredIs the system healthy?Where did this request spend time?Who did what, with what data, under which rules?
CompletenessSampling is fineSampling is normalEvery in-scope event, provably
MutabilityOverwritten or rolled upDropped after daysAppend-only, tamper-evident
ContentAggregatesTimings, some attributesIdentity, inputs, outputs, decisions, versions
ReadersOn-call engineersEngineersSecurity, compliance, legal, auditors
AccessBroadBroadNarrow, and every read is logged

Reuse trace IDs so audit events join to traces, but never let a sampler or log back-pressure decide whether an audit event exists.

The five questions every event must answer

Design the schema backwards from the investigation. For any consequential action, the record must let a stranger answer five questions without asking the engineers who built the system.

  1. Who. The authenticated principal, the tenant, and the actor on whose behalf it acted. For agents this is two identities: the human who started the session and the service identity the agent used to call tools.
  2. What it saw. The assembled input: system prompt version, user message, retrieved document IDs and their versions, tool results fed back into context, and prior turns.
  3. What it produced. The output, any refusal, and the verdicts of input and output filters.
  4. What happened as a result. Each tool call with its arguments, whether a policy engine or a human approved it, and the downstream effect, such as a ticket ID or a payment ID.
  5. Under which rules. Model ID and version, sampling parameters, policy bundle version, and the deployment or release identifier.

The fifth is most often missed, and without it you cannot tell a fixed bug from a live one.

One LLM request becomes two records: a lean audit event and a separately keyed content vaultGatewayauthn, policy, model callEvent builderhash, reference, classifyContent vaultper-subject key, encryptedKey serviceone key per subjectspansciphertextAudit streamseq no, hashes, idseventAppend-only storechained, object lockCoverage checkgateway count = eventsDetectionqueries over eventsInvestigationevent, then vault lookupRetentionexpire or legal holddecrypt if authorisedErasing a subject = destroying one key. The event stays, the hashes still verify, the content becomes unreadable.Every read of the vault is itself an audit event.
The audit stream carries identities, versions, hashes and references; the content itself sits in a vault encrypted under a key per data subject.
Advertisement

A schema that survives an investigation

One event per state change keeps the stream simple: llm.request, llm.response, tool.call, tool.approval, policy.decision and vault.read. All events from one interaction share an interaction ID and carry the trace ID. Content is never stored inline; it is referenced by a vault ID and pinned by a hash.

{
  "event_id": "01J9Z4K7Q8S2M6T0V3X5Y7Z9AB",
  "event_type": "tool.call",
  "occurred_at": "2026-09-14T10:42:07.311Z",
  "producer": {"service": "support-agent", "instance": "pod-7f9c", "seq": 1840221},
  "interaction_id": "int_5b1e...", "trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
  "actor": {"human": "u_4821", "tenant": "t_acme", "service_identity": "svc-support-agent"},
  "versions": {"model": "provider/model@2026-08-01", "system_prompt": "support-v41",
               "policy_bundle": "pol-2026.09.10", "release": "r-8812"},
  "tool": {"name": "issue_refund", "args_sha256": "9f2c...", "args_ref": "vault:7c1d...",
           "risk_tier": "high"},
  "approval": {"mode": "human", "approver": "u_1177", "decision_event": "01J9Z4K6..."},
  "effect": {"system": "payments", "external_id": "re_3Q8x...", "status": "succeeded"},
  "data_subjects": ["cust_99120"],
  "classification": ["financial", "personal"],
  "prev_hash": "b3a1...", "hash": "e47d..."
}

Four fields carry most of the investigative weight. producer.seq is a per-instance counter that makes gaps detectable. versions answers the fifth question. data_subjects lets you find every event about one person, which you will need for both access requests and erasure. args_sha256 lets anyone holding the vault content prove it is the content that was used, without the audit stream holding it.

Building events: hash, reference, encrypt

Inline prompt text in an audit log creates a permanent, widely replicated copy of every secret and every piece of personal data users type. The pattern that avoids this is envelope encryption with one data key per data subject. Content is encrypted with that subject's key and stored in a vault. The audit event holds only the hash and the vault reference. The sketch below uses the cryptography package's AES-GCM; in production the key map is a key management service, not a dictionary.

import hashlib, json, os, time
from cryptography.hazmat.primitives.ciphers.aead import AESGCM

KEYS = {}        # subject_id -> 256-bit data key (in production: KMS-wrapped)
VAULT = {}       # vault_id -> (subject_id, nonce, ciphertext)

def subject_key(subject):
    return KEYS.setdefault(subject, AESGCM.generate_key(bit_length=256))

def put_content(subject, payload: dict) -> tuple[str, str]:
    raw = json.dumps(payload, sort_keys=True, separators=(",", ":")).encode()
    digest = hashlib.sha256(raw).hexdigest()
    nonce = os.urandom(12)
    ct = AESGCM(subject_key(subject)).encrypt(nonce, raw, digest.encode())
    vault_id = "vault:" + hashlib.sha256(nonce + ct).hexdigest()[:32]
    VAULT[vault_id] = (subject, nonce, ct)
    return vault_id, digest

def tool_call_event(ctx, tool, args, subject, seq):
    ref, digest = put_content(subject, {"tool": tool, "args": args})
    return {
        "event_type": "tool.call", "occurred_at": time.time(),
        "producer": {"service": ctx["service"], "seq": seq},
        "interaction_id": ctx["interaction_id"], "actor": ctx["actor"],
        "versions": ctx["versions"],
        "tool": {"name": tool, "args_sha256": digest, "args_ref": ref},
        "data_subjects": [subject],
    }

def shred(subject):
    # Erasure: the events remain and still verify; the content is unreadable.
    KEYS.pop(subject, None)

Binding the hash as AES-GCM associated data means a vault entry cannot be swapped for different content that decrypts cleanly under the same reference. Content about several subjects needs a per-interaction key that lists every subject, so erasing one is a reviewed act.

Erasure and immutability can coexist

An append-only, object-locked store cannot delete a record, and data protection law can require you to erase personal data. Destroying a key is the common engineering answer, often called crypto-shredding. The event survives, the chain still verifies, the hash still proves what was there, and the content can no longer be read. Whether that satisfies an erasure obligation in your jurisdiction is a legal judgement, not a settled engineering fact, so get your counsel to sign off on the design before you rely on it. Legal hold overrides shredding: a key under hold must not be destroyable by the routine erasure job.

Integrity itself is well covered elsewhere. Chain each event to the hash of the previous one per producer, sign periodic checkpoints with a key the application cannot use, and store the stream under object lock. The ADK Java audit log article has a complete implementation of the chain and its verifier. Here it is enough to note that chaining must happen after hashing the content but must never require the content, or shredding would break verification.

Proving nothing is missing

An audit trail with silent gaps is worse than none, because it creates false confidence. Two cheap checks catch most gaps. First, the per-producer sequence number: any jump means lost events. Second, reconciliation: count model calls at the gateway, which already has metrics, and count llm.request events in the store over the same window. The counts should match exactly.

def find_gaps(events):
    """events: iterable of (producer_instance, seq). Returns missing ranges."""
    last, gaps = {}, []
    for inst, seq in sorted(events):
        prev = last.get(inst)
        if prev is not None and seq != prev + 1:
            gaps.append((inst, prev + 1, seq - 1))
        last[inst] = seq
    return gaps

def reconcile(gateway_counts, audit_counts, tolerance=0):
    """Both: dict minute -> count. Report minutes where audit lags the gateway."""
    return {m: (g, audit_counts.get(m, 0)) for m, g in gateway_counts.items()
            if g - audit_counts.get(m, 0) > tolerance}

Reconcile closed windows, such as minutes ending ten minutes ago, so late events do not alarm. Treat any gap as an incident until explained, and key sequences by instance and start time because restarts reset counters.

Detection queries worth running every day

Audit events are structured, so detection is SQL. Three queries catch common LLM-specific abuse. The first finds bulk retrieval, where one user surfaces far more distinct documents in a day than the typical user does, a typical sign of data harvesting through a copilot.

-- 1. Users retrieving unusually many distinct documents in a day
SELECT actor_human, COUNT(DISTINCT doc_id) AS docs
FROM audit_retrievals
WHERE occurred_at >= now() - interval '1 day'
GROUP BY actor_human
HAVING COUNT(DISTINCT doc_id) > 5 * (
  SELECT percentile_cont(0.95) WITHIN GROUP (ORDER BY daily_docs) FROM user_daily_baseline);

-- 2. High-risk tool calls retried after a policy denial in the same interaction
SELECT d.interaction_id, d.tool_name, MIN(c.occurred_at) AS retried_at
FROM policy_decisions d JOIN tool_calls c
  ON c.interaction_id = d.interaction_id AND c.tool_name = d.tool_name
 AND c.occurred_at > d.occurred_at
WHERE d.decision = 'deny' AND c.risk_tier = 'high'
GROUP BY d.interaction_id, d.tool_name;

-- 3. External-content retrieval followed by an outbound tool call
SELECT r.interaction_id
FROM audit_retrievals r JOIN tool_calls t ON t.interaction_id = r.interaction_id
WHERE r.source_trust = 'external' AND t.tool_name IN ('send_email', 'http_post')
  AND t.occurred_at > r.occurred_at;

The third query is the audit-log signature of indirect prompt injection: untrusted text entered context, then the agent sent data out. Expect false positives; route hits to review. The first query's thresholds are illustrative.

Retention and the regulatory floor

Retention has a floor and a ceiling. The floor comes from law and contracts. Under the EU AI Act, providers of high-risk AI systems must keep the automatically generated logs under their control for at least six months (Article 19), and deployers carry the same six-month minimum (Article 26(6)), unless other Union or national law says otherwise. Article 12 requires high-risk systems to support automatic logging; its specific list of minimum fields, such as each period of use and the people who verified results, applies to remote biometric identification systems rather than to every high-risk system. See EU AI Act compliance for which systems are in scope. Sector rules in finance and healthcare often require longer.

The ceiling comes from data minimisation: keep personal data no longer than needed. The two-tier design lets you set them separately. Lean events can be kept for years because they hold identifiers, versions and hashes. Vault content can expire at the floor unless a legal hold applies. Size it before you commit: a service doing 2 million model calls a day with about 8 KB of assembled context and output per call writes roughly 16 GB of vault content a day, about 480 GB a month, and close to 2.9 TB over a six-month window before compression.

Worked example: a customer says the assistant leaked their order history

A customer reports that another user's chat showed details of their orders. With the schema above, the investigation is a sequence of queries rather than guesswork. Search data_subjects for the complainant's ID across the reported window. This returns three interactions belonging to other users in which the customer's record was a retrieved document. The versions field shows all three ran with retrieval index build idx-0912, deployed that morning. The policy.decision events show the tenant filter returned allow. An authorised investigator reads the three vault entries, and each read emits a vault.read event. The content confirms the retrieved chunk carried the wrong tenant tag.

The root cause was an indexing job, not the model. Because events carried index versions and data subjects, the team could list every affected interaction for breach notification and prove through reconciliation that none were missing. The AI forensics article covers the wider investigation method.

Failure modes

  • Audit as a side effect of logging. Events flow through the application log pipeline, which drops under back-pressure. Use a dedicated, acknowledged path and decide explicitly whether high-risk actions fail closed when the audit write fails.
  • Raw content inline. The audit store becomes the largest copy of personal data you hold, readable by every analyst. Keep content in the vault; scrub secrets with the patterns in LLM PII protection before encryption.
  • Missing versions. You can see what happened but not under which prompt, policy or index, so you cannot tell whether it can still happen.
  • Single identity for agents. Logging only the service account hides which human started the chain of actions.
  • Unaudited reads. Investigators read the vault without leaving a trace, which turns the audit system into an exfiltration channel.

Trade-offs

Synchronous audit writes make the audit trail as reliable as the action and add latency to every call; asynchronous writes through a durable queue are faster but need reconciliation to prove completeness. Storing full content makes investigations easy and raises breach impact; storing only hashes makes the log safe and investigations impossible. The vault sits between them, at the cost of running a key service. Per-subject keys make erasure precise but multiply key counts; per-tenant keys are cheaper and make single-person erasure impossible. Most teams choose acknowledged asynchronous writes for ordinary calls and synchronous, fail-closed writes for high-risk tool calls.

What to do next

  1. Write the five questions on one page and map each to a field in your current logs; every unanswered question is a schema gap.
  2. Add versions (model, system prompt, policy bundle, retrieval index, release) to every model and tool event.
  3. Record two identities for agent actions: the human and the service identity.
  4. Move prompt and tool content out of the event into an encrypted vault, keyed per subject, referenced by ID and hash.
  5. Add per-instance sequence numbers and a daily reconciliation job against gateway call counts.
  6. Schedule the three detection queries and route hits to a review queue.
  7. Agree retention floors and ceilings per data class with legal, and implement legal hold before you implement shredding.
  8. Run one tabletop investigation using only the audit trail, and fix whatever you could not answer. Use the architecture overview to check the surrounding pipeline.
Key takeaway: An LLM audit log is a record built for a sceptical reader. Each event must say who acted and on whose behalf, what the model saw and produced, what happened as a result, and under which model, prompt, policy and index versions. Keep events lean and put prompt and tool content in a vault encrypted per data subject, so that destroying one key erases content without breaking an append-only, chained store. Prove completeness with sequence numbers and reconciliation, query the events daily for bulk retrieval, denied-then-retried actions and injection patterns, and set retention from legal floors such as the EU AI Act's six-month minimum for high-risk systems.