Credentials reach language models constantly. Users paste configuration files and stack traces into chat. Retrieval pulls wiki pages where someone once wrote down a database password. Agents run shell commands whose output includes environment variables. Models occasionally reproduce key-shaped strings from training data or from earlier in the conversation. Every one of those secrets can then flow into the model provider's request, your logs, your traces, an evaluation dataset or another user's answer.
A secret scanner is the detector that finds credentials in text so that something can be done about them. This page explains how scanners work inside, from keyword prefilters through entropy to checksums and live verification, then shows where to put them on an LLM application's data paths, how to scan a streaming response without letting a split secret through, what to do with a finding, which open-source tools to use and how to keep false positives at a level people will tolerate. Keeping secrets out of the context window by design is covered separately in secrets management architecture for LLM applications; this page is about the detector.
Why LLM traffic needs its own scanning
Classic secret scanning looks at code diffs. LLM traffic is high volume, mostly natural language, partly generated, often streamed token by token, and copied into many stores at once: the provider's API, your request log, an observability platform, a feedback queue and sometimes a fine-tuning dataset. A secret that enters at one point is replicated within seconds.
There are also LLM-specific ways for a secret to leave. A prompt-injected agent can be told to print its environment. A model can be tricked into repeating a credential that a tool result placed in its context. The scanner does not prevent any of these by itself, but it is the control that notices, and it is the last chance to redact before the text is stored or shown.
Anatomy of a detector
Every serious scanner layers the same techniques in order of cost: a keyword prefilter that skips a rule unless a cheap substring such as a token prefix appears, a regular expression for the credential's structure, an entropy filter for generic rules, a checksum where the format has one, and finally verification against the issuing service. How each stage decides, and how scanning works on repositories from pre-commit hooks to history scans, is covered in secrets scanning architecture. This section recaps only what matters for LLM traffic.
Prefixed formats are the scanner's best friend. AWS access key IDs start with AKIA for long-term keys and ASIA for temporary STS keys, followed by 16 upper-case alphanumerics; AWS's own documentation uses AKIAIOSFODNN7EXAMPLE as a sample. In 2021 GitHub moved to prefixed tokens, ghp_ for personal access tokens, gho_ for OAuth, ghu_ for user-to-server, ghs_ for server-to-server and ghr_ for refresh tokens, with an underscore separator chosen because it cannot appear in base64 or hex identifiers. GitHub's tokens also end in a 32-bit CRC32 checksum encoded in base62 in the last six characters, so a scanner can reject a made-up token offline without asking GitHub. When you issue credentials yourself, copy this design: a unique prefix plus a checksum makes your tokens nearly free of false positives to scan.
import math
import re
from collections import Counter
from dataclasses import dataclass
@dataclass(frozen=True)
class Rule:
name: str
keywords: tuple # cheap substring prefilter; skip the regex if none appear
pattern: re.Pattern
min_entropy: float = 0.0
RULES = [
Rule("aws-access-key-id", ("AKIA", "ASIA"), re.compile(r"\b(?:AKIA|ASIA)[A-Z0-9]{16}\b")),
Rule("github-token", ("ghp_", "gho_", "ghu_", "ghs_", "ghr_"),
re.compile(r"\bgh[pousr]_[A-Za-z0-9]{36}\b")),
Rule("private-key-block", ("PRIVATE KEY",),
re.compile(r"-----BEGIN [A-Z ]*PRIVATE KEY-----")),
# Generic rule: an assignment to a secret-sounding name, gated on entropy.
Rule("generic-assignment", ("key", "secret", "token", "passw"),
re.compile(r"(?i)\b[\w.-]{0,40}(?:api_?key|secret|token|passw(?:or)?d)[\w.-]{0,40}\s*[:=]\s*['\"]?([A-Za-z0-9+/_\-]{20,120})"),
min_entropy=3.5),
]
def shannon_entropy(s: str) -> float:
n = len(s)
return -sum(k / n * math.log2(k / n) for k in Counter(s).values())
def find(text: str):
for rule in RULES:
if not any(k.lower() in text.lower() for k in rule.keywords):
continue
for m in rule.pattern.finditer(text):
value = m.group(m.lastindex or 0)
if rule.min_entropy and shannon_entropy(value) < rule.min_entropy:
continue
yield rule.name, m.start(), m.end()The generic assignment rule needs the entropy gate, and the gate needs care: the entropy of a short string is capped at the log of its length, so a 20-character random token cannot score above about 4.3 bits per character. Treat 3.5 as a starting point to tune, not a constant. Every rule also has a bounded length, which the streaming redactor below depends on.
LLM output adds a false-positive source that code repositories rarely have: models write plausible-looking example credentials. Ask for a sample config and the answer may contain a key-shaped string that matches a prefix rule. Checksums reject most of these offline, which is one more reason to prefer formats that carry one. For the rest, redact anyway, since users copy examples into real files, but do not page anyone until verification says the key is live.
Where to scan an LLM application
Scan at every point where text crosses a trust boundary, and act before the text is persisted. In practice that is five places.
- Ingress, before the context is sent to a model: user prompts, retrieved documents and tool results. Redacting here keeps secrets out of the provider's logs and out of the model's view.
- Egress, on the model's response before the user or calling service receives it. This catches secrets the model repeats from context and key-shaped output that should never be shown.
- Tool arguments, before an agent's tool call runs, so an injected instruction cannot send a credential to an outside URL in a query string.
- Log and trace sinks, before persistence. Logs are copied widely and kept long; see audit logging for LLM systems.
- Offline, over exported transcripts, evaluation sets, prompt libraries and fine-tuning data, because these are often assembled from logs written before the inline scanners existed.
Ingress redaction replaces the secret with a typed placeholder, such as [REDACTED:aws-access-key-id], so the model can still reason about the text ('the config contains an AWS key') without seeing it. If an agent genuinely needs the credential, it should receive a handle resolved by the tool layer, never the value in context. Output scanning fits naturally in a general output-handling layer such as the one in LLM output handling.
Scanning a streaming response
Streaming breaks naive scanning. Tokens arrive a few characters at a time, and a 40-character token is split across many chunks. Scanning each chunk alone finds nothing; scanning the whole response at the end is too late, because the user has already seen it.
The fix is a hold-back buffer. Keep the last H characters unsent, where H is longer than the longest full match any rule can make, including context such as the key name in an assignment rule. That is why every rule needs a bounded length. On each new chunk, append it to a buffer of raw text, choose a cut point H characters from the end, move the cut back so it never splits a match, then redact and emit only the text before the cut. A match that is still arriving is always at the tail, and because it is shorter than H it stays in the buffer until it is complete.
class StreamingRedactor:
"""Redacts secrets in a token stream without letting a split secret slip through."""
def __init__(self, holdback: int = 256):
# holdback must exceed the longest FULL match any rule can make, keyword context included
self.holdback = holdback
self.buf = ""
def _redact(self, text: str) -> str:
# Two rules can match overlapping spans (a ghp_ token in a TOKEN= assignment): merge first.
spans = []
for name, start, end in sorted(find(text), key=lambda f: f[1]):
if spans and start < spans[-1][2]:
spans[-1][2] = max(spans[-1][2], end)
else:
spans.append([name, start, end])
# replace right to left so earlier offsets stay valid
for name, start, end in reversed(spans):
text = text[:start] + f"[REDACTED:{name}]" + text[end:]
return text
def feed(self, chunk: str) -> str:
self.buf += chunk # the buffer holds RAW text
if len(self.buf) <= self.holdback:
return ""
cut = len(self.buf) - self.holdback
ws = self.buf.rfind(" ", 0, cut) # prefer to cut at whitespace
cut = ws + 1 if ws > 0 else cut
moved = True
while moved: # never cut through a match
moved = False
for _, start, end in find(self.buf):
if start < cut < end:
cut, moved = start, True
out, self.buf = self.buf[:cut], self.buf[cut:]
return self._redact(out)
def flush(self) -> str:
out, self.buf = self._redact(self.buf), ""
return out
# usage inside a streaming endpoint
red = StreamingRedactor()
for chunk in model.stream(prompt):
if safe := red.feed(chunk):
yield safe
yield red.flush()Three details matter. Redact only what you emit: redacting the whole buffer on every chunk would redact a generic match as soon as it reached its minimum length and then emit the rest of the value. Merge overlapping matches from two rules before replacing. And the hold-back adds latency equal to H characters of generation, a few dozen tokens with H of 256, which users notice only at the end of a response, when flush sends the rest. If that is too slow, lower H to the longest full match of the rules you actually enable.
Encodings and evasion
Scanners match bytes, and the same secret can be written many ways: base64-encoded inside a Kubernetes manifest, URL-encoded in a query string, split across two lines of a JSON log, wrapped in quotes, or reformatted by a model that inserted spaces or a line break. Normalise before matching: decode base64 and URL-encoded substrings that are long enough to hold a credential, join wrapped lines in structured logs, and run the rules over both the original and the decoded form.
Be honest about the limit. A model that has been deliberately instructed to exfiltrate a secret can spell it backwards, translate it into words or hide one character per sentence, and no pattern scanner will catch that. Scanners catch accidents, which are most leaks. Deliberate exfiltration is stopped by not putting secrets in the context in the first place, by egress network controls on tools, and by canary credentials that alert when used, described in canary tokens.
What to do with a finding
| Where found | Immediate action | Follow-up |
|---|---|---|
| User prompt | Redact before the model call; tell the user a credential was removed | Suggest rotation; do not store the original |
| Retrieved document or tool result | Redact; continue the request | Open a ticket to remove it from the source system and rotate |
| Model output | Redact the stream; flag the conversation | Find how it entered context; treat as a possible injection |
| Log or dataset | Purge or rewrite the record | Rotate, then check every copy of that store |
Verification belongs off the hot path. Verifying a candidate means sending it to the issuer's API, which is slow and makes a network call per finding. Queue findings, verify asynchronously, and use the result to set priority: a verified live key is an incident, an unverified one is hygiene. Never log the raw secret in the finding. Rotation, not redaction, is the real fix for any secret that has already reached a log or a provider.
Tools and how to wire them in
You do not need to write detectors from scratch for offline work. Gitleaks in its current v8 command model scans git history, directories and standard input; the older detect and protect commands were deprecated in v8.19.0. TruffleHog ships hundreds of detectors and can verify candidates against the issuing service. Yelp's detect-secrets uses a baseline file so only new findings fail.
# Scan exported transcripts or eval datasets (gitleaks v8: git, dir and stdin commands)
cat transcripts.jsonl | gitleaks stdin -v --report-format json --report-path leaks.json
gitleaks dir -v ./prompt_library
# Scan a repository's history, including the prompts and fixtures committed to it
gitleaks git -v .
# TruffleHog: try to verify candidates against the issuing service, report live ones only
trufflehog filesystem ./exports --results=verified
# detect-secrets: record known findings once, then fail only on new ones
detect-secrets scan > .secrets.baseline
detect-secrets audit .secrets.baselineFor inline scanning in the request path, a subprocess per request is too slow. Either port the rules you need into an in-process library, as in the sketch above, or use a guardrail framework that includes a secrets scanner, as described in LLM Guard. Keep one versioned rule set for inline and offline scanning.
Tuning and operations
- Build a labelled corpus: real (rotated) secrets from past incidents, synthetic ones generated in each format, and a large sample of normal traffic. Measure recall and false positives per rule on every rule change.
- Prefer precise rules. Prefix and checksum rules can block; generic entropy rules should usually redact and flag, not block.
- Allowlist narrowly: documented sample keys such as AKIAIOSFODNN7EXAMPLE, test fixtures by path and specific hashes, never whole rules.
- Budget latency: keyword prefilters keep a few hundred rules within a millisecond or two per kilobyte; measure on your own traffic.
- Alert on rates: a spike of egress findings from one tenant suggests an injection campaign.
What to do next
- Map every path where text enters or leaves your models, including tool arguments, logs, traces and datasets.
- Add ingress redaction with typed placeholders on prompts, retrieval and tool results.
- Put a streaming egress redactor with a hold-back buffer longer than your longest rule in front of every response.
- Scan log and trace pipelines before persistence, and run gitleaks or TruffleHog over existing exports now.
- Queue findings for asynchronous verification and treat verified live credentials as incidents with rotation.
- Build a labelled test corpus and measure recall and false positives per rule before every rule change.
- If you issue tokens, give them a unique prefix and a checksum so they are cheap to detect.