LLM Guard is an open-source Python toolkit from Protect AI that wraps a call to a language model with two chains of scanners: one that inspects and rewrites the prompt before it reaches the model, and one that inspects and rewrites the response before it reaches the user. Each scanner is a small, focused check, some built from regular expressions and some from Hugging Face classifiers, and the library combines them into a pass or fail decision with a risk score per scanner.

Before anything else, one fact frames this whole article: the project is archived. On 2026-10-02 the GitHub repository reported archived: true with its last push on 2026-07-08, and its README states that the project and its Hugging Face models are no longer under active development or maintained. The last PyPI release is 0.3.16, uploaded in May 2025. The code still installs and runs, and many teams still depend on it, so this article explains how it works, what its source actually does in the places that matter for security, and how to run it, contain it and eventually replace it. For detector thresholds and evaluation in general, read prompt injection scanners alongside this page.

Advertisement

The architecture in one picture

The core API is two functions. scan_prompt(scanners, prompt, fail_fast=False) runs a list of input scanners sequentially, in the order given; each scanner receives the text as the previous scanner left it, so a redacting scanner early in the list changes what later scanners see. It returns three values: the sanitized prompt, a dictionary mapping scanner class name to a boolean validity flag, and a dictionary mapping scanner name to a risk score from 0 (no risk) to 1 (high risk). scan_output(scanners, prompt, output, fail_fast=False) does the same for the response and also passes the prompt, because some output checks, such as relevance, compare the two.

With fail_fast=True the chain stops at the first invalid result, which saves latency when you will block anyway, at the cost of incomplete scores for logging.

LLM Guard request path: two scanner chains around the model, one Vault linking themYour appbuilds promptscan_prompt(scanners, prompt)Anonymize, Secrets, InvisibleText, TokenLimit,PromptInjection ... each sees the previous outputLLMsanitized promptPolicyblock, flag, allowvalidis_valid falsescan_output(scanners, prompt, out)Sensitive, NoRefusal, Relevance, MaliciousURLs,Deanonymize (restores placeholders last)raw responseUseris_valid falsePolicyredact or refuseVault[REDACTED_PERSON_1] to real valuestorelookupScores per scanner are 0 (no risk) to 1 (high risk); the caller, not the library, decides what a failure means.
LLM Guard wraps the model with an input chain and an output chain. Anonymize writes placeholder mappings into the Vault and Deanonymize reads them back; the decision about what to do with an invalid result belongs to your code.

The scanner families and what each costs

The useful way to think about the catalogue is by mechanism, because mechanism determines latency, failure behaviour and how an attacker defeats it. Names below are the scanner classes from the source.

FamilyExamplesCostWeakness
Pattern and ruleBanSubstrings, Regex, Secrets, InvisibleText, TokenLimit, JSON, ReadingTimeMicroseconds to milliseconds, CPUTrivial to evade by paraphrase; precise when the target is literal
ClassifierPromptInjection, Toxicity, BanTopics, BanCode, Code, Gibberish, Language, Sentiment, Bias, NoRefusalOne transformer forward pass per windowFixed context window, distribution shift, adaptive attacks
TransformAnonymize, Deanonymize, SensitiveNamed-entity recognition (Presidio) plus rewritingMissed entities leak; restored entities need a trusted mapping
Comparison and networkRelevance, FactualConsistency, LanguageSame, MaliciousURLs, URLReachabilityModel pass or outbound HTTPURLReachability makes network calls from your guard tier

Two consequences follow. First, order the chain cheapest first: invisible-character stripping and secret detection before any classifier, so that a cheap block avoids an expensive model call. Second, every classifier has a bounded input length, so long prompts need a policy, which is what match types are for.

Advertisement

A worked integration

Here is a minimal, production-shaped wrapper. It builds the stateless, model-backed scanners once per process, creates a fresh Vault for every request, sets the PromptInjection threshold explicitly and logs scores rather than raw text.

from llm_guard import scan_prompt, scan_output
from llm_guard.input_scanners import Anonymize, InvisibleText, PromptInjection, Secrets, TokenLimit
from llm_guard.input_scanners.prompt_injection import MatchType
from llm_guard.output_scanners import Deanonymize, NoRefusal, Sensitive
from llm_guard.vault import Vault

# Model-backed scanners are expensive to build: create them once per process.
STATELESS_IN = [
    InvisibleText(),
    Secrets(redact_mode="all"),
    TokenLimit(limit=4096),
    PromptInjection(threshold=0.92, match_type=MatchType.SENTENCE),
]
STATELESS_OUT = [Sensitive(), NoRefusal()]

def guarded_chat(llm, user_text: str) -> str:
    vault = Vault()                       # one vault per request, never shared
    pii_in = Anonymize(vault)             # in production: cache the analyzer, not the vault
    prompt, valid, scores = scan_prompt([pii_in] + STATELESS_IN, user_text, fail_fast=True)
    if not all(valid.values()):
        log_decision("prompt", scores)    # scores only, never the raw text
        return "Sorry, I can't help with that request."
    raw = llm.complete(prompt)
    out, valid, scores = scan_output(STATELESS_OUT + [Deanonymize(vault)], prompt, raw)
    if not all(valid.values()):
        log_decision("output", scores)
        return "The answer was withheld by a safety check."
    return out

Three details are deliberate. The threshold is set explicitly because the source default (0.92) differs from the docstring (0.9) and from the documentation's example (0.5); never rely on a default you have not read. Deanonymize is last in the output chain so that detectors such as Sensitive see placeholders rather than restored personal data. And the refusal messages are generic, so a blocked user learns nothing about which scanner fired.

One caveat on the sketch: Anonymize binds its vault at construction and loads a recognizer, so constructing it per request is slow. In a real service you either pool scanner instances keyed by request, or subclass to pass the vault per call. What you must not do is the convenient thing, which is the next section's subject.

Anonymize, Deanonymize and the Vault

Anonymize uses Microsoft Presidio's analyzer to find entities such as names, email addresses, phone numbers, card numbers and IBANs, and replaces each with a placeholder of the form [REDACTED_PERSON_1]. It appends a tuple of placeholder and original value to the Vault. Deanonymize later walks every tuple in the vault and substitutes the original value wherever the placeholder appears in the model output, using an exact, case-insensitive or fuzzy matching strategy.

Reading the source makes the security property clear: the Vault is a plain in-memory list of tuples with no scoping, expiry or size limit. It is correct only if one vault serves one conversation. If two users share a vault, user B's response containing [REDACTED_PERSON_1], which a model can easily emit, is rewritten with whatever value user A's request stored under that placeholder. And because nothing is ever removed, the list grows for the life of the process.

Other practical points: the original PII stays in process memory, so a heap dump or debug log can contain it; Presidio recognises entity types it was configured for and misses others, especially in languages other than English; and fuzzy de-anonymization can restore a value into a sentence where the model used a similar-looking token for something else. Treat anonymization as data minimization toward the model provider, not as a guarantee, and pair it with the controls in PII protection for LLM systems.

PromptInjection: model, windows and match types

The PromptInjection scanner wraps ProtectAI/deberta-v3-base-prompt-injection-v2, a DeBERTa-v3 classifier with a binary label. Like any encoder, it reads a bounded number of tokens, so the match_type parameter decides how text longer than the window is handled:

  • FULL: classify the text as one input. Content beyond the window is not seen, so an attacker can pad the front and put the payload at the end.
  • SENTENCE: split into sentences and classify each; the risk is the highest sentence score. Better coverage, more forward passes.
  • CHUNKS: fixed-size chunks, useful for long retrieved documents.
  • TRUNCATE_HEAD_TAIL and TRUNCATE_TOKEN_HEAD_TAIL: classify the beginning and end only, by characters or by tokens. Cheap, and blind to the middle.

Two limits apply whatever you choose. The model was trained mostly on English and on attack styles known when it was built, and with the project archived it will not be retrained. And text you never pass to the scanner is never checked: documents fetched by retrieval and tool results need the same scan as user input, as defense in depth for LLM applications explains.

Running the API server, and its defaults

The repository also ships llm_guard_api, a FastAPI and Uvicorn service configured by config/scanners.yml. It exposes /analyze/prompt and /analyze/output, which run the chain sequentially and return sanitized text, /scan/prompt and /scan/output, which run scanners in parallel and return verdicts without sanitizing, plus /healthz, /readyz and /metrics. A hardened configuration looks like this:

app:
  scan_fail_fast: ${SCAN_FAIL_FAST:true}      # shipped default is false
  scan_prompt_timeout: ${SCAN_PROMPT_TIMEOUT:10}
  lazy_load: ${LAZY_LOAD:false}               # load models before /readyz goes green
auth:
  type: http_bearer
  token: ${AUTH_TOKEN:}                       # empty unless set: fail the deploy if unset
rate_limit:
  enabled: ${RATE_LIMIT_ENABLED:true}
input_scanners:
  - type: PromptInjection
    params:
      threshold: 0.92
      match_type: "sentence"

The source shows five behaviours you must handle before exposing it, even internally:

  1. One Vault for the whole process. The app creates a single Vault() at start-up and passes it to every scanner, so the cross-request de-anonymization and unbounded growth described above apply to the server as shipped. Do not use Anonymize and Deanonymize through the shared API for multi-user traffic; do anonymization in-process with per-request vaults.
  2. Callers can switch scanners off. Requests accept a scanners_suppress list that removes named scanners. Only your gateway should ever reach the service; never let end-user requests pass through to it.
  3. Authentication is opt-in. The bearer token defaults to empty and rate limiting to disabled.
  4. Debug logs contain raw text. At LOG_LEVEL=DEBUG the handlers log the full prompt and output. Keep production at INFO and treat log pipelines as sensitive.
  5. Timeouts return HTTP 408. Whether a timed-out scan blocks or passes the request is your gateway's decision. Make it explicitly.

Latency and capacity

Cost is roughly the sum of each scanner's latency in sequential mode. Pattern scanners are negligible; each classifier is a transformer forward pass per window, so SENTENCE mode on a long prompt multiplies the work. The project's own benchmark reports an average of 7.65 ms for PromptInjection on an AWS g5.xlarge GPU with ONNX Runtime; treat it as a best case for short inputs and measure your own traffic. ONNX support comes from the onnxruntime and onnxruntime-gpu extras and a use_onnx=True constructor flag.

Model loading dominates start-up: with lazy_load enabled, the first request pays it and may hit the timeout. Disable lazy loading, warm the models, and only then report ready. Size worker memory for every model in the chain, since each classifier is a separate copy of its weights.

Living with an archived dependency

An archived security library is not instantly unsafe, but its risk now grows over time. The concrete issues on PyPI metadata are a Python ceiling (requires_python >=3.10,<3.13, so it will not install on 3.13) and exact pins such as transformers==4.51.3 and presidio-analyzer==2.2.358, which conflict with newer ML stacks and will not receive security patches through this package.

  • Isolate it. Run it as its own service or container with its own lockfile, so its pins do not hold back the rest of your stack.
  • Mirror what you depend on. Vendor the wheel and copy the Hugging Face model files to your own storage with recorded hashes; an upstream deletion should not take down your guard tier.
  • Wrap it behind your own interface, for example check_prompt(text) -> Verdict, so that replacing it later is one adapter, not a rewrite.
  • Scan transitive dependencies for vulnerabilities and decide in advance who forks and patches if a serious one appears.
  • Keep your evaluation set, so any replacement can be measured against the same prompts and the same false-positive budget.

Failure modes

  • A shared Vault restores one user's PII into another user's response.
  • Long prompts in FULL mode hide payloads past the classifier window.
  • Retrieved documents and tool outputs bypass the input chain because only the user message is scanned.
  • Scanner timeouts fail open by accident because the 408 is treated as a transient error and retried without scanning.
  • Thresholds left at defaults that nobody reviewed, or copied from a documentation example.
  • Output scanners run on streamed text only after the whole response was already shown to the user.
  • The guard becomes the only control, so a missed detection becomes a tool call with real side effects. Scanners reduce volume; permissions and safe output handling contain impact.

What to do next

  1. Pin the exact LLM Guard version, Python version and model files you run today, and mirror the wheel and weights.
  2. Grep your code and config for any shared Vault and replace it with one vault per conversation.
  3. If you run the API server, set a bearer token, enable rate limiting, block scanners_suppress at your gateway and keep logs at INFO.
  4. Set every threshold and match type explicitly, and choose a policy for scanner timeouts.
  5. Scan retrieved documents and tool outputs, not only user messages.
  6. Measure per-scanner latency and false positives on a week of real traffic before enforcing.
  7. Put the library behind your own interface and schedule an evaluation of maintained replacements against your saved test set.
Key takeaway: LLM Guard wraps a model call with sequential input and output scanner chains that return sanitized text, validity flags and per-scanner risk scores. It is useful and still runs, but it is archived, pinned to older dependencies and capped below Python 3.13. Use one Vault per conversation, set thresholds and match types explicitly, harden the API server's defaults, scan retrieved content too, isolate and mirror the dependency, and wrap it behind your own interface so you can replace it.