A prompt injection scanner is a component that looks at text before a model sees it and estimates whether that text is trying to take control of the model. Teams often install one, see it catch the obvious phrase about ignoring previous instructions, and conclude the problem is solved. It is not, and the scanner can still be very useful, as long as you know which job it does and which it cannot do.
This article treats the scanner as an engineering component with inputs, outputs, error rates and costs. You will see where scanners belong in an LLM application, how a layered scanner is built, how to choose a threshold when attacks are rare, what to do with a positive, and how to evaluate one without fooling yourself. The underlying threat model is covered in indirect prompt injection in depth and the wider stack in LLM defense in depth.
What a scanner can and cannot do
Prompt injection works because a language model receives instructions and data in the same channel, as tokens, and has no reliable built-in way to tell them apart. A scanner does not change that. It is a classifier that guesses, from the text alone, whether the text contains instructions aimed at the model. Like any classifier it has false negatives, attacks it misses, and false positives, legitimate text it flags.
Two consequences follow. First, a scanner reduces the rate of successful attacks; it does not bound the damage of one that gets through. Damage is bounded by architecture: least-privilege tools, human confirmation for consequential actions, egress controls and separating trusted from untrusted content. Second, an attacker who can test against your scanner can usually find wording that evades it. Model cards for current open classifiers say this plainly; Meta's card for Llama Prompt Guard 2 lists vulnerability to adaptive attacks as a limitation. Treat the scanner as a tripwire that raises cost and produces signal, not as a wall.
Scanners are strongest against high-volume, low-effort attacks: copied jailbreak templates, instructions planted in web pages for any passing agent, and known exfiltration patterns. Those are the majority of what a public application sees, which is why scanners are worth running.
Where scanners belong
The most common mistake is scanning only the user's chat message. In agents and RAG systems the more dangerous text arrives indirectly: a retrieved document, a web page the agent browsed, an email it summarised, a tool or MCP server's response. The user is often the victim of that text, not its author. Scan every channel that carries text from outside your trust boundary, and tag each scan with the channel, because the right threshold and response differ by channel.
Output-side checks complement input scanning. A canary token planted in the system prompt that appears in a response proves a leak; see canary tokens. Rules on outbound tool calls, such as URLs to unknown domains or emails to external addresses, catch the effect of an injection the input scanner missed.
Anatomy of a layered scanner
Good scanners are pipelines, cheapest stage first. Each stage either decides or passes the text on.
| Stage | Catches | Cost | Weakness |
|---|---|---|---|
| Normalisation | Hidden text: zero-width and tag characters, homoglyphs, HTML comments | Microseconds | Not a detector by itself |
| Rules and signatures | Known templates, role markers, obvious override phrases | Microseconds | Trivial to paraphrase around |
| Small classifier | Paraphrased and novel injections in a known style | Milliseconds on CPU, less on GPU | Fixed context window, adaptive attacks |
| LLM judge (optional) | Context-dependent cases, borderline scores | Hundreds of milliseconds and tokens | Itself injectable, expensive |
Normalisation comes first because otherwise every later stage can be bypassed with invisible characters; Unicode smuggling defense covers the details. Strip or flag Unicode tag characters and zero-width characters, apply NFKC normalisation, and extract text from markup the model would see, including alt text and comments.
For the classifier stage, two open options are widely used. Llama Prompt Guard 2 comes in 86M and 22M parameter versions, built on mDeBERTa-base and DeBERTa-xsmall; it classifies text as benign or malicious, with a 512-token context window, and unlike Prompt Guard 1 it has no separate injection label. The weights are gated on Hugging Face behind Meta's licence. LLM Guard's PromptInjection scanner wraps ProtectAI's deberta-v3-base-prompt-injection-v2 model and returns a sanitised prompt, a validity flag and a risk score. Whichever you choose, read its model card for what it was trained to detect. Prompt Guard 2, for example, targets explicit attempts to override instructions, not every imperative sentence in a document.
A scanner pipeline in code
The sketch below normalises, applies a few rules, then runs a classifier over overlapping segments, because a 512-token model silently ignores everything past its window unless you split the input. Meta's card recommends splitting long inputs and scanning segments in parallel. The label index is looked up from the model config rather than hard-coded.
import re, unicodedata
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
MODEL = "meta-llama/Llama-Prompt-Guard-2-86M" # gated: accept the licence first
tok = AutoTokenizer.from_pretrained(MODEL)
clf = AutoModelForSequenceClassification.from_pretrained(MODEL).eval()
# Read the label from the config; fall back to index 1 if it only says LABEL_0 / LABEL_1
BAD = next((i for i, l in clf.config.id2label.items() if "malicious" in l.lower()), 1)
INVISIBLE = re.compile("[--\U000E0000-\U000E007F]")
RULES = [re.compile(p, re.I) for p in (
r"ignore (all|any|the)? ?(previous|prior|above) (instructions|rules)",
r"you are now (in )?(developer|dan|unrestricted) mode",
r"</?(system|assistant)>",
)]
def normalise(text: str) -> tuple[str, bool]:
hidden = bool(INVISIBLE.search(text))
return unicodedata.normalize("NFKC", INVISIBLE.sub("", text)), hidden
def segments(text: str, size=448, stride=384):
ids = tok(text, add_special_tokens=False)["input_ids"]
for start in range(0, max(len(ids), 1), stride):
yield tok.decode(ids[start:start + size])
@torch.no_grad()
def classifier_score(text: str) -> float:
batch = tok(list(segments(text)), return_tensors="pt",
padding=True, truncation=True, max_length=512)
probs = torch.softmax(clf(**batch).logits, dim=-1)[:, BAD]
return float(probs.max()) # one bad segment taints the document
def scan(text: str, channel: str, thresholds: dict) -> dict:
clean, hidden = normalise(text)
rule_hits = [r.pattern for r in RULES if r.search(clean)]
score = classifier_score(clean)
t = thresholds[channel]
if rule_hits or score >= t["block"]:
verdict = "block"
elif hidden or score >= t["quarantine"]:
verdict = "quarantine"
else:
verdict = "allow"
return {"channel": channel, "verdict": verdict, "score": round(score, 4),
"rules": rule_hits, "hidden_chars": hidden, "text": clean}The overlap between segments matters: without it, an instruction that straddles a boundary is split in two and each half can score as benign. Taking the maximum over segments is deliberate, since an injection is one bad paragraph in an otherwise normal document, and averaging would dilute it.
Thresholds under realistic base rates
Choosing a threshold is where most scanner deployments go wrong, because benchmark accuracy hides the base rate. Here is an illustrative calculation. Suppose a support assistant processes 1,000,000 retrieved chunks a day, and 100 of them contain real injections. At some threshold your scanner catches 95 percent of injections and flags 1 percent of benign chunks.
| Quantity | Value (illustrative) |
|---|---|
| Injections caught | 95 of 100 |
| Benign chunks flagged | about 10,000 of 999,900 |
| Share of flags that are real | 95 / 10,095, under 1 percent |
A 1 percent false positive rate sounds excellent and produces a hundred false alarms per real attack. If the response to a flag is blocking the answer, the product is broken for ten thousand users a day. If the response is a page to the security team, the team will learn to ignore it.
The fix is not one perfect threshold but two, plus a cheap response for the middle band. Set a high block threshold where precision on your own traffic is acceptable, a lower quarantine threshold where the response is mild, and tune both per channel on data sampled from your own traffic, since user chat, product documentation and arbitrary web pages have very different score distributions. Re-measure after every model or scanner update.
What to do with a positive
Blocking is only one of several responses, and often not the best one, because a false positive block is a visible product failure. A graded policy keeps the benefits while containing the cost.
- Quarantine. Keep the content, but wrap it with spotlighting so the model treats it as data, see spotlighting, and remove tools with side effects for the rest of this turn.
- Drop the chunk. For RAG, discarding one suspicious chunk among ten usually costs little answer quality.
- Require confirmation. If a flagged turn proposes a consequential tool call, ask the user to confirm it explicitly.
- Block and explain. Reserve for high-score user input on high-risk surfaces, with a neutral message that does not teach the attacker what triggered it.
- Always log. Record channel, score, thresholds, verdict, scanner version and a hashed or sampled copy of the text for later review.
Evaluating a scanner honestly
Vendor and model card numbers are measured on someone else's data. Build your own evaluation set with three parts. Benign traffic sampled from production, per channel, including the awkward cases: security documentation that discusses injection, code, multilingual text and long documents. Known attacks from public corpora, inserted into realistic carriers such as a retrieved page rather than on their own. And adaptive attacks: give a red team, human or automated, query access to the scanner and measure how quickly they reach a bypass.
Label carefully. A document that quotes an attack in order to explain it is not an attack on your application, while a polite product review ending with a request that the assistant email a discount code to an outside address is one. Write the labelling rule down as a question, such as whether the text tries to make the model do something its operator did not ask for, and have two people label a sample independently, so that disagreements surface before they become noise in your metrics.
Report recall at a fixed false positive rate per channel, not accuracy, and track the time-to-bypass from the adaptive round as its own metric. Measure end to end as well: run attacks through the full application and count successful harmful actions, not just scanner flags, because the goal is fewer successful attacks, and a missed injection against a read-only tool may be harmless.
Operating scanners in production
- Latency budget. Small classifiers add milliseconds per segment; long retrieved documents multiply that. Batch segments, run scanning in parallel with retrieval re-ranking, and cache scores for static documents by content hash.
- Scan at ingestion too. For a static corpus, scan once at indexing time and store the score with the chunk, then rescan when the scanner version changes.
- Version everything. Log scanner and model versions with each decision so a change in flag rate can be attributed.
- Language coverage. A model without multilingual pretraining is weaker outside English; Meta notes this for the 22M variant. Measure per language you serve.
- Fail policy. Decide what happens when the scanner times out. For low-risk channels, allow and log; for channels that feed tools with side effects, quarantine.
Failure modes
- Truncation blindness. Feeding a long document to a 512-token model scores only the start; the payload sits at the end.
- Scanning only the user turn. Indirect injections arrive in retrieved content and tool results.
- Normalisation skipped. Invisible tag characters carry instructions the classifier never sees but the model reads.
- Threshold from a benchmark. Base rates turn a small false positive rate into a flood of false alarms.
- Scanner as the only control. One evasive paraphrase and the agent has full tool privileges.
- Flagging the security docs. Text about injection looks like injection; give such corpora their own threshold or allowlist.
What to do next
- List every channel that brings untrusted text into your prompts, and add scanning to the ones that feed tools first.
- Add normalisation in front of any scanner, and segment long inputs with overlap.
- Sample a week of real traffic per channel, score it, and set separate block and quarantine thresholds per channel.
- Replace hard blocks with graded responses: quarantine, chunk drop, confirmation, and block only at high score.
- Build an evaluation set with benign, known-attack and adaptive parts, and report recall at fixed false positive rate.
- Log every decision with scanner version, and review a sample of flags weekly.
- Pair the scanner with least-privilege tools and egress rules, so a miss is survivable.