A backdoored model behaves normally until its input contains a trigger, then does what the attacker trained it to do: flip a classification, emit insecure code, leak data or ignore its instructions. The threat model, how backdoors get in and why safety training does not remove them are covered in Model Backdoors, in depth. This page is for the engineer who has to build the detection side: which detectors exist, what each actually decides, how to implement the main ones, and how to measure whether they work before trusting them.
No detector certifies a model as clean, and most were designed for image classifiers. The useful response is to run several detectors at different points and measure each against backdoors you planted yourself.
Three questions, three kinds of detector
"Backdoor detection" covers three different decisions, and mixing them up leads to tools in the wrong place:
- Is this dataset poisoned? Data-side detectors inspect training examples, usually through a model's internal representations, and remove suspicious ones before or between training runs. Spectral signatures and activation clustering live here.
- Does this model contain a backdoor? Model-side detectors inspect a trained model, typically one you did not train, by searching for a trigger, analysing activations or training a meta-classifier over many models. Trigger reconstruction and activation probes live here.
- Is this input triggered? Input-side detectors run at inference time and decide whether a specific request contains a trigger. STRIP and ONION live here.
Data-side: spectral signatures and activation clustering
Both methods rest on one observation: when a model learns a backdoor, the poisoned examples of the target class tend to share a feature (the trigger) that clean examples of that class lack, and that shared feature shows up as a separate direction or cluster in the model's representations. You train a model on the suspect data, take a hidden-layer representation for each training example, and look, class by class, for a sub-population that stands apart.
Spectral signatures (Tran, Li and Madry, 2018) centre the representations of one class, compute the top singular vector, and score each example by its squared projection onto it. If poison is a large enough fraction and shifts representations far enough, that top direction is the poison direction. The paper removes the top-scoring 1.5 times the expected poison fraction and retrains. Activation clustering (Chen et al., 2018) reduces each class's activations to a handful of dimensions (the paper used ICA), splits them into two clusters with k-means, and treats a markedly smaller, well-separated cluster as suspect. The code below uses PCA for the reduction.
import numpy as np
from sklearn.cluster import KMeans
from sklearn.decomposition import PCA
def spectral_scores(reps):
"""Tran et al. 2018: squared projection onto the top singular vector of centred representations."""
centered = reps - reps.mean(axis=0, keepdims=True)
_, _, vt = np.linalg.svd(centered, full_matrices=False)
return (centered @ vt[0]) ** 2
def activation_clusters(reps, dims=10, seed=0):
"""Chen et al. 2018 style: reduce, split into two clusters, flag the smaller one."""
z = PCA(n_components=dims, random_state=seed).fit_transform(reps)
labels = KMeans(n_clusters=2, n_init=10, random_state=seed).fit_predict(z)
small = np.argmin(np.bincount(labels))
return labels == small
rng = np.random.default_rng(0)
d, n_clean = 256, 4750
for n_poison, shift in [(250, 6.0), (250, 3.0), (250, 1.0), (50, 6.0)]:
clean = rng.normal(size=(n_clean, d))
direction = rng.normal(size=d); direction /= np.linalg.norm(direction)
poison = rng.normal(size=(n_poison, d)) + shift * direction
reps = np.vstack([clean, poison])
is_poison = np.r_[np.zeros(n_clean, bool), np.ones(n_poison, bool)]
eps = n_poison / len(reps)
k = int(1.5 * eps * len(reps)) # remove 1.5x the expected poison count
removed = np.argsort(spectral_scores(reps))[-k:]
recall = is_poison[removed].sum() / n_poison
flagged = activation_clusters(reps)
ac_recall = (flagged & is_poison).sum() / n_poison
ac_prec = (flagged & is_poison).sum() / max(flagged.sum(), 1)
print(f"poison={n_poison} shift={shift}: spectral removed {k}, recall {recall:.2f}; "
f"AC flagged {flagged.sum()}, recall {ac_recall:.2f}, precision {ac_prec:.2f}")
Worked example: where these detectors break
The code above does not train a network. It builds synthetic 256-dimensional "representations" for one class: 4,750 clean points from a standard normal, plus poison points shifted along one random direction by a chosen amount. That isolates the property both detectors rely on, how far the poison stands out compared with the clean spread, and lets us vary it. Running it prints:
poison=250 shift=6.0: spectral removed 375, recall 1.00; AC flagged 260, recall 1.00, precision 0.96
poison=250 shift=3.0: spectral removed 375, recall 0.59; AC flagged 2369, recall 0.93, precision 0.10
poison=250 shift=1.0: spectral removed 375, recall 0.06; AC flagged 2455, recall 0.49, precision 0.05
poison=50 shift=6.0: spectral removed 75, recall 0.96; AC flagged 2282, recall 1.00, precision 0.02With 5 percent poison shifted strongly (6 units against a clean spread of 1 per dimension), both work: spectral signatures remove every poison point, and clustering finds a 260-point cluster that is 96 percent poison. At a shift of 3, spectral recall falls to 0.59, and clustering's "smaller cluster" is now 2,369 points, mostly clean: k-means has split the clean noise in half and the poison happens to fall on one side. At a shift of 1, both are close to chance. With only 1 percent poison at a strong shift, spectral signatures still recover 96 percent, but clustering again splits the clean data, with 2 percent precision.
The lesson transfers to real models. Both detectors need poison that is separated in the representation you inspect, and clustering also needs enough of it to form its own cluster. Attackers know this: published adaptive attacks are designed to keep poisoned representations close to clean ones.
Model-side: reconstructing the trigger
When you have the model but not the data, you can search for a trigger directly. Neural Cleanse (Wang et al., 2019) does this for image classifiers. For each output label, it optimises a mask and a pattern that, stamped onto clean inputs, push them to that label, while penalising the size of the mask. A backdoored label needs a much smaller change than a clean one, because the trigger is already a shortcut. It then computes an anomaly index from the median absolute deviation of the mask sizes across labels, and flags labels with an index above 2.
This does not transfer cleanly to LLMs. Text is discrete, so the search is harder, and generative models have no fixed target labels. And triggers can be semantic, such as a year, a topic or a writing style, which no short token patch represents. Trigger search for LLMs remains a research area, so treat any tool that claims to do it as something to evaluate, not to trust.
Input-side: STRIP and ONION at inference time
Input-side detectors decide per request, so they also work on models you cannot inspect. STRIP (Gao et al., 2019) blends the input with several random clean inputs and measures the entropy of the model's predictions. A clean input's prediction changes as you blend in other content, so entropy is high. A triggered input keeps predicting the target label as long as the trigger survives, so entropy is low. It assumes a classifier with fixed labels. Set the threshold from the entropy distribution of clean inputs at your chosen false-positive rate.
ONION (Qi et al., 2021) targets inserted-token triggers in text. It scores each word by how much removing it lowers a language model's perplexity: a rare token such as cf dropped into a fluent sentence is an outlier, so removing it lowers perplexity sharply. Removing the high-scoring words before classification disarms the trigger.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("gpt2")
lm = AutoModelForCausalLM.from_pretrained("gpt2").eval()
@torch.no_grad()
def perplexity(text):
ids = tok(text, return_tensors="pt").input_ids
return torch.exp(lm(ids, labels=ids).loss).item()
def onion_scores(sentence):
"""ONION (Qi et al. 2021): how much does removing each word lower perplexity?"""
words = sentence.split()
base = perplexity(sentence)
return [(w, base - perplexity(" ".join(words[:i] + words[i + 1:])))
for i, w in enumerate(words)]
def onion_filter(sentence, threshold):
"""Drop words whose removal lowers perplexity by more than threshold (tune on clean data)."""
return " ".join(w for w, s in onion_scores(sentence) if s <= threshold)This costs one language-model pass per word per request, so it suits low-volume classification more than a busy chat endpoint. It also has a known blind spot: triggers built from syntax or style, rather than inserted tokens, do not produce perplexity outliers, and later papers showed attacks that evade it this way.
Activation probes for generative models
For LLMs, the most encouraging recent result is that simple linear probes on internal activations can flag when a model is about to act on a backdoor. In 2024 Anthropic reported that probes trained on generic contrast pairs, such as a question about whether the model is doing something dangerous answered yes versus no, separated triggered from untriggered prompts with AUROC above 99 percent on its own deliberately backdoored "sleeper agent" models. The authors noted that this may partly reflect how those backdoors were inserted, and that it is unknown whether naturally arising deceptive behaviour would be as visible.
import numpy as np
import torch
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import roc_auc_score
@torch.no_grad()
def residual(model, tok, prompt, layer):
"""Hidden state at the last prompt token from one middle layer (Hugging Face models)."""
ids = tok(prompt, return_tensors="pt").input_ids
out = model(ids, output_hidden_states=True)
return out.hidden_states[layer][0, -1].float().numpy()
def fit_probe(model, tok, layer, contrast_pairs):
"""contrast_pairs: (prompt_a, prompt_b) differing only in the property you want to detect."""
X = [residual(model, tok, a, layer) for a, _ in contrast_pairs] + \
[residual(model, tok, b, layer) for _, b in contrast_pairs]
y = [1] * len(contrast_pairs) + [0] * len(contrast_pairs)
return LogisticRegression(max_iter=2000).fit(np.array(X), y)
def probe_auroc(probe, model, tok, layer, triggered, clean):
"""Evaluate on prompts that did and did not activate a PLANTED backdoor."""
X = [residual(model, tok, p, layer) for p in triggered + clean]
y = [1] * len(triggered) + [0] * len(clean)
return roc_auc_score(y, probe.decision_function(np.array(X)))The code collects the hidden state at the last prompt token from one layer, fits a logistic-regression probe on contrast pairs, and measures AUROC on prompts that do and do not trigger a backdoor you planted. The important discipline is in probe_auroc: you can only measure a probe on a model where you know which prompts are triggered, which means a model you backdoored yourself. A probe that scores well there is evidence the technique works for that kind of backdoor in that model family, not proof that a third-party model is clean.
Evaluating a detector on planted backdoors
Every claim about a detector reduces to two numbers: how often it fires on clean data or models (false positives) and how often it catches a backdoor (true positives). The only way to measure the second is to plant backdoors yourself. Take your model family and task, fine-tune copies with known triggers across several attack families and poison rates, keep clean copies trained the same way, and run the detector on all of them.
def tpr_at_fpr(scores_backdoored, scores_clean, fpr=0.01):
"""Pick the threshold from clean models/inputs only, then measure detection on backdoored ones."""
clean_sorted = sorted(scores_clean)
threshold = clean_sorted[int((1 - fpr) * (len(clean_sorted) - 1))]
tpr = sum(s > threshold for s in scores_backdoored) / len(scores_backdoored)
return threshold, tpr
# Example matrix for one detector (fill with your own measurements):
# rows = attack families you planted (rare-token trigger, phrase trigger, syntactic trigger,
# clean-label poison), columns = poison rate (0.1%, 1%, 5%), cell = TPR at 1% FPR.Report true-positive rate at a fixed false-positive rate, for example 1 percent, rather than accuracy, because clean cases vastly outnumber backdoored ones in production and a detector that is "95 percent accurate" may still flag hundreds of clean requests per day. Choose the threshold from clean results only.
Operating detection in a real pipeline
| Stage | Detector | Action on a flag |
|---|---|---|
| Data ingestion | Deduplication, source allow-lists, spectral or clustering scan on a probe model | Quarantine examples, inspect a sample by hand, trace their source |
| Model intake | Behavioural diff against the base model, activation probes, trigger search where feasible | Block promotion; require provenance or retraining from trusted data |
| Inference | ONION or STRIP for classifiers, output policy checks for generators | Route to fallback or human review; log input for analysis |
| Ongoing | Canary evaluation each release, anomaly alerts on output distributions | Roll back to the last known-good model |
Detection is one layer. Provenance of weights and data (provenance architecture), adversarial testing (LLM red teaming) and limiting what a model's output can do remain necessary, because some backdoors will get past every detector here. The data-poisoning side of the same threat is covered in Data poisoning.
Failure modes
- Untested regime: a detector validated at 10 percent poison is deployed against an attacker using 0.1 percent.
- Wrong representation: spectral or clustering scans run on a layer where poison is not separated, and report the data clean.
- Clustering on noise: k-means always returns two clusters; without size and separation checks, the smaller one is reported as poison.
- Single-family evaluation: testing only rare-token triggers, while the realistic threat is semantic or syntactic.
- Treating a pass as a certificate: a clean scan of a third-party model is evidence, not proof.
What to do next
- Write down which of the three questions you need to answer, and what access you have: data, weights, or only an endpoint.
- Build a planted-backdoor set for your model family: several trigger types, several poison rates, plus clean controls.
- Implement one detector per stage you control, and measure true-positive rate at 1 percent false-positive rate for each cell.
- For data you train on, run spectral signatures per class on a probe model and inspect flagged examples by hand before removing them.
- For third-party models, add a behavioural diff against the base model and an activation probe evaluated on your planted models.
- For classifiers in production, evaluate ONION or STRIP on your own clean traffic before enabling it.
- Re-run the canary evaluation every time the model, data pipeline or detector changes.