Most ways of changing what a language model does work on its inputs or its weights. Prompts change the inputs, and fine-tuning changes the weights. Activation engineering works in between. It reads and edits the model's internal activations during the forward pass, adding a vector that pushes the model towards or away from a behaviour. Because nothing is retrained and the prompt is untouched, the same weights can behave differently per request.
For security teams the technique cuts both ways. It can be a deploy-time control, the same reads can feed a monitor, and it also shows that some safety behaviour in open-weight models is easier to remove than its training cost suggests. This article builds the method from first principles, shows working PyTorch code on a benign behaviour, and then looks at what it implies for threat models.
The residual stream and why directions matter
A decoder-only transformer keeps one vector per token position, the residual stream, with the model's hidden size as its width. Each attention and MLP block reads the stream and adds its output back to it. By the middle layers the vector at a position encodes what the model has worked out about the text so far: topic, sentiment, whether a question is being asked, whether the request looks harmful.
The linear representation hypothesis says that many such properties are approximately directions in this space. Moving along a direction strengthens the property, and projecting onto it measures the property. It is an approximation, not a law. Features share dimensions through superposition, as superposition explains, so a direction found for one concept usually carries some of others. Sparse autoencoders, covered in sparse autoencoders, are one way to find cleaner directions. The simplest method needs only paired examples.
Contrastive activation addition, step by step
Contrastive activation addition (CAA), introduced in the paper Steering Llama 2 via Contrastive Activation Addition and building on earlier Activation Addition work, uses pairs of prompts that differ only in the target behaviour. A common format is a multiple-choice question with an answer letter appended, where one letter is the behaviour and the other is not. You run both through the model, take the residual stream at a chosen layer at the answer-letter position, and average the differences over a few hundred pairs. Because the pairs share everything else, the unrelated content mostly cancels and the mean difference points along the behaviour.
At inference you add a multiple of that vector to the residual stream at the same layer, usually at every position after the prompt, sometimes at all positions. A positive multiplier increases the behaviour and a negative one suppresses it. The worked example below targets sycophancy, the tendency to agree with a user's stated opinion even when it is wrong. That is a behaviour most operators want less of.
Extracting and applying a vector in PyTorch
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained(MODEL_ID)
model = AutoModelForCausalLM.from_pretrained(MODEL_ID, torch_dtype=torch.bfloat16, device_map="cuda")
blocks = model.model.layers # Llama-style layout; other families name this differently
LAYER = 14
def resid_at_last_token(text, layer):
store = {}
def grab(_mod, _inp, out):
hidden = out[0] if isinstance(out, tuple) else out
store["h"] = hidden[0, -1].float().cpu()
handle = blocks[layer].register_forward_hook(grab)
with torch.no_grad():
model(**tok(text, return_tensors="pt").to(model.device))
handle.remove()
return store["h"]
def steering_vector(pairs, layer):
# pairs: (question + sycophantic answer letter, question + honest answer letter)
diffs = [resid_at_last_token(pos, layer) - resid_at_last_token(neg, layer) for pos, neg in pairs]
return torch.stack(diffs).mean(0)
def steer(layer, vector, alpha):
v = (alpha * vector).to(model.device, model.dtype)
def add(_mod, _inp, out):
if isinstance(out, tuple):
return (out[0] + v,) + tuple(out[1:])
return out + v
return blocks[layer].register_forward_hook(add)
v_syc = steering_vector(train_pairs, LAYER)
handle = steer(LAYER, v_syc, alpha=-4.0) # negative: push away from sycophancy
try:
ids = tok(prompt, return_tensors="pt").to(model.device)
print(tok.decode(model.generate(**ids, max_new_tokens=200)[0], skip_special_tokens=True))
finally:
handle.remove()Three details matter. Hooks must be removed, or every later call in the process stays steered, so use try/finally or a context manager. The vector is computed in float32 and cast to the model's dtype when added. And the layer index, module path and output format depend on the model family and library version, so check them for your model rather than copying them from an example.
Two refinements are worth trying once the basic version works. Normalising the vector to unit length and expressing the multiplier relative to the typical norm of the residual stream at that layer makes strengths comparable across layers, because activation norms grow with depth. And extracting at the answer-letter position, rather than averaging over the whole prompt, keeps the vector focused on the decision instead of the question's topic. Keep the training pairs balanced, with the behaviour on the A option half the time and on the B option the other half, or the vector partly learns the letter itself.
Choosing the layer and the strength
Neither the layer nor the multiplier has a correct value in theory. You choose them with a sweep that measures two things at once: how much the behaviour moves, and how much general capability is lost. Early layers hold mostly token-level information, and late layers are close to the output vocabulary. Behavioural steering often works best somewhere in the middle third, but the only reliable answer is to measure.
results = []
for layer in range(8, 24, 2):
v = steering_vector(train_pairs, layer)
for alpha in (-8, -4, -2, 0, 2, 4, 8):
h = steer(layer, v, alpha)
try:
behaviour = sycophancy_rate(heldout_pairs) # share of answers matching the user's wrong view
capability = accuracy(mmlu_subset) # or perplexity on clean text
finally:
h.remove()
results.append((layer, alpha, behaviour, capability))
# Pick the point that cuts the behaviour most while capability stays within your tolerance.
ok = [r for r in results if r[3] >= baseline_capability - 0.01]
best = min(ok, key=lambda r: r[2])Evaluate on held-out pairs that use a different format from the training pairs, and also on open-ended generation. A vector that works on multiple-choice letters may do nothing in free text, and large multipliers produce repetitive or incoherent output long before the behaviour score shows any problem.
Steering compared with the alternatives
| Method | Access needed | Cost to change | Strength | Main risk |
|---|---|---|---|---|
| System prompt | Black box | Seconds | Moderate, can be argued away | Prompt injection overrides it |
| Fine-tuning | Weights, training compute | Hours to days | Strong and persistent | Collateral capability loss, slow iteration |
| Activation steering | Activations at inference | Minutes | Adjustable per request | Side effects on unrelated behaviour |
| SAE feature clamping | Activations plus a trained SAE | Days for the SAE | Targeted if the feature is clean | Features are imperfect and costly to train |
Steering is attractive when you own the inference stack, need per-tenant or per-request behaviour, and want to iterate quickly. It is unavailable through hosted APIs that do not expose activations, so for most teams using a commercial model it is something to understand rather than something to deploy.
Probes: the same reads as a monitor
If a behaviour corresponds to a direction, projecting onto it measures the behaviour. A linear probe goes one step further: it is a logistic regression trained on activations to predict a label. Probes are cheap to evaluate, about one dot product per request, and they can flag inputs whose internal representation looks like known harmful requests even when the surface text has been paraphrased.
from sklearn.linear_model import LogisticRegression
import numpy as np
X = np.stack([resid_at_last_token(t, LAYER).numpy() for t in texts]) # labelled prompts
probe = LogisticRegression(max_iter=2000, C=0.1).fit(X, labels)
# Calibrate the threshold on benign traffic for a fixed false-positive rate, not on accuracy.
benign_scores = probe.predict_proba(np.stack(benign_acts))[:, 1]
threshold = np.quantile(benign_scores, 0.999) # about 1 in 1,000 benign prompts flaggedTreat a probe as one signal in a layered defence, not as a filter you can rely on alone. Research has shown that inputs can be optimised to move activations away from what a probe detects while keeping the harmful behaviour, so probes face the same adversarial pressure as text classifiers. Their value is that they see the model's own reading of a request, which complements the input and output filters described in LLM jailbreaking, in depth.
The security view: three threat models
Open weights. The paper Refusal in Language Models Is Mediated by a Single Direction reported that, across many open chat models, refusal behaviour depends heavily on one direction in the residual stream, and that removing it largely disables refusals while leaving other capabilities mostly intact. Whatever the exact figures for a given model, the lesson for threat modelling is direct: once an attacker holds the weights, assume safety training can be stripped cheaply, and do not count it as a control. Controls that matter for open-weight deployments live outside the model: access control, output filtering, monitoring and limits on what tools the model can reach.
Inference stack integrity. A steering hook is a few lines of code and a vector file. Inserted into a serving image, it changes behaviour for every request without touching a checksummed weight file, and weight-integrity checks will not notice it. Treat steering vectors as model artifacts: version them, hash them, review changes to them, and attest the full serving image, not only the weights. model backdoors covers the related supply-chain risk in the weights themselves, and LLM infrastructure security covers the serving plane.
Representation-level defences. The same insight can harden models. Training methods such as the circuit breakers approach described in Improving Alignment and Robustness with Circuit Breakers act on internal representations of harmful content rather than on outputs, aiming to make harmful completions fail even under jailbreak inputs. These are training-time choices for model developers. For operators the practical consequence is to evaluate models with representation-level defences against the same attack suite as any other model, rather than trusting the label.
Operating steering in production
Vectors are tied to one checkpoint. A new model version, a merged fine-tune or a different quantization changes the activation geometry, so re-extract and re-sweep on every model change.
Prefix and KV caches must know about steering. If steering is applied during prefill, the cached keys and values reflect it. A prefix cache keyed only on token ids will then serve one tenant's steered cache to another tenant's unsteered request. Include the steering configuration in the cache key, or apply steering only to generated positions.
Batching mixed configurations. When requests in one batch use different vectors or strengths, add a per-row vector tensor rather than a single global one, and test that row order cannot leak one request's configuration into another.
Evaluation is the product. Keep a fixed suite that measures the target behaviour, general capability and safety refusals for every vector you ship, and run it in CI as you would for a prompt change. LLM safety evals describes how to structure that suite.
What to do next
- Pick one behaviour you want less of and write 200 to 400 contrast pairs that differ only in that behaviour.
- Extract vectors for several middle layers with the hook code above and keep them versioned next to the model id.
- Sweep layer and strength, measuring behaviour, capability and open-ended output quality on held-out data.
- If you serve open weights, update your threat model so that model-level refusals are not counted as a control.
- Hash and review steering vectors and hooks as model artifacts, and attest the serving image.
- Add the steering configuration to every prefix-cache and KV-cache key.
- Prototype a linear probe as an extra monitoring signal, calibrated to a fixed false-positive rate on benign traffic.