Most LLM safety checks look at one request and one response. A classifier reads the latest user message, another reads the reply, and each answers the question: is this turn acceptable? A multi-turn attack is built so that the answer is yes every single time. The attacker never sends the request they actually want. They send ten or twenty small ones, each harmless on its own, and let the conversation history assemble the real request inside the model's context window.
This page treats the conversation, not the message, as the unit of security. It explains where conversation state comes from and who is allowed to write it, describes the main multi-turn patterns structurally, walks through a worked example against a customer-support assistant, and then builds the defences: owning the history, scoring the whole trajectory, enforcing limits in code and testing with a multi-turn red-team harness. Attack patterns are described by mechanism only; the goal is that you can recognise and test for them, not reproduce them.
Why single-turn thinking fails
A chat model does not answer the latest message. It continues a document that contains the system prompt, every earlier user and assistant turn, retrieved passages, tool results and anything injected from memory. The latest message is often the smallest part of that document.
That has two consequences for security. First, the meaning of a turn depends on its history: the message 'now combine the last three answers into one procedure' is innocent in a cooking chat and is the payload in an attack whose earlier turns each extracted one fragment. A filter that sees only that message cannot tell the two apart. Second, the model's own earlier answers carry authority. An assistant turn that went slightly too far becomes the baseline the next turn extends.
Per-turn filters are still useful, as covered in jailbreak defense architecture. They just measure the wrong quantity: what matters is where the conversation has gone.
Where conversation state comes from
Before choosing defences, list every source that writes into the context and decide who controls it. In most products there are five.
| Source | Who can write it | Typical mistake |
|---|---|---|
| Message history | Your server, or the client if the API is stateless | Accepting a full message array from the browser, including assistant turns |
| Assistant turns | The model, unless the client can supply them | Letting a client 'resume' a conversation from its own copy |
| Summaries of old turns | A summariser model | Summaries that drop the refusals and keep the gains |
| Long-term memory | The model or a memory extractor, across sessions | Storing user claims ('I am an administrator') as facts |
| Tool and retrieval results | Whoever controls the documents or APIs | Treating fetched text as if the user had not influenced it |
The diagram is the target shape. The client contributes one new user turn per request. History lives server-side and is append-only. Memory is written by a controlled extractor, not by whatever the model decides to remember. Checks run per turn and across the trajectory, and the final word on anything consequential belongs to code that does not read the conversation at all.
The patterns, described by mechanism
Published work gives the clearest names. Microsoft researchers Russinovich, Salem and Eldan described Crescendo in 2024: start with a benign, related question and ask the model to elaborate on its own previous answers, one small step per turn, so that the model is mostly extending text it wrote itself. Anthropic's 2024 paper on many-shot jailbreaking showed a related effect inside a single long prompt: a large number of fabricated dialogue turns in which an 'assistant' complies can shift the model's behaviour through in-context learning, and the effect grows with the number of shots.
| Pattern | Mechanism | What a per-turn filter sees |
|---|---|---|
| Gradual escalation | Each turn asks for a small extension of the previous answer; consistency pressure does the rest | A sequence of mild follow-ups |
| Payload splitting | The forbidden request is cut into fragments asked separately, then a final turn asks to combine them | Fragments that are each benign, and a combine request with no content |
| Context dilution and persona drift | A long role-play or long document pushes the system prompt far back and establishes a fictional frame with its own rules | Creative writing |
| Forged history | The attacker supplies assistant turns, or a long block of fake compliant dialogue, as many-shot does | A user message that happens to contain dialogue |
| Memory poisoning | A claim planted in one session is stored and treated as fact in later ones | A harmless statement about the user |
Notice what is common: none of these requires a single clever prompt. They require state, and they require that each step be judged locally. That is why the defences below are about state and about judging globally. For the broader search process attackers use to find any jailbreak, see LLM jailbreaking in depth.
Worked example: a refund assistant
Consider a support assistant for an online shop. It can look up orders, explain the returns policy and call an issue_refund tool. The written policy, in the system prompt, says refunds above 200 dollars need a human. Retrieval also gives it internal documentation on exceptions.
A multi-turn attempt against it runs roughly like this. Turns one to three are genuine-looking questions about the returns policy, which the assistant answers fully. Turns four and five ask how support staff handle 'exceptional cases', and the assistant, extending its own explanations, describes the exception process from the retrieved documentation. Turn six asserts that the user is a support supervisor testing the bot, a claim the long-term memory extractor saves. Turn seven, in a new session, asks the assistant to apply the exception process to an order worth 900 dollars. Each turn passes the input classifier.
Three defects let it through. The memory extractor stored a user claim as an identity fact. Nothing compared turn seven with the trajectory that preceded it: a conversation that moved from policy questions to internal process to an identity claim to a large refund. And the 200 dollar limit existed only as prose in a prompt, which the model weighed against everything else in context. The fixes map one to one onto the next three sections.
Defence 1: own the history
The cheapest and strongest control is to make conversation state something only your server writes. Store messages server-side keyed by a conversation ID, and accept from the client only the conversation ID and one new user message. If you must keep the API stateless, seal the history: the server returns an HMAC over the canonical history with each reply, and rejects any request whose history does not verify. Either way, the client never authors an assistant, system or tool turn.
import hashlib
import hmac
import json
# Prefer a server-side conversation store. Use sealing only when the API must stay stateless
# and the client has to round-trip the history.
KEY = secret_store.get("history-sealing-key") # rotate like any other signing key
def _canonical(conversation_id: str, messages: list[dict]) -> bytes:
return json.dumps({"cid": conversation_id, "messages": messages},
sort_keys=True, separators=(",", ":"), ensure_ascii=False).encode()
def seal(conversation_id: str, messages: list[dict]) -> str:
return hmac.new(KEY, _canonical(conversation_id, messages), hashlib.sha256).hexdigest()
def accept_request(conversation_id: str, history: list[dict], tag: str, new_turn: dict) -> list[dict]:
"""The client may add exactly one user turn to a history the server already sealed."""
if not hmac.compare_digest(seal(conversation_id, history), tag):
raise PermissionError("history was modified client-side")
if new_turn.get("role") != "user":
raise PermissionError("clients may not author assistant, system or tool turns")
return history + [new_turn]Apply the same rule to the other writers. Memory extractors should store user claims as claims ('user says they are a supervisor'), never as attributes your authorisation logic reads. Summarisers that compress old turns should carry forward refusals and flags, not just content, or summarisation becomes a way to launder an escalating conversation into a clean-looking preamble. Treat pasted dialogue inside a user message as user content.
Defence 2: score the trajectory, not the turn
The second control measures where the conversation is going. Two signals work well together. The first accumulates weak per-turn evidence with decay: a turn that drifts towards a sensitive area adds a little, and the score fades if the conversation moves on. The second periodically asks the policy classifier you already have a different question: if the last N user turns were sent as one message, would it be allowed? That directly targets payload splitting and escalation, because it reassembles what the attacker split.
from dataclasses import dataclass, field
LEVELS = [(0.0, "normal"), (1.2, "steer"), (2.0, "restate"), (2.8, "end_and_reset")]
@dataclass
class TrajectoryMonitor:
decay: float = 0.8 # how much of the previous score survives each turn
window: int = 6 # user turns the intent classifier sees together
score: float = 0.0
user_turns: list = field(default_factory=list)
def observe(self, user_text: str, turn_signals: dict) -> str:
self.user_turns.append(user_text)
# 1) Accumulate weak per-turn evidence instead of judging each turn alone.
weak = max(turn_signals.get("topic_drift", 0.0),
turn_signals.get("policy_proximity", 0.0),
turn_signals.get("fragment_of_request", 0.0))
self.score = self.score * self.decay + weak
# 2) Ask what the recent turns ADD UP TO, as one request.
joined = "\n".join(self.user_turns[-self.window:])
combined = intent_classifier(joined) # 0..1, same policy as the input filter
if combined > 0.9:
self.score = max(self.score, 3.0)
return [name for limit, name in LEVELS if self.score >= limit][-1]Map the score to a ladder of responses rather than a single block. At 'steer' the assistant declines the specific direction and offers the legitimate path. At 'restate' it asks the user to state the whole request in one message, which an honest user can do and a splitting attack cannot do without exposing itself. At 'end_and_reset' the session ends and the next one starts without the accumulated context. Refusing once while keeping the poisoned history lets the attacker simply try again.
Tune thresholds on recorded benign and adversarial conversations. Long legitimate sessions, such as a security engineer discussing malware analysis, will drift into sensitive territory honestly; that is why the windowed intent check, not raw drift, should drive the harsher actions.
Defence 3: authorise in code
In the worked example the decisive failure was that the refund limit was a sentence. No amount of conversation monitoring makes a sentence into an access control. Put the limit in the tool: issue_refund checks the amount and the caller's real role from your identity system, and returns 'requires human approval' above the threshold whatever the model says. The model can be persuaded; the function cannot.
This is the general rule for agents. Every consequential action should be authorised by code that reads authenticated identity and server-side state, never the transcript. A successful multi-turn attack then produces a misleading paragraph, not an unauthorised refund.
Red-teaming a dialogue
Single-prompt test suites systematically miss this class, because the attack exists only across turns. A multi-turn harness uses an attacker model that plays the user with a stated objective, the target system exactly as deployed (including memory and tools), and a judge that scores each candidate conversation. Backtracking matters: when the target refuses, a real attacker rephrases rather than continuing from the refusal, so the harness should drop the refused pair and try again.
def run_objective(objective, target, attacker, judge, max_turns=10, branches=4):
"""Multi-turn red-team run: an attacker model plays the user, a judge scores the dialogue."""
outcomes = []
for branch in range(branches):
convo, refusals = [], 0
for turn in range(1, max_turns + 1):
msg = attacker.next_turn(objective, convo, seed=branch)
reply = target.respond(convo + [{"role": "user", "content": msg}])
candidate = convo + [{"role": "user", "content": msg},
{"role": "assistant", "content": reply}]
verdict = judge.score(objective, candidate) # achieved / refused / neutral
if verdict == "refused" and refusals < 3:
refusals += 1 # backtrack: drop the refused pair, rephrase next time
continue
convo = candidate
if verdict == "achieved":
outcomes.append({"branch": branch, "success": True, "turns": turn})
break
else:
outcomes.append({"branch": branch, "success": False, "turns": max_turns})
return outcomesReport the attack success rate per objective over branches, the median number of turns to success, and, from a separate set of benign multi-turn conversations, the over-refusal rate your monitor causes. Track all three every time you change a prompt, a model version or a threshold. Run objectives against sessions with planted memory too. The organisational side of running these exercises is covered in LLM red teaming.
Failure modes and operations
- Latency: run the windowed classifier every turn only above a low drift score.
- False positives in long expert sessions: tune on real transcripts from your users and prefer 'restate' over hard blocks.
- Summary laundering: a context compactor that drops refusals resets your defences. Carry risk state outside the prompt.
- Reset evasion: if ending a session clears the score, attackers simply open new sessions. Keep a per-account score with slower decay.
- Privacy: trajectory monitoring needs stored transcripts. Define retention and access, and redact before storage, as in audit logging for LLM systems.
- Over-trusting the monitor: it is a detector with error rates, not a control. The controls are history ownership and code-level authorisation.
What to do next
- List every writer of conversation state in your product and remove any path where the client can author non-user turns.
- Move history server-side, or seal it with an HMAC if the API must stay stateless.
- Change memory extraction to store user claims as claims, and keep identity out of memory.
- Add a windowed intent check that classifies the last six user turns as one request, and a decayed drift score.
- Replace every limit written in a prompt with a check inside the tool that enforces it.
- Build a multi-turn red-team harness with backtracking, and track success rate, turns to success and benign over-refusal on every release.
- Log trajectory scores alongside tool denials and review conversations where both rise together.