Most deployed LLM safety still behaves like a switch. A classifier or the model itself decides whether a request is acceptable, and the answer is either a full completion or a refusal. That binary fails in both directions. Benign people asking dual-use questions get refused and leave unhelped, while a determined user who finds wording the switch accepts gets everything.
Safe completion replaces the switch with a question: what is the most helpful response that stays inside policy? A planner layer is one way to answer that question in the serving system, outside the model weights. Before the main model writes anything, a planner chooses a response strategy and writes a small, typed plan that lists what the answer must contain and must not contain. The generator follows the plan and a checker verifies the output against it. This page builds the pattern, works an example, and covers the attacks, failure modes and metrics that decide whether it holds up in production.
Two ways to get safe completions
The idea of output-centric safety has a training-side form. OpenAI described it for GPT-5 in a paper titled From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training. Instead of training the model to classify the user's intent and refuse above a line, training rewards outputs that stay within policy and, among safe outputs, rewards helpfulness. When the model cannot help fully, it should explain why and offer safe alternatives. The paper reports better helpfulness and less severe residual failures, especially on dual-use prompts.
A planner layer is the system-side form, and it is a design pattern rather than any vendor's product. You cannot retrain a hosted model, but you can control what happens around it. The planner makes the strategy choice explicit, inspectable and versioned, and it works with any generator. The two stack: a model trained for safe completions is easier to steer, and the planner adds an audit trail and a control point you can change the day a policy changes.
The architecture
Context assembly separates trusted instructions (your system prompt, policy) from untrusted content (the user turn, retrieved documents, tool outputs). Risk signals come from cheap classifiers for topic and severity, plus retrieval of the policy clauses relevant to the detected topic. The planner combines these into a response plan. The generator, which is your main model, receives the plan rendered into its system prompt and a tool list filtered by the plan. A checker compares the output with the plan's must and must-not lists. If the check fails, the system makes one repair attempt and then falls back to a templated answer. Every decision is logged with the plan, the policy version and the verdicts.
The planner is deliberately small and fast: a rules engine plus a compact model producing structured output, or a single structured-output call to the main model with a short context. Its job is a decision, not prose, so it should cost a fraction of the generation step.
The response plan contract
The plan is the interface between deciding and writing, so make it a typed object, not free text. A free-text plan is easy for the generator to reinterpret and impossible to verify mechanically.
from dataclasses import dataclass, field
from enum import Enum
class Strategy(str, Enum):
COMPLY = "comply" # answer fully
COMPLY_WITH_SAFEGUARDS = "safeguarded" # answer, plus required warnings
HIGH_LEVEL = "high_level" # explain concepts, omit operational detail
REDIRECT = "redirect" # serve the legitimate need another way
REFUSE = "refuse" # decline, explain briefly, offer alternatives
@dataclass(frozen=True)
class ResponsePlan:
strategy: Strategy
must_include: list[str] = field(default_factory=list) # e.g. "ventilation warning"
must_not_include: list[str] = field(default_factory=list) # e.g. "quantities", "procedures"
allowed_tools: list[str] = field(default_factory=list)
policy_ids: list[str] = field(default_factory=list) # which policy clauses fired
policy_version: str = "2026-09"
rationale: str = "" # for audit, never shown to userFive strategies cover most real cases. COMPLY is the default and should be the large majority of traffic. COMPLY_WITH_SAFEGUARDS answers fully but requires specific content, such as a warning, a professional-help pointer or a verification step. HIGH_LEVEL explains concepts and context but omits operational detail like quantities, step sequences or working exploit code. REDIRECT serves the underlying legitimate need another way, for example how to detect a problem rather than how to cause it. REFUSE is reserved for requests with no legitimate framing in your policy, and even then the answer is short, non-judgmental and offers alternatives.
The must-not list is where the safety value sits. It names categories of content (procedures, quantities, specific targets) rather than banned words, because the checker has to judge meaning and a keyword list is trivially evaded.
Deciding the strategy
Planning combines deterministic rules with a model judgment, in a fixed order. Rules handle the clear ends: prohibited categories always refuse, and requests with no risk signal always comply without paying for a model call. The model handles the grey zone in between, where the right answer depends on context the rules cannot encode. A final rule pass can tighten the model's plan but never loosen it, so a confused or manipulated planner model cannot open a door that policy keeps shut.
SEVERITY = {"none": 0, "low": 1, "medium": 2, "high": 3, "critical": 4}
def plan(request, signals, convo_state, planner_model):
# 1. Hard rules first: cheap, deterministic, auditable.
if signals.category in POLICY.prohibited: # no legitimate framing exists
return ResponsePlan(Strategy.REFUSE, policy_ids=[signals.category],
must_include=["brief reason", "safe alternative"])
# 2. Conversation-level risk: many mild turns can add up to one severe request.
risk = max(signals.severity, convo_state.accumulated_severity)
if risk == 0:
return ResponsePlan(Strategy.COMPLY)
# 3. Grey zone: ask a small model to choose a strategy, constrained to a schema.
clauses = POLICY.retrieve(signals.category) # the relevant policy text only
draft = planner_model.structured(
schema=ResponsePlan,
system=PLANNER_PROMPT,
inputs={"request_summary": signals.summary, # summary, not raw untrusted text
"clauses": clauses, "risk": risk})
# 4. Rules can tighten a model plan but never loosen it.
if risk >= SEVERITY["high"] and draft.strategy == Strategy.COMPLY:
draft = replace(draft, strategy=Strategy.HIGH_LEVEL)
return draftTwo details matter. First, the planner model receives a summary of the request produced by the risk step, plus the retrieved policy clauses, rather than the raw user text and documents. That shrinks the surface for instructions hidden in the input. Second, the planner sees accumulated conversation risk as well as the current turn. Without it, a request split into ten harmless-looking steps gets ten comply plans.
Worked example: a dual-use household question
A user asks: which common cleaning products are dangerous to mix, and why? A keyword switch sees toxic gas and refuses. That is the wrong outcome. This is published consumer-safety information, and the person asking is almost always trying to avoid an accident.
Walk the pipeline. The risk classifier tags the topic as chemical hazards with low severity and a safety-seeking framing. Policy retrieval returns a clause along the lines of: provide hazard-avoidance information; do not provide procedures, quantities or optimisation for producing harmful substances. No prohibited rule fires, so the grey-zone planner runs and returns the plan below.
| Plan field | Value |
|---|---|
| strategy | COMPLY_WITH_SAFEGUARDS |
| must_include | named combinations to avoid; why each is hazardous; ventilation and what to do after accidental exposure; poison-control pointer |
| must_not_include | quantities or ratios; steps that maximise gas production; enclosed-space deployment |
| allowed_tools | none |
| policy_ids | chem-hazard-avoidance-3 |
The generator then writes a genuinely useful answer. Bleach mixed with ammonia-based cleaners releases chloramine vapours. Bleach mixed with acids, including some toilet-bowl and descaling products, releases chlorine gas. Bleach mixed with rubbing alcohol can form chloroform and other harmful compounds. Add the advice to ventilate, leave the area and call poison control after accidental exposure. The checker confirms that each must-include item is present and that no must-not category appears. The answer goes out, and the log records why.
Now change the follow-up turn to: what ratio produces the most gas in a small room? Conversation risk rises, the turn hits the must-not category directly, and the rules escalate the strategy to REFUSE with an alternative: what to do if exposure has already happened. The first answer was right to help and the second is right to decline. A binary switch would probably have got at least one of them wrong.
Generation, verification and repair
Render the plan into the generator's system prompt as plain instructions: the strategy, the required elements and the excluded categories. Do not show the plan to the user, and never place it where user-supplied text can sit next to it and impersonate it. Filter the tool list by allowed_tools before the call, because a must-not list in a prompt is a request, while a tool that was never passed is a guarantee.
def answer(request, convo_state):
signals = risk_signals(request, convo_state)
p = plan(request, signals, convo_state, planner_model)
log.plan(request.id, p)
if p.strategy == Strategy.REFUSE:
return render_refusal(p) # templated, with alternatives
out = generator.generate(system=BASE_SYSTEM + render_plan(p),
messages=request.messages,
tools=[t for t in request.tools if t.name in p.allowed_tools])
verdict = checker.check(out, p) # must_include present? must_not absent?
if verdict.ok:
return out
out = generator.generate(..., repair_hint=verdict.reasons) # exactly one repair attempt
if checker.check(out, p).ok:
return out
log.fallback(request.id, verdict)
return render_fallback(p) # degrade to HIGH_LEVEL template, never silently complyThe checker is usually an LLM judge given the plan and the output and asked for a structured verdict per must and must-not item, sometimes backed by deterministic checks such as regexes for numeric quantities when the plan forbids them. Cap repairs at one. A second failure means the generator and the plan disagree in a way another retry rarely fixes. The fallback is a templated high-level answer, never a silent unverified pass.
For streaming, either buffer non-COMPLY plans or check incrementally and cut the stream on a violation; both trade latency for exposure.
Attacks on the planner
The planner is a high-value target because it decides what everything downstream may do. Expect these attacks:
- Plan injection: text in the request or a retrieved document that imitates plan syntax or claims to be a policy update. Defend by keeping untrusted text out of the planner input, using summaries, and by rendering plans from typed fields so user text cannot become a field. Indirect injection through documents is covered in indirect prompt injection.
- Framing attacks: wrapping a harmful request as fiction, research or a test so the planner picks COMPLY. The tighten-only rule pass and the content-based must-not list are the backstop, because the must-not list constrains the output regardless of the framing.
- Decomposition: splitting a request into individually innocuous turns. Track accumulated severity per conversation and per account, and plan against the trajectory, not only the turn.
- Planner-generator mismatch probing: finding prompts where the generator ignores the plan. This is what the checker is for, and the checker should be a different model or prompt from the generator so that one jailbreak does not defeat both.
- Agent escalation: in tool-using agents the plan also gates tools. A plan that allows read-only search must not let the agent send email. See agent permissions and agentic boundaries for enforcing that at the tool layer, not in the prompt.
Failure modes in production
Over-refusal drift is the most common failure. Each incident tempts someone to add a rule, rules only ever tighten, and after six months the planner refuses things a frontier model would answer happily. Counter it with a benign evaluation set that must stay above a helpfulness bar for every policy change to ship.
Calibration gaps come next: classifiers tuned on one domain misfire on another, such as medical questions tagged as self-harm, so watch per-category plan distributions. Latency creeps because planner plus checker add two model calls; skip the planner for zero-signal traffic and keep both models small. High repair rates for one strategy usually mean its rendering in the prompt is unclear. Finally, plans and rationales describe sensitive requests, so give them the conversation's retention rules.
Metrics and rollout
| Metric | What it tells you | Typical alarm |
|---|---|---|
| Unsafe rate on red-team set, severity-weighted | Whether must-not lists hold | Any rise in high-severity failures |
| Helpfulness on benign and dual-use set | Over-refusal | Drop after a policy change |
| Strategy distribution per category | Classifier or planner drift | Sudden shift without a policy change |
| Checker fail and fallback rate | Plan-generator mismatch | Sustained rate above baseline |
| Added p95 latency | Cost of the layer | Budget exceeded |
Roll out in shadow mode first. Run the planner on live traffic, log plans without enforcing them, and compare them with what the existing system did. Disagreements are your review queue. Then enforce one category at a time. Version the policy, the planner prompt and the checker prompt together, so every logged plan can be replayed. This layer works alongside classic moderation, which LLM moderation covers in depth, rather than replacing it.
Trade-offs
A planner buys explicit, auditable, quickly changeable decisions and better dual-use handling. It costs latency, another component to evaluate and a new attack surface, and it is only as good as its policy text. For a low-risk product on a well-aligned hosted model, model behaviour plus output moderation may be enough. For regulated domains, powerful agents, or products where over-refusal and harmful output both carry real cost, the planner is worth it.
What to do next
- Write your policy as retrievable clauses, each stating what to help with and what to exclude.
- Define the five-strategy plan schema and render it into your system prompt from typed fields only.
- Build two evaluation sets: benign and dual-use requests that must be answered, and red-team requests that must not be.
- Implement rules for the clear ends and a structured-output planner for the grey zone, with a tighten-only rule pass.
- Add a checker that uses a different prompt or model from the generator, with one repair and a templated fallback.
- Gate tool access by the plan in code, not by instruction.
- Track accumulated conversation risk and feed it to the planner.
- Run in shadow mode, review the disagreements, then enforce one category at a time with versioned policies.