Why raw reward-model scores can't drive the update
A reward model (RM) is trained on preferences — it learns to rank y_win above y_lose via a Bradley–Terry loss — not to emit a calibrated magnitude. Add any constant to every score and the ranking loss is unchanged, so the RM’s zero point is arbitrary, and its overall scale is whatever the optimization happened to settle on. One RM might output rewards clustered around +7.0 with a spread of ±0.3; another around −12 with a spread of ±4.
Feeding those numbers straight into a policy-gradient update is a problem on two fronts. A large constant offset injects a huge, meaningless baseline into every advantage estimate, inflating gradient variance. And a scale that doesn’t match your KL penalty means one term silently dominates the other. Worse, the distribution is non-stationary: as the policy shifts during training, the rewards it earns drift too, so any fixed rescaling you hard-code goes stale. The signal must be normalized online, against the running reality of what the current policy is producing.
Where reward enters the PPO update
RLHF with PPO maximizes expected reward while staying close to the supervised reference policy π_ref. The per-token reward actually optimized is not the RM score alone — it is the RM score with a KL penalty folded in:
r_t = r_RM(x, y) · [t = T] − β · ( log π_θ(y_t | .) − log π_ref(y_t | .) )The RM contributes a single scalar at the final token T of the completion; the KL term is subtracted at every token as a per-step penalty for drifting from the reference. Those per-token rewards feed a value function and a GAE advantage estimate A_t, which is what the clipped PPO surrogate actually multiplies. So there are two distinct quantities begging to be normalized: the reward r_RM before it is mixed with the KL term, and the advantage A_t before it scales the gradient. They are normalized for different reasons, and conflating them is the classic mistake.
Running mean/std reward whitening
The most common treatment is to whiten the RM score against a running estimate of its own mean and standard deviation, maintained across the whole run. Keep aggregate statistics updated each batch and transform:
r_norm = ( r_RM − μ_run ) / ( σ_run + ε )with ε a small constant (say 1e-8) guarding against division by zero. The running moments are typically tracked with a numerically stable online update — Welford’s algorithm, or an exponential moving average when you want the estimate to track a drifting policy. This does two things at once. Subtracting μ_run kills the meaningless offset, so the reward is centered near zero. Dividing by σ_run forces the reward onto a roughly unit scale, which is the key to the KL interaction: it keeps r_norm and β·KL in the same numerical ballpark, so a single β means the same thing across the run rather than being swamped as the raw reward scale wanders.
Per-batch whitening and baseline subtraction
An alternative (or complement) to a global running statistic is to whiten within each batch: compute the mean and std over the rewards in the current rollout and normalize against those. Per-batch mean subtraction is really a baseline in the REINFORCE sense — subtracting a constant that does not depend on the action leaves the policy gradient unbiased while shrinking its variance. The batch mean is a cheap, serviceable baseline.
The trade-off is bias versus responsiveness. A running statistic is smooth and stable but lags the policy; a per-batch statistic tracks the current distribution exactly but is noisy for small batches, and it introduces a subtle coupling — a completion’s normalized reward now depends on the other samples it was batched with. Many implementations subtract a running (or EMA) mean for the reward to keep the KL scale honest, then apply per-batch whitening to the advantages downstream. The two levers are not redundant: one conditions the signal, the other conditions the gradient.
Advantage normalization
After GAE turns the per-token rewards into advantages A_t, those advantages are whitened again — almost always per mini-batch:
A_norm = ( A_t − mean(A) ) / ( std(A) + ε )This is a different job from reward normalization. The advantage is what multiplies ∇ log π in the PPO surrogate, so its scale is the effective learning-rate multiplier on each step. If advantages have variance 100 one batch and 0.01 the next, the effective step size swings by four orders of magnitude and training becomes erratic. Forcing unit variance makes each update a well-scaled step regardless of how large the rewards happened to be, which is exactly what lets a fixed learning rate and a fixed PPO clip range ε_clip behave consistently across the run. Reward whitening keeps the signal commensurate with the KL penalty; advantage whitening keeps the gradient commensurate with the optimizer. You want both.
A worked numeric example
Take one rollout batch of four completions with raw RM scores [8.0, 6.0, 7.5, 6.5]. The batch mean is μ = 28.0 / 4 = 7.0. Deviations are [+1.0, −1.0, +0.5, −0.5]; squared, [1.0, 1.0, 0.25, 0.25], summing to 2.5, so the (population) variance is 2.5 / 4 = 0.625 and σ = √0.625 ≈ 0.79.
r_norm = (r − 7.0) / (0.79 + 1e-8)
8.0 → (+1.0)/0.79 ≈ +1.27
6.0 → (−1.0)/0.79 ≈ −1.27
7.5 → (+0.5)/0.79 ≈ +0.63
6.5 → (−0.5)/0.79 ≈ −0.63The offset of ~7 is gone and the values now sit in roughly [−1.3, +1.3] — the same order of magnitude as a KL penalty like β = 0.1 times a per-token KL of a few nats. Before whitening, a raw reward of +8 would have utterly dwarfed that penalty; after, the two terms genuinely trade off. Note how the best-of-four completion gets a positive signal and the worst a symmetric negative one — the batch mean is doing baseline duty.
Reward clipping
Whitening handles the typical case; clipping handles the tail. A reward model occasionally emits a wild outlier — an out-of-distribution completion it scores absurdly high or low — and a single such score can dominate a batch’s statistics and yank the update. Clipping bounds the signal:
r_clip = clip( r_norm, −c, +c ) # e.g. c = 5 (in std units, post-whitening)Applied after normalization, the clip is naturally expressed in standard deviations, so a bound of ±5 means “ignore anything beyond five sigma.” This is defense-in-depth alongside PPO’s own ratio-clipping ε_clip, which bounds how far the policy ratio can move per step; reward clipping instead bounds how extreme the signal can be before it enters the advantage. Set c too tight and you throw away real gradient on genuinely good or bad samples; too loose and outliers leak through. It is a safety rail, not a primary knob — the whitening does the everyday work.