From language model to reward model: the value head
A reward model is not trained from scratch — it is a pretrained language model with its head swapped out. A normal LM ends in a projection to vocabulary logits: the last hidden state h_t ∈ ℝ^d becomes a distribution over the next token. For a reward model you discard that [d, vocab] unembedding and bolt on a tiny value head: a single linear layer w ∈ ℝ^d (plus bias) that maps a hidden state to one number.
You feed the full sequence — prompt concatenated with response — through the transformer, take the hidden state at the final token (the position that has attended to everything before it), and project it: r(x, y) = w · h_last + b. Shapes: the transformer maps tokens → H: [N, d], you select h_last: [d], and the head produces a scalar [1]. The whole backbone is fine-tuned along with the head. Everything the pretrained model knows about language is reused; training only has to learn the far smaller thing — a direction in hidden space that correlates with human preference.
The Bradley-Terry preference model
Humans are unreliable at absolute scoring (‘rate this 7.3/10’) but reliable at comparison (‘A is better than B’). So the data is pairs: a prompt x, two responses, and a label saying which won. The Bradley-Terry model (1952) is the bridge from a latent scalar score to the probability of such a comparison. It posits that the odds of y_w beating y_l are the ratio of their exponentiated strengths:
P(y_w > y_l | x) = exp(r(x, y_w)) / ( exp(r(x, y_w)) + exp(r(x, y_l)) )
= σ( r(x, y_w) - r(x, y_l) )
where σ(z) = 1 / (1 + exp(-z)) is the logistic sigmoidThe second line is the key simplification: divide top and bottom by exp(r_w) and the pair-probability collapses to a sigmoid of the reward difference alone. This is exactly binary logistic regression where the ‘feature’ is r_w - r_l. Preference probability depends on the gap between rewards, never their individual magnitudes — a fact that governs everything that follows.
The pairwise ranking loss
Fitting Bradley-Terry by maximum likelihood gives the loss the whole RM is trained on. For a dataset D of triples (x, y_w, y_l) with y_w the human-chosen response, minimise the negative log-likelihood of the observed preferences:
L(θ) = - E_(x, y_w, y_l) ~ D [ log σ( r_θ(x, y_w) - r_θ(x, y_l) ) ]Read it plainly: push the chosen response’s reward above the rejected one’s, and the wider the correct margin the lower the loss. Because σ saturates, the penalty for getting a pair confidently wrong (large negative margin) grows nearly linearly, while a pair already ranked correctly with a big margin contributes almost nothing — the gradient concentrates on the pairs the model still gets wrong or is unsure about. Both responses pass through the same network with shared weights θ, so one training step nudges the reward surface for both at once. In practice this is one forward pass per response, one subtraction, and a logsigmoid — computed as -log σ(z) = softplus(-z) to avoid overflow.
Why only relative reward is learned
Look again at the loss: it depends on r_w - r_l and nothing else. Add any constant c to every reward the model outputs and the difference is unchanged: (r_w + c) - (r_l + c) = r_w - r_l. The loss is shift-invariant. There is no term anywhere that pins down where zero sits, so the training signal simply cannot determine absolute reward — only reward differences are identifiable.
The practical consequences are sharp. First, an RM output of +4.2 means nothing on its own; only r(A) - r(B) is meaningful. Second, two correctly-trained RMs can disagree wildly in absolute scale and offset while ranking every pair identically. Third — and this is why it matters downstream — whatever consumes the reward (a PPO loop) must not depend on the offset; it typically normalizes rewards to zero mean per batch. You can think of the RM as learning a potential surface defined only up to a constant, like altitude measured without a fixed sea level: the slopes and relative heights are real, the absolute number is a free gauge.
A worked example
Take one prompt and two candidate responses. The RM, in a forward pass, emits r(x, y_w) = 2.0 for the human-preferred response and r(x, y_l) = 1.0 for the other. The margin is z = r_w - r_l = 1.0. Then:
σ(z) = 1 / (1 + e^-1.0) = 0.731 # model says 73% chance chosen > rejected
loss = -log(0.731) = 0.313 # per-pair NLL
gradient of loss w.r.t. z: d L / d z = σ(z) - 1 = -0.269
-> descent moves z UP by 0.269 * lr : widen the gapNow the instructive cases. If the RM had scored them equal (z = 0): σ(0) = 0.5, loss = -log 0.5 = 0.693 — maximum uncertainty. If it got the pair backwards (z = -1.0, chosen scored lower): σ(-1) = 0.269, loss = 1.313 — over four times the correct-margin loss, and the gradient σ(z) - 1 = -0.731 is much larger, hauling the two rewards apart hard. Notice the gradient magnitude is exactly 1 - σ(z): the model’s own error probability on that pair. Confident-correct pairs (z large) yield near-zero gradient; the learning budget flows to the mistakes.
Training dynamics and data
Training is standard supervised learning over the preference set: sample a batch of pairs, forward both responses, compute the mean pairwise loss, backprop into the shared backbone and head. A few things bite in practice. One epoch is common — reward models overfit fast, memorising annotator quirks and surface features (length, formatting) rather than quality, so held-out preference accuracy is watched closely and training stopped early. Pairs from the same prompt are the unit of signal; the loss never compares responses across different prompts, which is another way to see why absolute scores across prompts are not mutually calibrated. When a prompt has a full ranking of K responses, the common trick (InstructGPT) is to expand it into all C(K,2) pairs and average their losses within the prompt — more sample-efficient and less overfit-prone than treating each pair independently. Learning rates stay small: you are steering a large pretrained model with a weak scalar signal, not reshaping it.