From paired preferences to a single bit
Preference tuning grew up on pairwise data. RLHF trains a reward model on pairs (a winner y_w and a loser y_l for the same prompt x); DPO collapses that into one loss but still consumes the same pair. The pair is the expensive part — getting two responses to one prompt and asking which is better is slow, and much real-world feedback is not shaped that way: a user upvotes one reply, flags another, abandons a third, each an isolated verdict on a single output.
KTO throws the pairing requirement away. Its unit of data is one (x, y) plus one bit: y is desirable or undesirable given x. You can have a thousand desirable examples and fifty undesirable ones, from entirely different prompts, and KTO still trains. This is the data-flexibility win — the format matches how feedback is actually logged (accepts, rejects, ratings binarized around a threshold) rather than forcing it into curated A-vs-B pairs. The question is how to build a sensible gradient from an unpaired signal, and the answer comes from behavioural economics.
The Kahneman-Tversky value function
Prospect theory’s central claim is that humans judge outcomes not on absolute wealth but on gains and losses relative to a reference point. Its value function v(z) passes through that point at z = 0, is concave for gains (z > 0 — the second won dollar thrills less than the first) and convex for losses (z < 0), and is steeper on the loss side: a loss hurts more than an equal gain pleases — loss aversion.
KTO’s insight is that a language model has its own notion of gain and loss: how much more (or less) probability it puts on an output than a frozen reference model. Turn that into the value-function input z, pass it through an S-shaped curve, and you get a per-example objective that makes desirable outputs feel like gains and undesirable ones like losses — without ever comparing two outputs. The logistic sigmoid σ(z) = 1 / (1 + e^(-z)) supplies the S-curve: concave for z > 0, convex for z < 0, exactly the gains/losses geometry prospect theory describes.
The implied reward
KTO, like DPO, never trains a separate reward model. It reads the reward straight off the policy through the KL-regularized-RL solution: the optimal policy’s log-ratio against the reference is the reward up to scale. Define the implied reward of output y for prompt x as
r_θ(x, y) = β · log( π_θ(y|x) / π_ref(y|x) )Here π_θ is the model being trained, π_ref is the frozen starting checkpoint (usually the SFT model), and β > 0 controls how hard the KL leash pulls the policy back toward the reference. When the policy raises an output’s probability above the reference, r_θ > 0; when it lowers it, r_θ < 0. This scalar is all KTO needs from the network, computed from ordinary log-probabilities in one forward pass through each model. The next step decides what counts as ‘above’ — the reference point.
The reference point: the z_0 baseline
A reward is meaningless in isolation; prospect theory says value is measured against a reference. KTO’s reference point is z_0, an estimate of the typical KL divergence between the policy and the reference over the current batch:
z_0 = KL( π_θ(y’|x) || π_ref(y’|x) )
≈ max( 0, (1/m) Σ_i r_θ(x_i, y’_i) )In practice it is estimated cheaply: shuffle outputs against prompts within the batch so each y’ is mismatched, average the implied rewards of those pairs, and clamp at zero. This z_0 is detached from the gradient — it shifts the origin of the value function but is not itself optimized, which keeps training stable. Read it as a floating baseline for how far the policy has drifted from the reference on average. An output whose implied reward sits above z_0 is a gain; one below it is a loss. That single comparison, r_θ − z_0, is the argument the value function shapes.
The KTO value function and loss
Now assemble the value. For a desirable output we want its reward pushed above the baseline (a gain); for an undesirable one, below (a loss). The sigmoid turns the signed distance from z_0 into a bounded value in (0, 1):
v_KTO(x, y) = σ( r_θ(x, y) − z_0 ) if y is desirable
v_KTO(x, y) = σ( z_0 − r_θ(x, y) ) if y is undesirableIn both cases v is near 1 when the model does the right thing (desirable reward high, undesirable reward low) and near 0 otherwise. The loss is the shortfall from a perfect value, weighted per class:
L_KTO = E_(x,y ~ D) [ w(y) · ( 1 − v_KTO(x, y) ) ]
where w(y) = λ_D if y desirable
w(y) = λ_U if y undesirableMinimizing 1 − v drives v → 1: for desirable examples the gradient raises log π_θ(y|x), for undesirable ones it lowers it — each example moved on its own, no pair in sight.
Gains, losses, and loss aversion via &amp;amp;lambda;_D, &amp;amp;lambda;_U
The two weights λ_D (desirable) and λ_U (undesirable) are where KTO encodes prospect theory’s asymmetry — and where it earns its robustness to imbalance. The sigmoid curve itself is symmetric; loss aversion does not come from the curve, it comes from making λ_U weigh undesirable examples more heavily than λ_D weighs desirable ones, so a ‘loss’ moves the loss function more than an equal ‘gain.’
These same knobs absorb class imbalance. With n_D desirable and n_U undesirable examples, if n_U dwarfs n_D the raw loss is swamped by undesirable examples and the model just learns to suppress everything. KTO counters this by choosing weights so the two classes contribute comparably — the authors recommend keeping
(λ_D · n_D) / (λ_U · n_U) ∈ [ 1 , 4/3 ]So with ten times as many undesirable as desirable examples, you raise λ_D (or lower λ_U) to rebalance. DPO has no equivalent lever: its pairs are, by construction, one winner and one loser each, so it cannot express ‘I have mostly negative feedback.’