05 · Reward model Learned from comparisons only

Distil taste
into a number.

You can't write down "pleasant." So instead of a rule, we show the model pairs and, for each, which one was preferred — here, the warm line (love, heart, grace) over the grim one (blood, death, war). No person picked these: a fixed count of warm words against grim ones stands in for the human, and the model never sees that rule. From comparisons alone it learns to hand out a score by itself. That score is the kind of judge the next page trains a model to chase.

It learns to agree
Share of kept-back pairs it ranks the same way the rule did. A coin toss would score 50%.
—
training steps →
It scores text it never saw
Each dot is a held-out sample. Hover to read it.
↑ learned score  ·  hidden pleasantness → warm grim
⚠️
The crack in the judge

It never understood "pleasant". All it ever saw was which of two texts had more warm words and fewer grim ones — so a wall of one warm word is, by that measure, ideal:

—reward score

A reward model is just another model — with its own blind spots. It learned to rank real samples well, yet a wall of love love love — pure nonsense — scores higher than most genuine text. Nothing optimising against this judge is told the difference. That gap is exactly what the next station drives a truck through. The judge you train against, and the safety you actually want, are not the same thing.