06 · RLHF / Reward hacking Real fine-tuning run

Watch it learn
to cheat.

We take the trained model and reward it for writing pleasant text — measured, crudely, by how often it uses warm words like love, sweet, heart. Then we fine-tune it to score higher. Same starting model, two leashes. Drag through the rounds and watch what "maximise the reward" actually does.

Training round round 1
FREEreward only
0.00reward
0.00drift from language
LEASHEDreward − base-surprise penalty
0.00reward
0.00drift from language
Reward — what you optimise
Both climb. By this number, FREE is winning.
freeleashed
Drift — what you didn't measure
FREE runs away from real language. That's the cheat.
freeleashed

This is the whole problem in one screen. Look only at the reward — the thing the judge measures — and the FREE model looks like the better student: it scores higher. But read its best line of the round — the highest-scoring of the 64 it wrote each time. It found the crack in the reward and jammed it, love love love, until the words stopped meaning anything. That is reward hacking. The leash — a penalty for drifting from the base model — is the only reason the other one stayed fluent. Training a model against a score like this, with a leash holding it near the language it started from, is what the field calls RLHF.

The catch AgentSure exists for: a model that aces its training reward has not been shown to be safe. It has been shown to be good at the reward. Telling those apart takes a measurement the reward itself never made.