Evaluation · Contamination Two real runs, identical diets

It didn't get smarter.
It had seen the paper.

Two models, same size, same number of training steps, and — this is the part that matters — exactly the same number of characters to eat. The only difference is that one of them had the exam paper quietly mixed into its food, the way a leaked test set ends up inside a web crawl. Now give them both the exam.

trained clean
exam leaked in

trained clean
exam leaked in

Read those two screens together. On the exam the contaminated model looks a full better — the kind of jump that gets written up. On a paper it hasn't seen, it is . The number moved. The model didn't. That is what a contaminated benchmark buys you, and from the outside the two are indistinguishable — you cannot tell them apart by looking at the score.

It isn't reciting. The obvious guess is that contamination looks like plagiarism, so we checked: handed both models a passage from the middle of the exam and let them continue greedily. Neither reproduces it — neither followed the exam's own wording for more than characters before going off inventing. Contamination doesn't have to look like copying to inflate a score. It just has to have made the paper a little less surprising, and that is invisible in anything except a paper the model has genuinely never met.

Why the diets had to match. The first version of this experiment simply added the exam on top of the training text, and the leaked model improved on the unseen paper almost as much as on the exam — because it had been fed more, not because it had cheated. That result proved nothing and was thrown away. Here a slice of real training text is held back and swapped for the exam, character for character, so the only thing that differs between the two runs is what was in the food. Next: how big does it have to get before any of this pays off.