Frontier · Reasoning Not the model you trained — a full-size one

Learning
to think.

This is not the small model you trained — that one is far too small for this. Here is a full-size commercial model, asked the same question two ways. Same model, same question — one switch. With thinking off, it blurts the first answer that comes to mind — and on the questions built to catch you out, that blurt can be confidently wrong. Turn thinking on and it works through the trap out loud — and gets it right. It didn't get smarter. It learned to draft before it answers.

Thinking

This is the frontier, and the whole point of the switch. The snap answer isn't stupidity — it's what any of us blurts before checking. Thinking out loud is a learned habit: the model was trained to draft, catch its own trap, then answer. How it learned that — rewarding it only when the final answer checks out (RLVR / GRPO) — is the deep end.