Where You Pause Changes What You Forget
TL;DR for operators Kim and colleagues’ study of pause-token fine-tuning1 points to a training rule that is easy to miss if pause tokens are treated mainly as extra thinking time. At an equal pause-token budget, putting pauses at semantic boundaries is the only tested placement that consistently improves both math and code averages over ordinary supervised fine-tuning. The stronger variant also masks the loss on the pause tokens themselves. ...