Reasoning Is a Configuration, Not a Switch
TL;DR for operators A legal-translation team may assume it can train a model normally and later enable extra intermediate text whenever higher quality is needed. In this experiment, that deployment-only change produced nearly the same translation quality at far greater output volume. For Qwen3.5 9B, using reasoning during both training and inference reached COMET 82.50 with 6.27 million output tokens. Enabling reasoning only at inference reached COMET 82.32 with 20.67 million tokens—more than three times the output for slightly lower quality. ...