More Thought Is Not Always More Reliable: Routing Reasoning for Social Judgment
TL;DR for operators Extra inference compute is not a monotonic reliability upgrade for socially ambiguous tasks. Across three benchmarks, reasoning-focused models sometimes outperform non-reasoning counterparts, sometimes underperform them, and can become less accurate when deliberation is pushed harder on difficult cases. The practical lesson is not to suppress reasoning. Moderate reasoning, token limits, and adaptive stopping can improve results. Instead, treat reasoning depth as a control variable: decide when to invoke it, when to stop it, and whether the prompt format itself is steering the model toward shortcuts. ...