More Critics, Less Gain: Self-Questioning Has a Stability Limit
TL;DR for operators When a model can check its own reasoning, more self-checks are not automatically better. On GSM8K, Llama-3.2-1B rises from a 33.14% chain-of-thought baseline to 35.28% with one alternative critique and 35.84% with two, but falls back to 33.43% with three. The broader analysis links higher disagreement among these self-generated alternatives to greater reward variance and less stable policy updates. ...