More Compute, Different Jobs: Choosing Inference-Time Reliability Controls
TL;DR for operators When one model answer can trigger a consequential downstream step, spending more inference compute before trusting that answer can help—but the form of that spending matters. In the reported experiment, generating several reasoning attempts and aggregating the answer that repeatedly emerged increased verified acceptance from 56.2% to 64.9%. Asking the same model to critique and revise itself produced a smaller increase, from 47.2% to 50.6%. Adding a second model did not improve the reported acceptance rate: 47.4% of outputs survived cross-model verification versus a 48.7% single-model baseline. ...