More Shots, More Coverage? Measure the Set, Not the Sample
TL;DR for operators If your system produces five patches, ten evidence candidates, or twenty molecular designs, choosing its model and inference settings from single-answer benchmarks can select the wrong configuration for the workflow you actually run. The central measurement problem is redundancy. Several outputs can each be valid and individually strong while repeatedly landing on the same task-relevant result. Conversely, outputs that look different in wording or structure may add no genuinely new option. ...