Confidence Is Not a Stop Signal: Test Whether the Model Knows When Information Is Missing
TL;DR for operators A model can be given an explicit way to say “the available information is insufficient” and still choose an unsupported answer most of the time. Tahermazandarani, Mahmood, Islam, and Sheng test this directly across five LLMs.1 They remove the correct answer from medical multiple-choice questions, replace it with an insufficient-information option, and observe abstention rates ranging from just 0.156 to 0.382. Reported unsafe rates range from 0.186 to 0.828. In a separate experiment, progressively stronger warnings that the clinical information may be incomplete or ambiguous also produce little reduction in model confidence. ...