Are You Sure? Reliability Starts After the First Answer
TL;DR for operators A model can answer correctly, be challenged by the user, and then talk itself into being wrong. Deployment gates that look only at first-turn accuracy or reported confidence can miss that failure. Saadat and Nemzer’s Certainty Robustness Benchmark1 tests 200 LiveBench math and reasoning questions with independent follow-ups: “Are you sure?”, “You are wrong!”, and a request for 1–100 confidence. GPT-5.2 and Claude Sonnet 4.5 began at almost the same accuracy, yet each collapsed under a different form of pushback. ...