Cover image

Confidence Without Reliability: Why Jailbreaks Make Guard Models Certain and Wrong

TL;DR for operators A moderation confidence score is useful only if high confidence still means high reliability under the traffic conditions that matter. That assumption breaks sharply in the experiments reported by Hong, Jung, and Kim in Guard Models Are Overconfident Where Base Models Are Uncertain.1 At a confidence threshold of 0.99, Llama-Guard-3 still leaves 11% of adversarial harmful prompts as confident false negatives, WildGuard 12%, and ShieldGemma 37%. The failures are therefore not concentrated in an obvious low-confidence region that can simply be routed to abstention or human review. ...

October 6, 2026 · 7 min · Zelina