TL;DR for operators

A moderation confidence score is useful only if high confidence still means high reliability under the traffic conditions that matter.

That assumption breaks sharply in the experiments reported by Hong, Jung, and Kim in Guard Models Are Overconfident Where Base Models Are Uncertain.1 At a confidence threshold of 0.99, Llama-Guard-3 still leaves 11% of adversarial harmful prompts as confident false negatives, WildGuard 12%, and ShieldGemma 37%. The failures are therefore not concentrated in an obvious low-confidence region that can simply be routed to abstention or human review.

The paper also compares each guard with its corresponding base language model. On many prompts where the guard confidently predicts “safe” incorrectly, the base model remains more uncertain. That comparison does not make the base model a better moderation system. It indicates that uncertainty visible in the underlying model can become suppressed or distorted after safety specialization.

For teams that use guard confidence to allow, block, or escalate interactions, the immediate implication is narrower but consequential: validate calibration under adversarial inputs, not just clean traffic. Base-model uncertainty and internal representation diagnostics may eventually provide auxiliary signals, but this paper does not establish a deployment-ready correction.

A 0.99 threshold does not isolate the dangerous cases

Suppose a moderation system treats uncertain classifications conservatively. Low-confidence decisions are escalated; high-confidence decisions are allowed to proceed automatically. That design works only if mistakes become more concentrated as confidence falls.

The experiments show why jailbreak traffic can violate that assumption.

The study evaluates five generative guard models on a pool containing 713 clean harmful prompts, 500 benign prompts, and 3,965 adversarial harmful prompts produced using six attack methods. Several guards look highly reliable on clean inputs. Llama-Guard-3, for example, has clean expected calibration error of 0.013, while WildGuard reaches 0.004.

Expected calibration error measures the gap between stated confidence and observed accuracy across confidence ranges. A well-calibrated system should make predictions with roughly 90% empirical accuracy among cases it assigns around 90% confidence.

Under adversarial input, those same guards deteriorate substantially. Llama-Guard-3’s ECE rises from 0.013 to 0.308; WildGuard’s rises from 0.004 to 0.196. ShieldGemma reaches an adversarial ECE of 0.684, although its clean calibration was already comparatively poor at 0.325.

More consequentially for routing systems, some errors remain nearly certain. Even at a 0.99 confidence threshold, meaningful shares of adversarial harmful prompts remain confident false negatives.

This changes the role confidence can safely play. A threshold calibrated on ordinary traffic cannot be assumed to identify the risky tail once adversarial prompting changes the relationship between confidence and correctness.

The base model is often uncertain where the guard is not

High-confidence failures leave an important ambiguity. Perhaps the jailbreak simply transforms a harmful prompt into something intrinsically difficult to classify. If so, guard overconfidence might reflect an unusually ambiguous input rather than something introduced by safety fine-tuning.

The matched base-model comparison tests that possibility.

For each guard, the researchers evaluate the corresponding base LM through the same binary safe/unsafe decision framework. They then examine adversarial prompts on which the guard produces a false negative.

For Llama-Guard-2, the base model has higher binary verdict entropy on 82% of these failures. The share is 80% for Llama-Guard-3 and 93% for WildGuard. ShieldGemma is much weaker at 51%.

Here entropy is simply a measure of uncertainty over the two possible verdicts. A positive base-minus-guard entropy gap means the base model is less certain than the specialized guard on exactly the same input.

A second comparison reaches a compatible result. The researchers measure how much each model’s safety-verdict distribution changes when a clean harmful prompt receives an adversarial wrapper. Across four matched pairs used for this analysis, guards shift more than their base models, with mean attack-sensitivity gaps ranging from +1.71 for Llama-Guard-3 to +4.91 for ShieldGemma.

The interpretation needs discipline. The base LM is not presented as a substitute safety classifier. Instead, it provides a diagnostic reference: on many failures, uncertainty has not disappeared from the underlying model family. The guard has become more decisive.

The divergence develops in later layers

The next question is where that increased certainty appears.

The paper projects intermediate hidden states through the model’s final normalization and language-model head, allowing the researchers to track an implied unsafe probability across network depth. On adversarial false negatives, the difference between base and guard probabilities grows progressively in later layers. At the final layer, the reported gaps reach roughly 0.45 for Llama-Guard-3, 0.7 for WildGuard, and 0.9 for ShieldGemma.

Representation analysis adds another pattern. Guard models develop sharper separation between clean safe and harmful inputs, while their late-layer representations occupy lower-dimensional effective spaces than those of their matched base models. At the penultimate layer, base-to-guard effective-rank ratios are approximately 1.4× for Llama-Guard-2, 1.7× for Llama-Guard-3, 3.4× for WildGuard, and 4.0× for ShieldGemma.

Effective rank summarizes how many spectral directions meaningfully contribute to the representation. A lower value indicates that representations have become concentrated into fewer dominant directions.

The adversarial harmful inputs are often positioned closer to the clean-safe side of the guard’s representation geometry. Together, these observations are consistent with a specialized classifier that has developed a sharper decision structure but can map adversarially altered harmful content confidently onto the wrong side.

They do not establish causality. The paper does not intervene on the compressed representations or manipulate the identified directions to show that compression causes overconfidence.

The robustness checks narrow the claim rather than universalize it

Several appendix analyses test whether the main discrepancy is an artifact of experimental choices.

The base-model result survives substantial interface variation. Across 97 valid alternative prompting settings, the base model has higher entropy on guard false negatives in 79 cases and lower adversarial ECE in 90. That makes a single five-shot classification prompt an unlikely explanation for the main pattern.

Attack type matters, however. For Llama-Guard-3, GCG and AutoDAN produce comparatively small guard-base shifts, while TAP, PAIR, AutoDAN-Turbo, and the custom attack produce larger output and intermediate-layer differences.

Scale also fails to give a simple explanation. ShieldGemma remains worse calibrated than its base model from 2B through 27B parameters, while Llama-Guard-3-1B reverses the ordering seen in the larger Llama-Guard-3 comparison.

These are meaningful boundaries. The study supports an adversarial calibration problem across the tested guards, but not a universal mechanism that behaves identically across attacks, families, or scales.

Moderation teams need an adversarial calibration test, not just an accuracy test

The direct paper result concerns measurement. The business consequence appears when confidence is connected to an action.

A product may use a guard’s probability to decide whether to block a request automatically, allow it without further inspection, invoke another model, or send the interaction to human review. In that workflow, confidence is not descriptive metadata. It allocates review and determines which cases receive additional scrutiny.

Cognaptus therefore infers three practical evaluation requirements from the study:

Evaluation question What the paper indicates Boundary
Does clean calibration remain valid under jailbreak traffic? Not reliably; ECE can increase sharply after adversarial wrapping. Magnitudes differ across guards and attacks.
Can very high confidence be used to bypass additional review? High-confidence false negatives survive thresholds as high as 0.99. The paper does not specify an optimal replacement policy.
Could another uncertainty signal help? Matched base LMs often remain more uncertain on guard failures. Base-model uncertainty is diagnostic here, not a validated production safeguard.

This suggests that red-team and governance dashboards should separate detection accuracy from confidence reliability and should report high-confidence false negatives under adversarial stress. If a threshold controls escalation capacity, its performance needs to be evaluated on the attacked distribution where the threshold will be relied upon.

The paper diagnoses the failure but does not repair it

The strongest evidence is comparative: adversarial attacks can produce severely overconfident guards, and matched base models often preserve more uncertainty on the same failures.

The internal explanation is less settled. Late-layer divergence, sharper class separation, and lower effective rank accompany the observed errors, but the study does not demonstrate that changing those properties will correct calibration.

Nor does it establish that querying both a guard and its base model is the right production architecture. Doing so could introduce latency, cost, interface sensitivity, and new failure modes that are outside this evaluation.

The useful next step is therefore not to replace confidence with a different unvalidated score. It is to stop assuming that clean confidence calibration transfers automatically to adversarial traffic.

A guard that says “safe” with 99% confidence can still be exactly the case that required another check.

Cognaptus: Automate the Present, Incubate the Future.


  1. Jonghyun Hong and MinJae Jung and Minwoo Kim (2026). Guard Models Are Overconfident Where Base Models Are Uncertain. arXiv:2609.36477. https://arxiv.org/abs/2609.36477 ↩︎