TL;DR for operators
A retrieved document does not need to look malicious to corrupt a generative-search answer. In Counter-GEO-Bench, carefully constructed documents remain fluent and topically relevant while inserting a targeted false claim; in the undefended pipeline, 55.7% of these attacks move the generated answer toward that claim across three tested models.1
The evaluated general-purpose guardrails do not solve the problem well. Granite Guardian lowers pooled attack success rate (ASR) from 55.7% to 54.0%, while Llama Guard 3 reaches 52.5%. Both block less than 1% of attacked chunks. A much lower result from one NeMo configuration is misleading because the system refuses almost every query, including benign ones.
The paper’s proposed C-GEO Guard is more effective within this benchmark: pooled ASR falls to 29.2%, a 26.5 percentage-point reduction, while average clean and legitimate optimized-query accuracy changes by only +0.2 percentage points on the primary evaluation. That leaves substantial residual attack success, but it shows why factual manipulation should be treated as its own retrieval-security problem rather than assumed to be covered by toxicity or policy filters.
For operators, the key measurement change is as significant as the detector: evaluate attacked, clean, and legitimately optimized content together. A defense that suppresses attacks by refusing broadly or deleting useful contradictory evidence is not delivering the same protection as one that selectively identifies manipulation.
Retrieved evidence can look normal and still change the answer
A generative-search system typically retrieves several relevant sources, ranks passages, and asks an LLM to synthesize an answer. The difficult case is not necessarily a page containing prompt injection, toxic language, or an obvious policy violation. It can be a page that still looks like ordinary informational content but has been optimized so that a targeted false claim survives retrieval and enters the synthesis context.
Counter-GEO-Bench studies this narrower problem. The authors begin with GEO-Bench queries and create paired versions of a target document: an information-preserving (IP) rewrite that applies optimization without changing the underlying information, and an information-distorting (ID) rewrite that introduces targeted misinformation. A quality gate checks semantic similarity, length deviation, perplexity, and judged naturalness. Of 1,000 starting queries, 250 pass both IP and ID gates; human verification removes three defective cases, leaving 247 paired benchmark instances.
That pairing matters because the experiment is not asking whether optimization itself is harmful. It is asking what changes when factual distortion is introduced while much of the document’s legitimate, optimized surface remains comparable.
Under the undefended condition, the three victim models average 55.7% ASR, with a 95% bootstrap confidence interval of 53.1%–58.2%. In this controlled setting, more than half of the accepted single-document attacks are therefore sufficient to pull the answer toward the intended false claim.
Generic safety filters barely engage with the manipulated content
The first practical question is whether an existing safety layer already covers the threat.
The benchmark suggests a mismatch. Granite Guardian and Llama Guard 3 are applied as chunk filters, yet they block only 0.61% and 0.83% of ID chunks respectively. Their pooled ASRs remain 54.0% and 52.5%, compared with 55.7% undefended. Granite Guardian’s estimated 1.6 percentage-point reduction is not statistically significant in the paired bootstrap test; Llama Guard’s 3.2-point reduction is significant but small relative to the baseline attack rate.
| Defense | Average ID ASR | ID chunks blocked | Primary interpretation |
|---|---|---|---|
| Undefended | 55.7% | — | Baseline vulnerability |
| Granite Guardian | 54.0% | 0.61% | Little measurable protection |
| Llama Guard 3 | 52.5% | 0.83% | Small but significant reduction |
| C-GEO Guard | 29.2% | 10.27% | Substantially stronger task-specific filtering |
The mechanism is consistent with the benchmark’s threat model. These documents need not contain the kinds of safety-taxonomy violations a general guard model is trained to identify. They can remain fluent, relevant, and informational while changing a fact that matters downstream.
For a product team, this means a policy filter and a factual-manipulation control should not automatically be treated as interchangeable components. They may inspect the same text while solving different classification problems.
A lower ASR can still represent a failed defense
Attack-only benchmarks create another measurement problem: they can reward systems for suppressing output rather than correctly identifying manipulated evidence.
The paper’s clearest example is NeMo Self-Check on Llama-4. Its ID ASR falls to 0.4%, which would appear exceptional if viewed alone. But clean accuracy falls to 1.8%, IP accuracy to 1.4%, and the configuration blocks roughly 98% of queries. The low attack rate is therefore dominated by near-total refusal.
The paired design exposes a second failure mode. Granite Guardian slightly increases Llama-4 ASR, from 54.7% undefended to 56.9%. The paper attributes this case to filtering that can remove legitimate evidence contradicting the injected claim. Once disagreement among sources is reduced, the remaining manipulated evidence can become more influential during synthesis.
This changes the evaluation target. A useful defense needs to suppress factual manipulation while preserving enough benign evidence for the generator to answer correctly. ASR alone cannot distinguish selective detection from broad damage to the evidence pipeline.
GEO-specific detection performs better, but attacks remain
After establishing the mismatch with general guardrails, the paper introduces C-GEO Guard, a 184M-parameter DeBERTa-based detector trained contrastively to separate manipulated chunks from paired benign material. Each attack class is represented by a prototype centroid. At inference, the system embeds a candidate chunk, measures its maximum cosine similarity to those prototypes, and blocks the chunk when the score exceeds threshold $\tau$.
At the main operating threshold, C-GEO Guard lowers average ASR from 55.7% to 29.2%. The 26.5 percentage-point reduction has a paired-bootstrap 95% confidence interval of 23.8–29.3 points and $p < 0.001$.
The utility result is equally relevant to interpretation. Across clean and IP conditions, average accuracy changes by +0.2 percentage points versus undefended inference. ID-condition answer quality is 4.48, compared with 4.49 undefended. The detector blocks 10.27% of ID chunks versus 2.30% of clean and 2.16% of IP chunks.
These numbers support selective detection within the benchmark, not elimination of the threat. An ASR of 29.2% still means that a substantial share of attacks succeed after filtering.
Transfer tests support the signal, not universal robustness
The paper also tests whether the detector has simply learned wording peculiar to the original rewriting procedure.
With all 247 instances rewritten using GPT 5.5 under the original attack instructions, C-GEO Guard reduces Qwen ASR from 55.7% to 22.1%, a 60.4% relative reduction. This transfer comes with some utility cost in that experiment: IP accuracy falls from 86.8% to 82.6%.
A stricter robustness test uses an independently written GPT 5.5 attack template that omits several features of the original instructions. Here, ASR falls from 50.6% to 34.4%, a 32.0% relative reduction, while IP accuracy changes from 87.0% to 88.3%.
These experiments provide evidence that the detector is not limited to exact Sonnet-generated wording. They do not establish robustness to arbitrary manipulation. Detector-aware attacks, manual adversarial editing, style-transfer attacks, coordinated multi-document attacks, and open-set manipulation remain outside the evaluation.
What changes for RAG and generative-search operators
What the paper directly shows: within a controlled English-language generative-search harness, factual manipulation embedded in otherwise plausible optimized content is frequently effective; the tested generic guardrails provide limited usable protection; and a GEO-specific detector performs substantially better while preserving aggregate benign utility on the main benchmark.
Cognaptus inference: retrieval security should distinguish at least three decisions that are often collapsed into one “guardrail” layer: whether content violates safety policy, whether retrieved evidence may be factually manipulated, and whether the evidence set still contains enough independent disagreement or provenance to support synthesis. A GEO-specific detector can address the second decision, but residual attacks make complementary provenance, cross-source consistency, or external verification controls plausible system-level additions.
The affected users are teams operating RAG or generative-search systems where retrieved third-party content can influence factual answers. The affected decision is how to evaluate and place retrieval-time defenses. The relevant condition is an environment in which ordinary-looking content can be adversarially optimized. The boundary is that this paper tests one controlled implementation rather than production search platforms.
The benchmark is component evidence, not a commercial safety result
The evaluation contains 247 English-language instances, three open-weight victim LLMs, one manipulated target document per query, and a fixed retrieval-and-reranking stack. It does not test proprietary models, multilingual behavior, coordinated poisoning across several sources, arbitrary attack classes, or adaptive adversaries targeting the detector itself.
Those limits do not erase the benchmark result. They determine what decision it can support.
For a product team, the paper provides a credible reason to test factual manipulation separately from conventional content safety and to measure attack suppression alongside clean and legitimate-optimization utility. It does not establish that C-GEO Guard, its threshold, or its measured ASR will transfer unchanged into a commercial generative-search product.
The broader contribution is therefore diagnostic as much as defensive: once retrieved misinformation is evaluated together with the evidence a defense accidentally removes, “lower attack success” stops being a sufficient definition of success.
Cognaptus: Automate the Present, Incubate the Future.
-
Bing Zheng and Zongyao Zhao and Wenming Yang (2026). Counter-GEO-Bench: Evaluating Defenses Against Information-Distorting Generative Engine Optimization. arXiv:2609.02316. https://arxiv.org/abs/2609.02316 ↩︎