TL;DR for operators
An enterprise assistant can retrieve three passages that all look relevant and still answer from a corrupted evidence base. Once the generator has been instructed to rely on retrieved material, the integrity of that material becomes part of the system’s reliability boundary.
Iliano Fasolino’s 2026 experiment1 measures this directly. In a small local RAG setup, poisoned-answer accuracy fell from 69.4% in the nominal 0-of-3 corruption condition to 43.5% when all three retrieved passages were corrupted. The rate of runs where a clean-context answer was correct but its poisoned counterpart was not rose from 9.5% to 34.0%.
The practical implication is not that retrieval makes systems less reliable. It is that retrieval shifts part of reliability from model parameters into the evidence pipeline. Production controls therefore need to cover provenance, source independence, conflicting evidence, and what happens when the model refuses to answer. The paper also contains a warning for benchmark design: even the 0-of-3 condition produced substantial disagreement because the clean and comparison generations were sampled independently.
Retrieval creates an evidence trust boundary
A standard RAG system retrieves relevant passages and gives them to a generator as evidence. The expected benefit is straightforward: the model can answer from information that is newer, private, or more specific than what is encoded in its parameters.
But the generator does not independently certify that evidence before using it.
If a relevant passage is modified while preserving enough surface plausibility to remain useful-looking, grounding can transmit the corrupted claim instead of correcting it. This is the threat examined in the paper: retrieved passages are altered after retrieval but before prompt construction. Retrieval ranking, the retriever itself, and the model weights are left unchanged.
The experiment calls this document poisoning. Three rule-based corruptions are tested: replacing entities with plausible alternatives, replacing numbers with nearby incorrect values, and inserting negations that reverse factual meaning.
The setup uses a 4-bit Llama 3.1 8B Instruct model, MiniLM embeddings, a FAISS index, and three retrieved passages per query. Forty-nine indexed query slots are crossed with three corruption strategies and four poisoning levels, producing 588 runs and 1,176 clean and poisoned generations.
Reliability falls as corrupted evidence occupies more of the context
The clearest result is the poisoning-level sweep.
| Poisoned passages | Poisoned accuracy | Correct → incorrect | Abstention |
|---|---|---|---|
| 0 / 3 | 69.4% | 9.5% | 21.1% |
| 1 / 3 | 63.9% | 15.0% | 29.3% |
| 2 / 3 | 53.1% | 25.9% | 38.8% |
| 3 / 3 | 43.5% | 34.0% | 51.7% |
The clean arm stays near 78% across these conditions. As corrupted passages take a larger share of the supplied context, poisoned-answer accuracy falls while the probability of converting a previously correct answer into an incorrect one rises.
The paper calls the latter quantity the fooled rate: a concrete paired failure in which the answer generated from clean context is judged correct while the answer generated from the corresponding poisoned context is not.
The progression is substantial, but neighbouring bootstrap intervals overlap. The experiment therefore supports a degradation pattern rather than a precise universal response curve.
There is another measurement issue hidden in the first row. No passage is actually corrupted in the 0-of-3 condition, yet its paired generation still shows a 9.5% fooled rate, and only 76.9% of paired answers are identical. The clean and comparison answers were generated independently at temperature 0.1, while correctness was determined with a brittle keyword matcher.
So 69.4% is not the correct no-attack performance baseline. The roughly 78% clean arm is. More broadly, an adversarial evaluation should establish how reproducible its clean pairs are before interpreting small differences as attack effects.
Poisoning produces refusal as well as wrong answers
A simple mental model of poisoned RAG is that corrupted evidence causes the model to repeat false claims. The recorded behaviour here is more varied.
Across poisoned runs, abstention averages 35.2%, compared with a 21.1% fooled rate. Under full entity-swap poisoning, abstention approaches 69%.
In other words, conflicting or suspicious-looking evidence often causes the model to decline to answer rather than commit to a false response.
This is useful operational information, but abstention should not be counted as successful robustness. A system that refuses instead of returning a false claim has avoided one failure mode while entering another: it has not completed the requested task.
For production monitoring, however, a shift in refusal or “not enough information” rates can become a diagnostic signal. A sudden increase may indicate degraded evidence quality, contradictory retrieval, or changing source composition. Such monitoring needs a normal-behaviour baseline because abstention can also occur for benign reasons.
Three passages do not automatically provide three independent checks
The numeric corruption experiment gives a more specific reason to care about evidence composition.
With number swaps, the fooled rate is 10.2% when zero passages are corrupted and remains 10.2% when one of three is corrupted. It rises to 26.5% when two passages are corrupted and reaches 34.7% when all three are corrupted.
The increase from one poisoned passage to two is 16.3 percentage points, with a bootstrap 95% interval of 2.0 to 30.6 percentage points. Paired McNemar tests are significant at two-of-three and three-of-three poisoning, but not at one-of-three.
The likely mechanism is redundancy: one false number can be countered by two clean passages, while two agreeing corrupted passages can dominate the context presented to the model.
This is evidence for a majority-leaning pattern, not a universal majority threshold. The experiment is too small and noisy to establish a sharp phase transition.
For business systems that answer questions about prices, dates, quantities, balances, limits, or measurements, the stronger design inference is about source independence. Retrieving three passages is weak protection if all three ultimately repeat the same upstream record or can be modified through the same path.
What changes in a production RAG design
The paper directly shows that, in this configuration, answer reliability deteriorates as relevant retrieved evidence becomes increasingly corrupted. It also shows that failures are not homogeneous: some answers flip, some abstain, and a crude lexical-overlap measure captures yet another category imperfectly.
Cognaptus inference begins from that distinction.
For systems using editable internal knowledge bases, cached content, web retrieval, or externally supplied documents, evidence controls should sit between retrieval and answer generation. Provenance can establish where a passage came from and whether it changed. Independent-source validation can reduce dependence on repeated copies of the same underlying claim. Cross-passage disagreement detection can trigger secondary checks for sensitive facts.
Abstention also needs a policy. A refusal may be preferable to confidently adopting corrupted evidence, but a production workflow still needs to decide whether to ask for clarification, retrieve again, consult a second source, escalate to a human, or return an explicit uncertainty state.
Higher-risk evaluations should likewise distinguish several outcomes that this experiment does not fully separate: adoption of the poisoned claim, contradiction of it, abstention, and independent unsupported invention. Keyword accuracy and lexical overlap are too coarse for that classification.
The strategy rankings are provisional
Entity swaps produce the highest average fooled rate among the three tested corruptions: 23.5%, versus 20.4% for number swaps and 19.4% for negation. Entity swaps also produce the highest abstention rate.
Those differences can guide further testing, but they should not become production assumptions about which attack type is intrinsically strongest.
The study uses one quantized 8B model, one primary corpus, 37 distinct query texts represented through 49 indexed slots, rule-based corruptions, and independent paired decoding. Entity replacement can also create internally inconsistent passages that encourage refusal for reasons beyond the falsified fact itself.
The same boundary applies to the paper’s unsupported-generation measure. Its lexical-overlap proxy falls under poisoned context, but this does not establish that hallucination falls. Faithful paraphrases can have low overlap, while an answer that repeats poisoned evidence can have high overlap and still be false.
Grounding needs its own integrity controls
RAG is often treated as a mechanism for improving factuality by connecting a model to external evidence. This experiment shows the complementary engineering problem: once evidence becomes influential enough to ground an answer, corrupted evidence becomes influential for the same reason.
The useful production response is not to reduce reliance on retrieval indiscriminately. It is to make the evidence path auditable and testable: preserve provenance, detect disagreement, check whether sources are genuinely independent, monitor abstention, and measure baseline generation instability before interpreting small robustness changes.
Retrieval can extend what a model knows. It also extends the system boundary that has to be trusted.
Cognaptus: Automate the Present, Incubate the Future.
-
Iliano Fasolino (2026). In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning. arXiv:2609.09243. https://arxiv.org/abs/2609.09243 ↩︎