TL;DR for operators

Publishing images that remain useful to people while making them poor training material for unauthorized models creates an awkward engineering constraint: stronger interference often comes with more visible image degradation.

The paper Leveraging Imperfect Restoration for Data Availability Attack1 shows that those two objectives need not move together. Its proposed method deliberately restores some of the visual damage created by an existing convolution-based poisoning technique while leaving enough class-specific structure to continue disrupting learning.

On ImageNet-100, the method improves all five reported image-quality metrics relative to CUDA. At the same time, it produces lower clean-data accuracy in the paper’s principal supervised and self-supervised comparisons, where lower accuracy means a stronger data-availability attack. The practical implication is not that restoration removes poisoning, but that incomplete restoration can preserve the training-time signal while recovering perceptual quality.

For organizations considering this kind of protection, two constraints are decisive. Effectiveness falls sharply when only part of a dataset is poisoned, and defenses such as AVATAR and COIN recover substantial clean accuracy. For organizations training on externally sourced images, the inverse lesson also matters: visual normality is not evidence that an image is harmless training data.

Visible damage and training damage are different variables

CUDA applies a different convolutional filter to each class. The resulting images contain class-correlated patterns that a model can exploit during training instead of learning the semantic features that should generalize to clean data.

That protection has a cost: convolutional filtering can visibly degrade the images.

The new paper asks a more specific question than simply whether CUDA can be strengthened. If the visible distortion is partly reversed, does the poisoning necessarily disappear with it?

The reported results say no.

On ImageNet-100, CUDA records an LPIPS of 0.272, while Imperfect Restoration Poisoning, or IRP, reduces it to 0.142. SSIM rises from 0.561 to 0.775, and MS-SSIM from 0.831 to 0.903. CLIP-IQA improves from 0.528 to 0.706, while BRISQUE falls from 29.570 to 17.318. Every reported quality measure moves in IRP’s favor.

Yet the images remain difficult training material. With ResNet-18, clean test accuracy after supervised training is 1.98% for IRP versus 6.10% for CUDA on ImageNet-100. Under the paper’s SSL setting, Table 2 reports 9.30% for IRP versus 26.12% for CUDA. Lower is stronger because the attack is evaluated by how badly a model trained on poisoned data performs on clean test images.

The paper therefore separates two quantities that are easy to conflate: how damaged an image looks and how damaging it is to learning.

CUDA’s effect comes from structured class bias, not arbitrary corruption

The paper’s theoretical work focuses primarily on CUDA rather than IRP. That distinction is worth preserving because the mechanism is what makes imperfect restoration plausible.

Its first argument concerns optimization. CUDA filters each training image with a class-specific kernel before the network sees it. The first convolutional layer therefore receives a different gradient from the one it would receive on the clean images. The paper argues that this poisoned gradient deviates from the direction that would improve performance on clean inputs.

The second argument explains why that distortion can become class-specific.

For a CUDA filter correlated with itself, the strongest correlation remains centered. Cross-correlating filters from different classes usually shifts the dominant peak away from the center and produces different blur patterns. Under the paper’s assumptions, the probability that a cross-class correlation peak remains centered is

$$ \frac{1}{\kappa^2}, $$

where $\kappa$ is the filter width and height. For the ImageNet configuration with $\kappa=9$, the paper reports a 0.988 probability that cross-class information is shifted away from the correct location.

The consequence is not merely noisier pixels. The filters introduce class-correlated spatial structure that can function as a shortcut during training.

That mechanism changes how restoration should be interpreted. If the poisoning effect depended only on visible corruption, improving the image should weaken the attack in direct proportion. If it depends instead on structured residual signals, perceptual recovery and training recovery need not coincide.

Imperfect restoration deliberately stops before inversion

IRP starts from a CUDA-filtered image and learns a separate linear restoration filter for each class.

For every class, the method extracts local patches from filtered images and fits coefficients that predict the corresponding original center pixel through least squares:

$$ \widehat{\alpha_c} = \underset{\alpha_c}{\operatorname{argmin}} \sum_i \sum_j \left\| y_{cij} - \alpha_c^T \eta_{cij} \right\|_2^2. $$

The learned restoration filter is then composed with the original CUDA filter. In the paper’s notation,

$$ P_c = R_c \star A_c, \qquad X_{ci}^{\mathrm{IRP}} = X_{ci} \star P_c. $$

The objective is not perfect reconstruction. The restoration recovers enough image structure to improve visual quality while leaving residual class-dependent patterns behind.

This is the paper’s central design contribution: restoration becomes part of the poisoning construction rather than an attempt to remove it.

A filter ablation reinforces that interpretation. Random Sharpness and Double Blur enlarge the effective kernel but do not reproduce IRP’s poisoning performance. The result argues against explaining IRP as a simple consequence of kernel size; the learned restoration procedure itself appears relevant.

The broader experiments test robustness, not a second mechanism

The empirical program then asks whether the effect survives changes in training conditions.

On CIFAR-10, IRP produces 10.39% clean accuracy under supervised learning and 43.24% under SSL, compared with CUDA at 21.89% and 66.58%. Across SimCLR, MoCoV3, SimSiam, and BYOL on ImageNet-100, IRP produces clean accuracies of 9.30%, 10.58%, 7.82%, and 11.60%.

Under adversarial training across six architectures, IRP’s worst reported clean accuracy is also below CUDA’s on all three tested datasets: 35.39% versus 49.44% on CIFAR-10, 28.85% versus 37.95% on CIFAR-100, and 35.58% versus 45.99% on STL-10.

These experiments support robustness across the tested learning algorithms, architectures, augmentations, and defenses. They do not establish a universal poisoning mechanism for arbitrary models or modalities.

There is also a numerical inconsistency worth retaining rather than smoothing over: Table 2 reports CUDA’s ImageNet-100 SSL accuracy as 26.12%, while nearby prose reportedly gives 21.89%. The structured table value is the safer basis for comparison.

Coverage is the deployment constraint that changes the result most

The paper’s partial-poisoning experiment puts an important boundary around the headline numbers.

With the full CIFAR-10 training set poisoned, IRP drives clean accuracy to 10.39%. With 80% poisoned data mixed with clean images, accuracy jumps to 85.70%. At 60% coverage it reaches 89.89%, at 40% 92.56%, and at 20% 93.07%.

That is not a small deterioration. It means benchmark results obtained under complete dataset protection should not be treated as evidence that the same protection will remain effective when protected images are surrounded by abundant clean substitutes.

Defenses matter as well. On CIFAR-10, AVATAR raises clean accuracy on IRP-poisoned data to 54.78%, UEraser to 26.55%, and COIN to 53.32%. IRP remains stronger than CUDA in those reported comparisons, but the recovered accuracy is large enough that “defense-proof” would overstate the evidence.

Two different operational decisions follow from the same result

For creators and platforms, the paper reframes the engineering target. Maximizing visible corruption is not necessarily the correct objective. A more relevant product question is how far learnability can be reduced while maintaining the perceptual quality required for normal viewing, distribution, or commercial use.

For organizations training models from public or externally supplied images, the risk points in the opposite direction. An ingestion pipeline cannot assume that an image is benign because it looks normal. Deliberately structured transformations may remain difficult to detect visually while materially changing training behavior. Purification, provenance controls, augmentation, and robustness testing therefore belong in the same evaluation pipeline rather than being treated as independent safeguards.

The evidence does not establish that IRP generalizes beyond the tested image-classification settings, nor that its learned filters admit the same closed-form explanation the paper develops for CUDA. It does show something narrower and operationally consequential: restoring perceptual quality does not necessarily restore normal learnability.

For data-protection systems, that makes imperfect restoration a design variable rather than a contradiction.

Cognaptus: Automate the Present, Incubate the Future.


  1. Yi Huang and Jeremy Styborski and Mingzhi Lyu and Fan Wang and Adams Kong (2026). Leveraging Imperfect Restoration for Data Availability Attack. arXiv:2609.04627. https://arxiv.org/abs/2609.04627 ↩︎