Forgotten Until Asked Differently: Unlearning Needs an Adversarial Sign-Off
TL;DR for operators A model can pass an ordinary unlearning evaluation while still yielding supposedly forgotten information when the request is reformulated strategically. Gupta and colleagues demonstrate this gap in a controlled benchmark of LLM unlearning: on the 1% forget split, four fine-tuning-based methods produced average adversarial recovery rates between 72.8% and 84.3%, compared with 87.5% for the unprotected model.1 ...