Data Unlearning via Inverse Distillation

Aleksei Leonov, Nikita Kornilov, Zhenhe Zhang, Evgeny Burnaev, Iaroslav Koshelev, Alexander Korotin

arXiv:2609.36099 · 2026-10-06 공개 · arXiv · PDF

diffusion-models flow-matching distillation cifar-10 mnist generation-quality inverse-distillation data-unlearning

Abstract

Multi-step matching models, including flow and diffusion models, produce high-quality outputs but incur substantial inference costs and may reproduce unwanted components of their training datasets. We introduce Inverse Distillation Unlearning (IDU), a unified framework that simultaneously distills a teacher multi-step matching model into an efficient one-step student generator and suppresses outputs corresponding to a designated training subset. We first formulate distillation as a min-max objective over a data distribution and then represent this distribution as a mixture of the forget-set and the generated distributions. This allows us to compare this mixture with the teacher's training distribution and recover only the retained data at the optimum. Our method requires only a pretrained full-data teacher and data from the forget set, without access to retained training examples, extra feature extractors or classifiers. Extensive experiments on MNIST and CIFAR-10 datasets under flow-matching and score-based diffusion settings demonstrate that IDU substantially reduces the generation frequency of forgotten classes while preserving generation quality on the retained classes. To the best of our knowledge, IDU is the first unified framework for simultaneous unlearning and distillation in unconditional flow-matching and score-based models.

한국어 요약

한 줄 요약

IDU는 다단계 매칭 모델을 효율적인 단계 생성자로 증류하면서 지정된 학습 집합을 억제하는 통합 프레임워크이다.

핵심 기여도

핵심 아이디어

IDU는 기존 다단계 매칭 모델(예: flow, diffusion)이 생성 과정에서 불필요한 학습 데이터를 재현하는 문제를 해결하기 위해, `unlearning`과 `distillation`을 동시에 수행하는 새로운 접근법을 제안한다. 이는 `min-max objective`를 사용하여 데이터 분포를 최적화함으로써 이루어진다. 구체적으로, `forget-set`과 생성된 분포의 혼합을 정의하고, 이를 `teacher`의 학습 분포와 비교하여 최적화한다. 이 과정에서 `forget-set`에 해당하는 데이터는 억제되고, 나머지 데이터만 복구된다. 이는 기존의 `data unlearning`과 `class unlearning` 접근법과 달리, `one-step generator`에 적용 가능한 통합적 해결책을 제공한다.

기술적 접근법

주요 결과

의의 및 한계

IDU는 기존 `unlearning`과 `distillation`을 결합한 첫 번째 통합 프레임워크로, `flow-matching`과 `score-based` 모델에서 효과적으로 작동함을 입증하였다. 특히, `one-step generator`에 적용 가능하며, 기존 학습 예제나 추가 분류기 없이도 작동한다는 점에서 실용적 가치가 높다. 그러나, `IDU`는 `unconditional models`에 초점을 맞추고 있으며, 조건부 모델(`conditional models`)에 대한 확장 가능성은 명시되지 않았다. 또한, `FGR` 감소폭은 데이터셋과 모델 종류에 따라 변동이 있을 수 있으므로, 보다 다양한 실험 환경에서의 검증이 필요하다.

실용적 활용

IDU는 개인 정보 보호, 저작권 문제, 부적절한 콘텐츠 재현을 방지하는 데 활용될 수 있다. 예를 들어, NSFW 이미지나 특정 얼굴을 포함한 데이터를 학습한 모델에서 해당 데이터를 효과적으로 억제할 수 있으며, 이는 콘텐츠 생성 AI, 이미지 생성 서비스, 디지털 마케팅 등 다양한 산업 분야에서 활용 가능하다.