MACE: Mass Concept Erasure in Diffusion Models

Shilin Lu, Zilan Wang, Leyang Li, Yanzhu Liu, A. Kong

arXiv:2403.06135 · 2026-07-27 공개 · arXiv · PDF

diffusion-models text-to-image cross-attention lora-finetuning concept-erasure object-erasure celebrity-erasure explicit-content

Abstract

The rapid expansion of large-scale text-to-image diffusion models has raised growing concerns regarding their potential misuse in creating harmful or misleading content. In this paper, we introduce MACE, a finetuning framework for the task of MAss Concept Erasure. This task aims to prevent models from generating images that embody unwanted concepts when prompted. Existing concept erasure methods are typically restricted to handling fewer than five concepts simultaneously and struggle to find a balance between erasing concept synonyms (generality) and maintaining unrelated concepts (specificity). In contrast, MACE differs by successfully scaling the erasure scope up to 100 concepts and by achieving an effective balance between generality and specificity. This is achieved by leveraging closed-form cross-attention refinement along with LoRA finetuning, collectively eliminating the information of undesirable concepts. Furthermore, MACE integrates multiple LoRAs without mutual interference. We conduct extensive evaluations of MACE against prior methods across four different tasks: object erasure, celebrity erasure, explicit content erasure, and artistic style erasure. Our results reveal that MACE surpasses prior methods in all evaluated tasks. Code is available at https://github.com/Shilin-LU/MACE.

한국어 요약

한 줄 요약

MACE는 100개 이상의 개념을 확장적으로 제거하면서 일반성과 특이성을 균형 있게 유지하는 텍스트-이미지 디퓨전 모델 개념 제거 프레임워크이다.

핵심 기여도

핵심 아이디어

기존의 개념 제거 방법은 특정 단어나 문장에만 집중하거나, 초기 디노이징 단계에서 전체적인 맥락을 손상시키는 경향이 있었다. MACE는 이러한 문제를 해결하기 위해, 개념이 잔존 정보를 다른 단어에 숨기는 방식을 고려하여 cross-attention 계층을 closed-form 방식으로 정제한다. 이는 단순히 특정 단어를 제거하는 것이 아니라, 해당 개념이 문맥 내에서 표현되는 방식 자체를 제거하는 데 기여한다. 또한, 각 개념에 대해 독립적인 LoRA 모듈을 학습하고, 이를 통합할 때 상호 간섭을 방지하는 loss 함수를 도입하여, 100개 개념을 동시에 제거하면서도 다른 개념에 영향을 주지 않도록 설계되었다.

기술적 접근법

주요 결과

의의 및 한계

MACE는 대규모 개념 제거를 가능하게 하며, 기존 방법이 해결하지 못했던 일반성과 특이성의 균형 문제를 해결함으로써 텍스트-이미지 모델의 안전성 향상에 기여한다. 특히, closed-form 기반의 접근법은 계산 효율성과 모델 성능을 동시에 확보할 수 있다는 장점이 있다. 그러나 MACE는 훈련 데이터 없이 개념을 제거하는 방식이므로, 제거 대상 개념이 모델에 포함되어 있지 않은 경우 효과가 제한될 수 있다. 또한, 제거된 개념이 문맥에서 완전히 사라지지 않고 다른 표현 방식으로 나타날 가능성도 존재한다.

실용적 활용

MACE는 디퓨전 모델이 생성하는 콘텐츠의 부정적 영향을 줄이기 위해 사용될 수 있으며, 특히 저작권 보호, 성인 콘텐츠 필터링, 인물 이미지 제거 등에 활용 가능하다. 또한, 모델 공급자가 새로운 버전의 모델을 출시할 때 기존 모델의 부정적 콘텐츠 생성 경향을 제거하는 데 유용할 수 있다.