foundation-model transfer-learning privacy-preserving cross-dataset object-based spectral-artifact identity-independent replay-detection
Abstract
Face presentation attack detection (PAD) is traditionally formulated as a face-specific problem, although many of the visual artifacts introduced by print, replay, and recapture processes are not inherently tied to facial appearance. In this work, we investigate whether transferable PAD representations can be learned without using faces during downstream PAD training. To this end, we introduce TPO, a controlled face-free presentation attack dataset consisting of bona fide, print, and replay recordings of, almost randomly chosen, tomatoes, potatoes, and onions acquired under protocols that closely mirror conventional face PAD datasets. Using a foundation-model-based PAD architecture, we demonstrate that a detector trained on TPO achieves an average AUC of 92.70% across four standard cross-dataset face PAD benchmarks, outperforming training on synthetic faces and remaining competitive with models trained on real face datasets. Conversely, models trained on face PAD datasets transfer consistently above chance to TPO, suggesting that the learned representations capture characteristics of the presentation process rather than object semantics. Furthermore, incorporating TPO into conventional face PAD training consistently improves cross-dataset performance under fixed optimization budgets, indicating that face-free data provides complementary information rather than simply additional training samples. Finally, representation and frequency analyses provide further evidence that transferable PAD representations cannot be explained by a single spectral artifact but instead encode richer presentation cues shared across object categories. Together, these results provide empirical evidence that transferable presentation attack representations can be learned independently of facial content, opening new opportunities for privacy-preserving and identity-independent PAD development.
한국어 요약
한 줄 요약
토마토, 감자, 마늘을 사용한 무표현 얼굴 PAD 데이터셋 TPO를 통해 얼굴 정보 없이도 높은 성능의 PAD가 가능함을 실증.
핵심 기여도
- TPO: 무얼굴, 제어된 환경의 12,480개 샘플을 포함한 새로운 PAD 데이터셋 제시.
- TPO 기반 모델이 4개 얼굴 PAD 벤치마크에서 평균 AUC 92.7% 달성.
- 얼굴 기반 모델이 TPO에서 평균 AUC 78.4% 달성, 무얼굴 데이터가 공통된 공격 특성을 학습함을 입증.
- TPO와 얼굴 데이터를 혼합 학습할 경우 성능 향상, 무얼굴 데이터가 보완 정보를 제공함을 밝힘.
핵심 아이디어
전통적으로 얼굴 PAD는 얼굴 데이터에만 의존하며, 공격 신호가 얼굴 자체에서 비롯된다고 가정해왔다. 그러나 본 연구는 공격 신호가 **재촬영 파이프라인**(recapture pipeline)에서 비롯된다는 가설을 제시한다. 이는 디스플레이의 무어 패턴, 인쇄물의 반점 구조, 재촬영 시 생기는 감마 왜곡 등과 같은 **공격 기구**(PAI)의 특성이다. 따라서 얼굴이 아닌 다른 객체(예: 채소)를 사용해도 공격 신호를 학습할 수 있다고 주장한다. 이를 검증하기 위해 TPO 데이터셋을 구성하고, 얼굴 없이도 높은 성능을 내는지 실험적으로 검증했다.
기술적 접근법
- **TPO 데이터셋**: 채소(토마토, 감자, 마늘)의 정품 및 인쇄/재생 공격 샘플 12,480개.
- **모델 아키텍처**: Foundation 모델 기반 PAD 시스템, CLIP 사전학습 모델 사용.
- **학습 프로토콜**: TPO 기반 모델을 4개 얼굴 PAD 벤치마크(예: MSU-MFSD, OULU-NPU 등)에 cross-dataset 테스트.
- **비교 실험**: TPO → 얼굴, 얼굴 → TPO, 혼합 학습 실험 수행.
- **분석 방법**: 임베딩 분석, 주파수 잔차 분석 등을 통해 공격 신호의 공통성을 검증.
주요 결과
- TPO 기반 모델이 4개 얼굴 PAD 벤치마크에서 평균 AUC 92.7% 달성.
- 합성 얼굴 기반 모델 대비 1~2% 이상 우수.
- 얼굴 기반 모델이 TPO에서 평균 AUC 78.4% 달성, zero-shot 성능(67.7%)보다 높음.
- TPO와 얼굴 데이터를 혼합 학습할 경우, 동일 최적화 예산 하에서 성능 향상.
- 주파수 분석 결과, 단일 스펙트럼 특징이 아닌, 다양한 공격 힌트를 학습함을 밝힘.
의의 및 한계
본 연구는 얼굴 정보 없이도 PAD가 가능하다는 사실을 실증적으로 입증함으로써, **개인 정보 보호 및 ID 독립적 PAD 개발**의 가능성을 열었다. TPO 데이터셋은 얼굴 없이도 공격 신호를 학습할 수 있음을 보여주는 실험적 기반을 제공하며, 기존 얼굴 중심 데이터 수집 방식의 한계를 드러낸다. 그러나 본 연구는 **3D 마스크나 적외선/깊이 센서 기반 공격**에는 적용되지 않을 수 있으며, CLIP 모델이 얼굴 기반 사전학습을 포함하고 있어, 무얼굴 학습에 필요한 최소 표현 사전(minimal representation prior)에 대한 연구는 여전히 필요하다.
실용적 활용
본 연구는 **개인 정보 보호가 중요한 환경**(예: 공공 장소, 금융 인증)에서 얼굴 데이터 없이도 신뢰할 수 있는 PAD를 구현할 수 있음을 시사한다. 또한, **다양한 객체에 대한 공격 신호를 학습**함으로써, 더 넓은 범위의 공격 탐지가 가능해질 수 있다. TPO 데이터셋은 향후 무얼굴 중심의 PAD 연구 및 개발에 활용될 수 있다.