denoising image-restoration low-light-enhancement all-in-one dehazing deraining frequency-mining adaptive-network
Abstract
In the image acquisition process, various forms of degradation, including noise, haze, and rain, are frequently introduced. These degradations typically arise from the inherent limitations of cameras or unfavorable ambient conditions. To recover clean images from degraded versions, numerous specialized restoration methods have been developed, each targeting a specific type of degradation. Recently, all-in-one algorithms have garnered significant attention by addressing different types of degradations within a single model without requiring prior information of the input degradation type. However, these methods purely operate in the spatial domain and do not delve into the distinct frequency variations inherent to different degradation types. To address this gap, we propose an adaptive all-in-one image restoration network based on frequency mining and modulation. Our approach is motivated by the observation that different degradation types impact the image content on different frequency subbands, thereby requiring different treatments for each restoration task. Specifically, we first mine low- and high-frequency information from the input features, guided by the adaptively decoupled spectra of the degraded image. The extracted features are then modulated by a bidirectional operator to facilitate interactions between different frequency components. Finally, the modulated features are merged into the original input for a progressively guided restoration. With this approach, the model achieves adaptive reconstruction by accentuating the informative frequency subbands according to different input degradations. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance on different image restoration tasks, including denoising, dehazing, deraining, motion deblurring, and low-light image enhancement. Our code is available at https://github.com/c-yn/AdaIR.
한국어 요약
한 줄 요약
AdaIR은 주파수 분해와 조절을 통해 다양한 이미지 복원 작업에서 최신 성능을 달성한 적응형 통합 복원 모델이다.
핵심 기여도
- **AFLB (Adaptive Frequency Learning Block)**를 도입하여 주파수 영역 정보를 활용한 복원 구조를 제안.
- **FMiM (Frequency Mining Module)**과 **FMoM (Frequency Modulation Module)**를 통해 저주파 및 고주파 성분을 추출하고 조절.
- **All-in-One 설정**에서 단일 모델로 denoising, dehazing, deraining, deblurring, low-light enhancement 등 다중 작업에서 SOTA 성능 달성.
- **30.52 dB PSNR** 성능을 기록하며, 기존 방법 대비 **3.03 dB 개선**.
핵심 아이디어
기존의 통합 이미지 복원 방법은 공간 영역에서만 작동하여, 다양한 종류의 퇴화가 서로 다른 주파수 대역에 영향을 미친다는 사실을 고려하지 못했다. AdaIR은 이 점을 해결하기 위해, 입력 이미지의 주파수 스펙트럼을 적응적으로 분해하여, 저주파와 고주파 정보를 추출하고, 이를 **FMiM**을 통해 가이드받아 복원 과정에 활용한다. 추출된 주파수 성분은 **FMoM**을 통해 상호작용하며 조절되어, 각 퇴화 유형에 맞는 복원이 가능하도록 한다. 이는 기존의 고정된 주파수 분해 방식과 달리, 입력에 따라 유연하게 조절되는 주파수 학습 기법이다.
기술적 접근법
- **AFLB (Adaptive Frequency Learning Block)**:
- **FMiM (Frequency Mining Module)**: 입력 이미지의 주파수 스펙트럼을 분해하여 저주파와 고주파 성분을 추출.
- **FMoM (Frequency Modulation Module)**: 추출된 주파수 성분 간 상호작용을 촉진하여 복원 성능 향상.
- **4-level encoder-decoder 구조**: 각 레벨에 **Transformer Block (TB)**을 사용하며, 데코더 사이에 AFLB를 삽입.
- **하이퍼파라미터**: Adam 최적화기 (β1=0.9, β2=0.999), 학습률 2e-4, 150 에포크, 배치 크기 32 (All-in-One), 8 (Single-task).
- **입력 크기**: 128×128 픽셀 크롭 패치.
주요 결과
- **All-in-One 설정**에서 **SOTS, Rain100L, BSD400, WED** 등 다중 데이터셋에서 SOTA 성능 달성.
- **Image Dehazing** 작업에서 **30.52 dB PSNR** 기록 (기존 방법 대비 **+3.03 dB**).
- **FMiM** 도입으로 기존 기준 대비 **+1.58 dB PSNR** 향상.
- **FMiM + FMoM** 결합 시 **+3.03 dB PSNR** 개선.
- **AFLB**는 2.64M 파라미터, 6.21 GFlops의 계산량 증가를 동반.
의의 및 한계
- **의의**: 단일 모델로 다양한 퇴화 유형을 처리할 수 있는 **All-in-One 복원**을 가능하게 하며, 주파수 영역 정보를 활용한 새로운 접근법을 제시.
- **한계**: AFLB는 계산량 증가 (2.64M 파라미터, 6.21 GFlops)를 동반하며, 고성능 하드웨어가 필요. 또한, 주파수 분해는 훈련 데이터의 품질에 민감할 수 있음.
- **추가 실험**: **Average pooling**과 **Gaussian filtering** 대비 **31.24 dB PSNR** 기록하며, 마스크 기반 주파수 분해가 가장 효과적임을 입증.
실용적 활용
- **보안 카메라, 의료 영상, 원격 감시** 등 다양한 분야에서 단일 모델로 퇴화 이미지를 복원할 수 있어, 모델 관리 및 배포 비용을 절감.
- **엣지 기기**에서의 실시간 이미지 복원에도 적용 가능하나, 계산량 증가에 따라 하드웨어 최적화가 필요.
- **영상 편집, 드론 영상 처리, 모바일 카메라** 등에서 실용적 활용 가능.