Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion

Xunpeng Yi, Han Xu, Hao Zhang, Linfeng Tang, Jiayi Ma

arXiv:2403.16387 · 2026-07-27 공개 · arXiv · PDF

text-guided multi-modal-fusion image-fusion semantic-encoder infrared-visible degradation-aware interactive-fusion sota-methods

Abstract

Image fusion aims to combine information from different source images to create a comprehensively representative image. Existing fusion methods are typically helpless in dealing with degradations in low-quality source images and non-interactive to multiple subjective and objective needs. To solve them, we introduce a novel approach that leverages semantic text guidance image fusion model for degradation-aware and interactive image fusion task, termed as Text-IF. It innovatively extends the classical image fusion to the text guided image fusion along with the ability to harmoniously address the degradation and interaction issues during fusion. Through the text semantic encoder and semantic interaction fusion decoder, Text-IF is accessible to the all-in-one infrared and visible image degradation-aware processing and the interactive flexible fusion outcomes. In this way, Text-IF achieves not only multi-modal image fusion, but also multi-modal information fusion. Extensive experiments prove that our proposed text guided image fusion strategy has obvious advantages over SOTA methods in the image fusion performance and degradation treatment. The code is available at https://github.com/XunpengYi/Text-IF.

한국어 요약

한 줄 요약

Text-IF는 텍스트 유도를 활용해 저품질 이미지 복원과 상호작용 기반 퓨전을 통합한 새로운 이미지 퓨전 모델이다.

핵심 기여도

핵심 아이디어

기존 이미지 퓨전 방법은 저품질 이미지 처리와 사용자 상호작용에 제한적이었다. Text-IF는 텍스트 유도를 통해 이러한 문제를 해결한다. 텍스트 세마틱 인코더는 사전 학습된 비전-언어 모델을 활용해 텍스트의 의미적 정보를 추출하고, 세마틱 인터랙션 퓨전 디코더는 텍스트와 이미지 퓨전 특성을 결합하여 사용자의 요구에 맞는 퓨전 결과를 생성한다. 이는 텍스트와 이미지의 다중 모달 정보 퓨전을 가능하게 하며, 유연한 퓨전 결과를 제공한다.

기술적 접근법

주요 결과

의의 및 한계

Text-IF는 텍스트 유도를 통해 사용자 요구에 맞는 유연한 퓨전 결과를 제공하며, 기존 방법의 제한적인 퓨전과 상호작용 문제를 해결한다. 특히, 텍스트와 이미지의 다중 모달 정보 퓨전을 가능하게 하여, 실용적이고 이론적으로 중요한 기여를 한다. 그러나 텍스트 입력의 질에 따라 퓨전 결과가 달라질 수 있으며, 텍스트 해석 오류 시 퓨전 품질 저하 가능성은 남아 있다.

실용적 활용

Text-IF는 보안 감시, 군사 탐지, 야간 관찰 등 다양한 분야에서 사용자의 텍스트 입력을 기반으로 유연한 퓨전 이미지를 생성할 수 있다. 특히, 전문 지식 없이도 사용자가 원하는 퓨전 결과를 생성할 수 있어, 비전문가 사용자에게도 실용적이다.