ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang, Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang
arXiv:2607.24743 · 2026-07-28 공개 · arXiv · PDF
retrieval-augmented instruction-following multimodal-llm report-generation radiologist-evaluation region-of-interest vision-centric medical-images
Abstract
Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical understanding that systematically addresses these limitations. We propose a compositional and cascaded vision encoder architecture featuring a Cascade Spatial-Aware Locality Fusion operator that unifies diverse 2D and native 3D medical image understanding within a fused encoder. We further introduce a vision-grounded evaluation framework, including MedIF-Bench for instruction-following assessment and a region-of-interest-grounded method for clinically aligned and factualness-driven report generation evaluation. We show that ClinFusion sets a new state-of-the-art across a comprehensive suite of 2D and 3D multimodal medical benchmarks---spanning visual question answering, report generation, and instruction following---as well as textual medical tasks, outperforming leading open-source medical MLLMs (e.g., Hulu-Med, Lingshu) on 20 out of 24 benchmarks and demonstrating multimodal capabilities better than powerful proprietary models such as GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks, and can be further augmented with agentic tool use for retrieval-augmented and tool-assisted clinical workflows. A blinded evaluation by board-certified radiologists confirms that ClinFusion produces the highest-ranked reports, and validates our RoI-grounded metric as achieving the strongest correlation with expert judgment among all automatic evaluation metrics examined.
한국어 요약
한 줄 요약
ClinFusion은 2D 및 3D 의료 이미지를 통합 처리하는 Vision-Centric MLLM으로, 의료 분야에서 최고 성능을 보인다.
핵심 기여도
- **Cascade Spatial-Aware Locality Fusion** 연산자를 도입하여 2D 및 3D 의료 이미지를 통합 처리.
- **MedIF-Bench**와 **RoI-grounded 평가 방법**을 제안하여 의료 보고서 생성 평가의 정확도를 높임.
- 24개 의료 MLLM 벤치마크 중 20개에서 기존 오픈소스 모델(Hulu-Med, Lingshu)을, 16개 중 13개에서 GPT-5.2, Gemini-3-Flash를 초과.
- 보드 인증 방사선과 전문의의 맹목 평가에서 최고 보고서를 생성함이 확인됨.
핵심 아이디어
의료 분야에서 MLLM을 성공적으로 활용하려면, 2D 및 3D 의료 이미지의 이질적인 시각 정보를 효과적으로 흡수하고, 보고서 생성과 같은 작업을 정확하고 사실 기반으로 평가하는 것이 필수적이다. ClinFusion은 이 두 가지 핵심 문제를 해결하기 위해 **Compositional Vision Encoder**를 설계하여 2D 및 3D 이미지를 통합 처리하며, **Cascade Spatial-Aware Locality Fusion** 연산자를 통해 공간 정보를 효과적으로 결합한다. 또한, 기존 평가 방식의 한계를 극복하기 위해 **MedIF-Bench**와 **RoI-grounded 평가**를 도입하여 보고서 생성 시 정확성과 사실성 중심의 평가를 가능하게 한다.
기술적 접근법
- **Vision Encoder**: 2D 및 3D 의료 이미지를 통합 처리하는 **Compositional and Cascaded Vision Encoder**를 사용.
- **Fusion Operator**: **Cascade Spatial-Aware Locality Fusion**을 통해 2D 및 3D 이미지의 공간 정보를 결합.
- **평가 프레임워크**:
- **MedIF-Bench**: 복잡한 의료 관련 지시사항을 따르는 능력을 평가.
- **RoI-grounded 평가**: 보고서 생성 시 특정 해부학적 영역을 기반으로 정확도를 평가.
- **이미지 처리**: 3D 볼륨 입력 시 기본적으로 32개의 256×256 크기의 슬라이스를 샘플링.
주요 결과
- **2D 의료 VQA 벤치마크**: ClinFusion-8B가 8개 벤치마크 중 7개에서 1위 또는 2위를 기록.
- **의료 보고서 생성**: CheXpert-Plus에서 37.8 F1 점수 (Hulu-Med-7B 대비 +5.9), IU-XRAY에서 57.3 F1 점수 (Hulu-Med-7B 대비 +10.8).
- **의료 지시사항 따르기 (MedIF-Bench)**: 전체 지시사항 준수 점수(Overall-IF)에서 뛰어난 성능.
- **전문가 평가**: 보드 인증 방사선과 전문의에 의해 ClinFusion이 생성한 보고서가 가장 높은 순위를 기록.
의의 및 한계
ClinFusion은 의료 분야에서 MLLM의 시각 정보 처리와 평가의 한계를 체계적으로 해결하며, 의료 진단 및 보고서 생성에 있어 높은 정확도와 사실성을 보인다. 특히, **RoI-grounded 평가**는 전문가 판단과 가장 강한 상관관계를 보여, 의료 AI 평가의 새로운 기준을 제시한다. 그러나 모델은 여전히 고정된 파라미터 기반으로 작동하며, 특정 분야에서는 전용 AI 모델보다 한계가 있을 수 있다.
실용적 활용
ClinFusion은 병원의 방사선과, 병리학, 진단 지원 시스템 등에서 의료 이미지 분석 및 보고서 생성에 활용 가능하다. 또한, **Agentic Tool Use**를 통해 검색 증강 및 도구 지원 클리니컬 워크플로우에 통합될 수 있어, 의료 AI의 실용성과 확장성을 높인다.