ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang, Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang

arXiv:2607.24743 · 2026-07-28 공개 · arXiv · PDF

retrieval-augmented instruction-following multimodal-llm report-generation radiologist-evaluation region-of-interest vision-centric medical-images

Abstract

Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical understanding that systematically addresses these limitations. We propose a compositional and cascaded vision encoder architecture featuring a Cascade Spatial-Aware Locality Fusion operator that unifies diverse 2D and native 3D medical image understanding within a fused encoder. We further introduce a vision-grounded evaluation framework, including MedIF-Bench for instruction-following assessment and a region-of-interest-grounded method for clinically aligned and factualness-driven report generation evaluation. We show that ClinFusion sets a new state-of-the-art across a comprehensive suite of 2D and 3D multimodal medical benchmarks---spanning visual question answering, report generation, and instruction following---as well as textual medical tasks, outperforming leading open-source medical MLLMs (e.g., Hulu-Med, Lingshu) on 20 out of 24 benchmarks and demonstrating multimodal capabilities better than powerful proprietary models such as GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks, and can be further augmented with agentic tool use for retrieval-augmented and tool-assisted clinical workflows. A blinded evaluation by board-certified radiologists confirms that ClinFusion produces the highest-ranked reports, and validates our RoI-grounded metric as achieving the strongest correlation with expert judgment among all automatic evaluation metrics examined.

한국어 요약

한 줄 요약

ClinFusion은 2D 및 3D 의료 이미지를 통합 처리하는 Vision-Centric MLLM으로, 의료 분야에서 최고 성능을 보인다.

핵심 기여도

핵심 아이디어

의료 분야에서 MLLM을 성공적으로 활용하려면, 2D 및 3D 의료 이미지의 이질적인 시각 정보를 효과적으로 흡수하고, 보고서 생성과 같은 작업을 정확하고 사실 기반으로 평가하는 것이 필수적이다. ClinFusion은 이 두 가지 핵심 문제를 해결하기 위해 **Compositional Vision Encoder**를 설계하여 2D 및 3D 이미지를 통합 처리하며, **Cascade Spatial-Aware Locality Fusion** 연산자를 통해 공간 정보를 효과적으로 결합한다. 또한, 기존 평가 방식의 한계를 극복하기 위해 **MedIF-Bench**와 **RoI-grounded 평가**를 도입하여 보고서 생성 시 정확성과 사실성 중심의 평가를 가능하게 한다.

기술적 접근법

주요 결과

의의 및 한계

ClinFusion은 의료 분야에서 MLLM의 시각 정보 처리와 평가의 한계를 체계적으로 해결하며, 의료 진단 및 보고서 생성에 있어 높은 정확도와 사실성을 보인다. 특히, **RoI-grounded 평가**는 전문가 판단과 가장 강한 상관관계를 보여, 의료 AI 평가의 새로운 기준을 제시한다. 그러나 모델은 여전히 고정된 파라미터 기반으로 작동하며, 특정 분야에서는 전용 AI 모델보다 한계가 있을 수 있다.

실용적 활용

ClinFusion은 병원의 방사선과, 병리학, 진단 지원 시스템 등에서 의료 이미지 분석 및 보고서 생성에 활용 가능하다. 또한, **Agentic Tool Use**를 통해 검색 증강 및 도구 지원 클리니컬 워크플로우에 통합될 수 있어, 의료 AI의 실용성과 확장성을 높인다.