SentZero: An Enhanced Sentence-Centric Vision-Language Pretraining for Multi-Task Zero-Shot Chest X-Ray Analysis

Hangyul Yoon, Hyungyung Lee, Edward Choi, Eunho Yang

arXiv:2609.34479 · 2026-09-29 공개 · arXiv · PDF

vision-language zero-shot contrastive-learning pretraining llm-based multi-task chest-x-ray radiology-reports

Abstract

Vision-language (VL) pretraining using paired chest X-ray (CXR) images and radiology reports has shown strong potential for medical image understanding. However, existing methods often remain dependent on task-specific finetuning because radiology reports are lengthy, clinically dense, and difficult to align with simple zero-shot prompts. Recent sentence-level approaches partially address this limitation using clinical phrases extracted by large language models (LLMs), but they largely overlook the intrinsic characteristics of radiology discourse. In particular, limited positive-pair diversity constrains further gains, while clinically equivalent sentences frequently recur across patients, creating false negatives in contrastive learning. To address these issues, we propose SentZero, an enhanced sentence-centric VL pretraining framework for zero-shot, multi-task CXR analysis. SentZero introduces LLM-based abstract-level sentence structuring and mapping to expand positive-pair diversity, together with an additional loss term to mitigate false negatives. We further introduce sentence-conditioned residual modulation of visual embeddings, enabling visual features to adapt to the semantic characteristics of each input sentence. Across diverse downstream tasks and datasets, SentZero improves zero-shot generalization and outperforms prior multi-task zero-shot methods.

한국어 요약

한 줄 요약

SentZero는 병원성 문장 구조를 활용한 새로운 시각-언어 사전 학습 프레임워크로, 병원 흉부 X선 분석에서 제로샷 성능을 향상시킨다.

핵심 기여도

핵심 아이디어

기존 시각-언어 사전 학습은 병원 보고서의 복잡성과 반복성을 고려하지 못해 제로샷 성능에 한계가 있었다. SentZero는 병원 보고서의 문장 구조와 의미적 중복성을 활용하여 구조화된 감독 정보를 생성한다. 예를 들어, "There is mild opacity in the bilateral lung base"와 같은 구체적 문장을 "There is opacity"와 같은 추상적 문장으로 매핑함으로써, 동일한 임상 개념을 가진 문장 간의 일관성을 강화한다. 또한, 반복되는 문장이 거짓 음성으로 간주되는 문제를 보완하기 위해 패치 수준의 선택적 끌어당김(selective patch-level attraction)을 도입한 새로운 손실 함수를 제안한다. 이는 기존의 쌍 재라벨링 방식과 달리, 전체 대비 학습 목표를 유지하면서도 정확도를 향상시킨다.

기술적 접근법

주요 결과

의의 및 한계

SentZero는 병원 보고서의 반복성과 의미 구조를 활용한 새로운 제로샷 접근법을 제시하며, 기존 대비 더 높은 일반화 능력을 보여준다. 특히, 반복되는 문장이 학습에 부정적 영향을 줄 수 있는 문제를 해결함으로써, 의료 영상 분석에서의 대비 학습을 더욱 안정적으로 수행할 수 있다. 그러나, 본 연구는 특정 병원성 문장 집합에만 적용되었으며, 다른 의료 영상 분야로 확장 가능성은 추가 연구가 필요하다. 또한, LLM 기반 추출 과정에서 발생할 수 있는 오류나 편향도 한계로 지적될 수 있다.

실용적 활용

SentZero는 병원 흉부 X선 분석에서 제로샷 분류, 공간 정착, 이상 탐지 등 다양한 다중 태스크에 적용 가능하다. 특히, 병원 보고서를 기반으로 한 자동 진단 시스템 개발, 의료 이미지 해석 보조 도구, 또는 대규모 비정형 데이터 처리에 유용하게 활용될 수 있다.