LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model

Dilxat Muhtar, Zhenshi Li, Feng Gu, Xue-liang Zhang, P. Xiao

arXiv:2402.02544 · 2026-07-27 공개 · arXiv · PDF

vision-language large-language-models curriculum-learning remote-sensing multimodal-language-models volunteered-geographic-information image-text-dataset rs-image-understanding

Abstract

The revolutionary capabilities of large language models (LLMs) have paved the way for multimodal large language models (MLLMs) and fostered diverse applications across various specialized domains. In the remote sensing (RS) field, however, the diverse geographical landscapes and varied objects in RS imagery are not adequately considered in recent MLLM endeavors. To bridge this gap, we construct a large-scale RS image-text dataset, LHRS-Align, and an informative RS-specific instruction dataset, LHRS-Instruct, leveraging the extensive volunteered geographic information (VGI) and globally available RS images. Building on this foundation, we introduce LHRS-Bot, an MLLM tailored for RS image understanding through a novel multi-level vision-language alignment strategy and a curriculum learning method. Additionally, we introduce LHRS-Bench, a benchmark for thoroughly evaluating MLLMs' abilities in RS image understanding. Comprehensive experiments demonstrate that LHRS-Bot exhibits a profound understanding of RS images and the ability to perform nuanced reasoning within the RS domain.

한국어 요약

한 줄 요약

LHRS-Bot은 VGI와 RS 이미지를 기반으로 한 대규모 RS 전용 MLLM으로, LHRS-Align과 LHRS-Instruct 데이터셋을 통해 뛰어난 RS 이미지 이해 능력을 보인다.

핵심 기여도

핵심 아이디어

기존 MLLM은 RS 이미지의 다양한 지형과 객체를 충분히 고려하지 못한다는 문제를 해결하기 위해, LHRS-Bot은 VGI와 전 세계 RS 이미지를 기반으로 데이터셋을 구축하고, 다단계 시각-언어 정렬 전략을 도입했다. LHRS-Bot은 CLIP의 ViT-L/14 시각 인코더를 사용하며, 다양한 레이어의 히든 피처를 유지해 더 풍부한 이미지 표현을 가능하게 한다. 또한, 이미지 특징의 중복성이 네트워크 깊이에 따라 증가한다는 관찰을 바탕으로, 더 얕은 레벨에 더 많은 쿼리를 할당하는 'descending query allocation' 전략을 제안한다. 이는 시각 정보의 효율적 요약과 언어 모델과의 정렬을 동시에 달성한다.

기술적 접근법

주요 결과

의의 및 한계

LHRS-Bot은 RS 이미지 이해를 위한 체계적인 데이터셋과 평가 벤치마크를 제공하며, 기존 MLLM이 고려하지 못한 다단계 시각-언어 정렬 전략을 도입한 점에서 학술적 의의가 크다. 또한, VGI를 활용한 데이터셋 구축 방식은 RS 분야의 대규모 학습 자료 확보에 기여할 수 있다. 그러나, LLM 일반적인 한계인 환상(hallucination)에 취약하며, RS 특화 MLLM의 발전을 위해서는 더 높은 품질의 데이터셋과 훈련 전략이 필요하다.

실용적 활용

LHRS-Bot은 지구 관측, 도시 계획, 재해 대응 등 다양한 RS 응용 분야에서 활용 가능하다. 특히, GPT-4와 유사한 성능을 보이는 점에서, 전문가가 아닌 일반 사용자도 RS 이미지 해석에 접근할 수 있는 인터페이스로 활용될 수 있다.