RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation

arXiv:2609.29028 · 2026-09-27 공개 · arXiv · PDF

multimodal dataset semantic-segmentation high-fidelity annotation rgb-d large-scale fine-grained

Abstract

In this paper, we propose RGBD20K, a novel dataset for facilitating the development of more robust and general RGB-D semantic segmentation by encompassing abundant categories and high-quality annotations. RGBD20K possesses several attractive properties: (1) Expanded Semantic Space. In particular, it covers 160 fine-grained categories, largely surpassing the category diversity of existing popular RGB-D benchmarks (e.g., NYUv2 with 40 classes and SUN RGB-D with 37 classes). With such enriched semantic coverage, we expect to promote the learning of more generalizable segmentation models. (2) Larger Scale. Compared with current benchmarks, RGBD20K offers 20,000 RGB-D image pairs, providing a substantially larger training resource that benefits the development of more powerful deep models. (3) High-Fidelity Annotation. We perform rigorous re-evaluation and correction of existing labels to resolve long-standing annotation noise, resulting in a clean and reliable ground-truth foundation. Furthermore, we propose a novel score-purified fusion (SPF) method, which achieves state-of-the-art performance across all evaluated benchmarks, demonstrating the effectiveness of our approach in leveraging high-quality multimodal information for RGB-D semantic segmentation. The dataset is here: https://github.com/ShaohuaDong2021/RGBD20K/.

한국어 요약

한 줄 요약

RGBD20K는 160개의 세분화된 카테고리와 20,000개의 RGB-D 이미지 쌍을 포함한 대규모 높은 품질의 RGB-D 세분화 벤치마크 데이터셋이다.

핵심 기여도

핵심 아이디어

기존 RGB-D 세분화 데이터셋은 카테고리 수가 제한적(40~37개)이며, 어노테이션 품질이 낮아 모델의 일반화 능력 향상에 한계가 있었다. 본 연구는 이를 해결하기 위해 160개의 세분화된 카테고리를 포함한 RGBD20K 데이터셋을 제안한다. 이는 기존 데이터셋 대비 4배 이상의 세분화된 세멘틱 공간을 제공하여, 더 일반적인 세분화 모델 학습을 촉진한다. 또한, 20,000개의 RGB-D 이미지 쌍을 통해 더 강력한 딥러닝 모델 개발에 기여한다.

또한, 기존 라벨의 오류를 철저히 재검토하고 SPF라는 새로운 퓨전 방법을 제안하여, 높은 품질의 멀티모달 정보를 효과적으로 활용하는 방식을 제시한다.

기술적 접근법

주요 결과

의의 및 한계

RGBD20K는 RGB-D 세분화 분야에서의 모델 일반화 능력 향상과 더 강력한 딥러닝 모델 개발에 기여할 수 있는 중요한 기초 자료이다. 특히, 160개의 카테고리와 20,000개의 이미지 쌍은 기존 데이터셋의 한계를 극복하며, SPF 방법은 멀티모달 정보 활용의 효과성을 입증한다.

한편, 본 연구는 SPF 방법의 구체적인 알고리즘 디테일이나 다른 퓨전 방법과의 비교 분석을 명시하지 않았으며, 특정 도메인에서의 성능 평가도 제시되지 않았다.

실용적 활용

RGBD20K는 로봇 비전, 자율 주행, 스마트 홈 등에서 활용되는 RGB-D 세분화 모델 개발에 유용한 자료가 될 수 있다. 특히, 다양한 객체 카테고리와 높은 품질의 어노테이션은 실제 환경에서의 정확한 객체 인식과 상호작용을 가능하게 한다.