Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current MLLM agents can turn local urban perception into reliable action in a complicated real-scale city. We propose UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data. UrbanGround supports closed-loop interaction from a first-person view and provides an interactive map for navigation. Agents can directly enter the 3D city and explore from a first-person view. Our analysis follows the growth of the spatial problem through three research questions. We first test whether an agent can ground a local scene well enough to answer spatial questions after active observation. Then we ask whether that grounding supports navigation as destinations become farther away and less explicit. Finally, we examine whether the resulting behavior survives changes in route availability and pedestrian motion. Contemporary MLLM agents usually show useful atomic abilities in visual recognition and short-range spatial reasoning, while orientation and pedestrian-aware movement remain unreliable. Their central failure emerges over extended exploration, where local abilities do not compose into sustained goal-directed behavior and errors accumulate without effective correction. We hope UrbanGround will support broader study of how far current MLLM agents can explore reliably in complex, open-ended urban environments.
한 줄 요약
UrbanGround는 홍콩의 실제 3D 지리 데이터를 기반으로 MLLM 에이전트의 도시 내 지속적 탐색 능력을 평가하는 첫 번째 샌드박스 환경이다.
핵심 기여도
- UrbanGround는 홍콩의 3D 시각화 지도와 3D 보행자 네트워크 데이터를 기반으로 구축된 첫 번째 실규모 도시 샌드박스 환경이다.
- MLLM 에이전트가 단일 장면 인식에서부터 장거리 탐색, 환경 변화 대응까지의 능력을 평가하는 3가지 연구 질문을 제시한다.
- 현재 MLLM은 단거리 시공간 추론은 가능하지만, 장거리 탐색 시 오류 누적이 발생하며, 보행자와 환경 변화에 대한 적응력이 부족하다는 것을 밝힘.
핵심 아이디어
도시 내에서 MLLM 에이전트가 지속적으로 탐색하고 목표를 달성하려면 단순한 시각 인식을 넘어, 지속적인 공간 추정과 경로 수정 능력이 필요하다. UrbanGround는 홍콩의 실제 지형과 보행자 네트워크를 반영한 3D 환경을 Unity에 스트리밍하여, 에이전트가 첫인상 시점에서의 관찰을 바탕으로 지속적인 탐색을 수행하도록 설계되었다. 이는 단일 장면 인식에서 도시 전체의 공간적 이해로의 전환을 연구하는 데 중요한 기반이 된다.
기술적 접근법
- UrbanGround는 홍콩 정부의 Lands Department에서 제공한 3D Visualisation Map과 3D Pedestrian Network 데이터를 기반으로 구축됨.
- Unity 엔진을 사용하여 지리적으로 등록된 도시를 스트리밍하며, 물리적 충돌 처리와 지리 좌표 기반의 경로 기록을 지원.
- 조명, 날씨 조건, 도로 폐쇄, 보행자 이동 등의 변화를 시뮬레이션하여 동적 환경에서의 탐색을 테스트.
- 에이전트는 첫인상 시점의 관찰, 물리적 제어, 대화형 지도를 통해 도시를 탐색.
주요 결과
- MLLM은 단거리 시공간 추론에서 유용한 능력을 보이지만, 장거리 탐색 시 오류 누적이 발생하며, 지속적인 탐색 능력이 부족함.
- 환경 변화(예: 도로 폐쇄, 보행자 이동)에 대한 적응력이 낮아, 실시간 업데이트와 제약 준수에 실패함.
- 홍콩의 실제 지형과 보행자 네트워크를 반영한 환경에서, 에이전트는 단일 장면 인식에서 도시 전체 탐색으로의 전환에 실패함.
의의 및 한계
UrbanGround는 MLLM 에이전트의 도시 내 지속적 탐색 능력을 평가하는 데 중요한 기반을 제공하며, 단일 장면 인식에서 도시 전체의 공간적 이해로의 전환을 연구할 수 있는 첫 번째 샌드박스 환경이다. 그러나 현재 MLLM은 장거리 탐색 시 오류 누적이 발생하며, 환경 변화에 대한 적응력이 부족하다는 한계를 드러낸다. 이는 MLLM이 단순한 인식 능력을 넘어, 지속적인 공간 추정과 경로 수정 능력을 갖추어야 한다는 점을 시사한다.
실용적 활용
UrbanGround는 도시 내 자율 이동 시스템, 스마트 도시 관리, 로봇 탐색 연구 등에 활용될 수 있다. 특히, MLLM 기반 에이전트가 실제 도시 환경에서 지속적으로 탐색하고 적응하는 능력을 향상시키는 데 기초적인 연구 도구로 활용 가능하다.