UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City

Tianjie Ju, Zheng Wu, Yueqing Sun, Yuhan Cui, Bobo Li, Shengqiong Wu, Pengzhou Cheng, Haodong Zhao, Zongru Wu, Xinbei Ma, Doris Zhang, Kunling Li, Mong-Li Lee, Wynne Hsu, Hao Fei, Qi Gu, Gongshen Liu, Zhuosheng Zhang

arXiv:2608.27456 · 2026-08-28 공개 · arXiv · PDF

spatial-reasoning multimodal-llm navigation agent-orientation pedestrian-awareness closed-loop-interaction local-perception urban-agency

Abstract

Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current MLLM agents can turn local urban perception into reliable action in a complicated real-scale city. We propose UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data. UrbanGround supports closed-loop interaction from a first-person view and provides an interactive map for navigation. Agents can directly enter the 3D city and explore from a first-person view. Our analysis follows the growth of the spatial problem through three research questions. We first test whether an agent can ground a local scene well enough to answer spatial questions after active observation. Then we ask whether that grounding supports navigation as destinations become farther away and less explicit. Finally, we examine whether the resulting behavior survives changes in route availability and pedestrian motion. Contemporary MLLM agents usually show useful atomic abilities in visual recognition and short-range spatial reasoning, while orientation and pedestrian-aware movement remain unreliable. Their central failure emerges over extended exploration, where local abilities do not compose into sustained goal-directed behavior and errors accumulate without effective correction. We hope UrbanGround will support broader study of how far current MLLM agents can explore reliably in complex, open-ended urban environments.

한국어 요약

한 줄 요약

UrbanGround는 홍콩의 실제 3D 지리 데이터를 기반으로 MLLM 에이전트의 도시 내 지속적 탐색 능력을 평가하는 첫 번째 샌드박스 환경이다.

핵심 기여도

핵심 아이디어

도시 내에서 MLLM 에이전트가 지속적으로 탐색하고 목표를 달성하려면 단순한 시각 인식을 넘어, 지속적인 공간 추정과 경로 수정 능력이 필요하다. UrbanGround는 홍콩의 실제 지형과 보행자 네트워크를 반영한 3D 환경을 Unity에 스트리밍하여, 에이전트가 첫인상 시점에서의 관찰을 바탕으로 지속적인 탐색을 수행하도록 설계되었다. 이는 단일 장면 인식에서 도시 전체의 공간적 이해로의 전환을 연구하는 데 중요한 기반이 된다.

기술적 접근법

주요 결과

의의 및 한계

UrbanGround는 MLLM 에이전트의 도시 내 지속적 탐색 능력을 평가하는 데 중요한 기반을 제공하며, 단일 장면 인식에서 도시 전체의 공간적 이해로의 전환을 연구할 수 있는 첫 번째 샌드박스 환경이다. 그러나 현재 MLLM은 장거리 탐색 시 오류 누적이 발생하며, 환경 변화에 대한 적응력이 부족하다는 한계를 드러낸다. 이는 MLLM이 단순한 인식 능력을 넘어, 지속적인 공간 추정과 경로 수정 능력을 갖추어야 한다는 점을 시사한다.

실용적 활용

UrbanGround는 도시 내 자율 이동 시스템, 스마트 도시 관리, 로봇 탐색 연구 등에 활용될 수 있다. 특히, MLLM 기반 에이전트가 실제 도시 환경에서 지속적으로 탐색하고 적응하는 능력을 향상시키는 데 기초적인 연구 도구로 활용 가능하다.