Research Radarcs.CVAug 27, 2026classified

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City

Tianjie Ju, Zheng Wu, Yueqing Sun, Yuhan Cui, Bobo Li, Shengqiong Wu, Pengzhou Cheng, Haodong Zhao, Zongru Wu, Xinbei Ma, Doris Zhang, Kunling Li, Mong-Li Lee, Wynne Hsu, Hao Fei, Qi Gu, Gongshen Liu, Zhuosheng ZhangarXivPDF
cs.CV

Paper Guide Brief

Reading Brief

UrbanGround is a real-scale urban sandbox built from Hong Kong 3D geospatial data, designed to evaluate multimodal large language model (MLLM) agents on tasks ranging from local scene understanding to long-range navigation, multi-stop planning, and dynamic replanning. The paper introduces a five-level evaluation ladder and reports that current MLLM agents perform well on visual recognition and short-range spatial reasoning but fail on sustained goal-directed behavior, with errors accumulating during extended exploration.

Central Claim

Introduces UrbanGround, the first real-scale urban sandbox for evaluating MLLM agents in closed-loop first-person navigation, with a five-level task ladder and comprehensive evaluation of state-of-the-art models.

Contribution

Introduces UrbanGround, the first real-scale urban sandbox for evaluating MLLM agents in closed-loop first-person navigation, with a five-level task ladder and comprehensive evaluation of state-of-the-art models.

Why It Matters

UrbanGround provides a physically constrained, real-scale city replica with interactive map and first-person control, enabling systematic study of how local perception composes into long-horizon spatial agency in MLLM agents.

Prerequisites

multimodal large language models, closed-loop interaction, first-person navigation, interactive map, task ladder

Atlas Placement

Computer Vision (subfield)

Read If

You care about multimodal large language models, closed-loop interaction, first-person navigation.

Skip If

You only care about UrbanGround, five-level evaluation ladder.

Methods
multimodal large language modelsclosed-loop interactionfirst-person navigationinteractive maptask laddervisual recognitionorientation understandingactive exploration
Tasks
visual recognitionorientation understandingactive explorationshort-range goal navigationlong-range goal navigationinstructional navigationconstrained navigationplace-type search
Datasets
Hong Kong 3D geospatial dataUrbanGround sandbox
Benchmarks
UrbanGroundfive-level evaluation ladder

Noosaga Placements

  • Computer Visionsubfield90%
    The paper focuses on evaluating MLLM agents' visual perception and spatial reasoning in a photorealistic urban environment, with tasks centered on visual recognition and scene understanding.
    Multimodal large language models (MLLMs) can interpret a street viewtest whether an agent can ground a local scene well enough to answer spatial questions after active observation
  • Supervised Deep Learningframework80%
    The benchmark evaluates MLLM agents that are trained with supervised learning on large-scale multimodal data, and the tasks involve supervised visual recognition.
    Visual Recognition (VR)multimodal large language models
  • The work evaluates general AI agents (MLLMs) on embodied navigation and planning tasks, contributing to the broader study of AI agency in complex environments.
    how far current MLLM agents can turn local urban perception into reliable actionsupport broader study of how far current MLLM agents can explore reliably
  • Learning-Based Roboticsframework70%
    The agents are learning-based MLLMs that use perception and reasoning to act in the environment, fitting the learning-based robotics paradigm.
    MLLM agentsclosed-loop interaction from a first-person view
  • The benchmark includes tasks requiring route planning, scheduling, and replanning under constraints, which are core to planning and search.
    Multi-stop Route PlanningTime-window SchedulingDynamic Road-closure Replanning
  • Self-Supervised Learningframework50%
    MLLMs often rely on self-supervised pretraining, and the benchmark tests their ability to generalize to novel urban environments without task-specific fine-tuning.
    current MLLM agentsexplore reliably in complex, open-ended urban environments
  • Markov Decision Processesframework40%
    Navigation tasks can be modeled as Markov decision processes, and the benchmark implicitly evaluates agents' ability to make sequential decisions under uncertainty.
    closed-loop interactiongoal-directed behavior
  • The paper evaluates single-agent navigation, but the presence of pedestrians and dynamic obstacles touches on multi-agent interaction, though not the central focus.
    Navigation among Pedestrians

Abstract

Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current MLLM agents can turn local urban perception into reliable action in a complicated real-scale city. We propose UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data. UrbanGround supports closed-loop interaction from a first-person view and provides an interactive map for navigation. Agents can directly enter the 3D city and explore from a first-person view. Our analysis follows the growth of the spatial problem through three research questions. We first test whether an agent can ground a local scene well enough to answer spatial questions after active observation. Then we ask whether that grounding supports navigation as destinations become farther away and less explicit. Finally, we examine whether the resulting behavior survives changes in route availability and pedestrian motion. Contemporary MLLM agents usually show useful atomic abilities in visual recognition and short-range spatial reasoning, while orientation and pedestrian-aware movement remain unreliable. Their central failure emerges over extended exploration, where local abilities do not compose into sustained goal-directed behavior and errors accumulate without effective correction. We hope UrbanGround will support broader study of how far current MLLM agents can explore reliably in complex, open-ended urban environments.

Paper Context

Source ContextWhole paper
Budget100,000 tokens
Coverage112,704 chars

Classified from the full extracted paper text (112,704 characters). The Paper Guide brief above is the user-facing synthesis; raw context is kept out of the page.

Full-paper context sent 112,704 of 112,704 extracted characters to classification.

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City | Research Radar