UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City
Paper Guide Brief
Reading Brief
UrbanGround is a real-scale urban sandbox built from Hong Kong 3D geospatial data, designed to evaluate multimodal large language model (MLLM) agents on tasks ranging from local scene understanding to long-range navigation, multi-stop planning, and dynamic replanning. The paper introduces a five-level evaluation ladder and reports that current MLLM agents perform well on visual recognition and short-range spatial reasoning but fail on sustained goal-directed behavior, with errors accumulating during extended exploration.
Central Claim
Introduces UrbanGround, the first real-scale urban sandbox for evaluating MLLM agents in closed-loop first-person navigation, with a five-level task ladder and comprehensive evaluation of state-of-the-art models.
Contribution
Introduces UrbanGround, the first real-scale urban sandbox for evaluating MLLM agents in closed-loop first-person navigation, with a five-level task ladder and comprehensive evaluation of state-of-the-art models.
Why It Matters
UrbanGround provides a physically constrained, real-scale city replica with interactive map and first-person control, enabling systematic study of how local perception composes into long-horizon spatial agency in MLLM agents.
Prerequisites
multimodal large language models, closed-loop interaction, first-person navigation, interactive map, task ladder
Atlas Placement
Computer Vision (subfield)
Read If
You care about multimodal large language models, closed-loop interaction, first-person navigation.
Skip If
You only care about UrbanGround, five-level evaluation ladder.
Noosaga Placements
- The paper focuses on evaluating MLLM agents' visual perception and spatial reasoning in a photorealistic urban environment, with tasks centered on visual recognition and scene understanding.Multimodal large language models (MLLMs) can interpret a street viewtest whether an agent can ground a local scene well enough to answer spatial questions after active observation
- Supervised Deep Learningframework80%The benchmark evaluates MLLM agents that are trained with supervised learning on large-scale multimodal data, and the tasks involve supervised visual recognition.Visual Recognition (VR)multimodal large language models
- The work evaluates general AI agents (MLLMs) on embodied navigation and planning tasks, contributing to the broader study of AI agency in complex environments.how far current MLLM agents can turn local urban perception into reliable actionsupport broader study of how far current MLLM agents can explore reliably
- Learning-Based Roboticsframework70%The agents are learning-based MLLMs that use perception and reasoning to act in the environment, fitting the learning-based robotics paradigm.MLLM agentsclosed-loop interaction from a first-person view
- The benchmark includes tasks requiring route planning, scheduling, and replanning under constraints, which are core to planning and search.Multi-stop Route PlanningTime-window SchedulingDynamic Road-closure Replanning
- Self-Supervised Learningframework50%MLLMs often rely on self-supervised pretraining, and the benchmark tests their ability to generalize to novel urban environments without task-specific fine-tuning.current MLLM agentsexplore reliably in complex, open-ended urban environments
- Markov Decision Processesframework40%Navigation tasks can be modeled as Markov decision processes, and the benchmark implicitly evaluates agents' ability to make sequential decisions under uncertainty.closed-loop interactiongoal-directed behavior
- The paper evaluates single-agent navigation, but the presence of pedestrians and dynamic obstacles touches on multi-agent interaction, though not the central focus.Navigation among Pedestrians
Abstract
Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current MLLM agents can turn local urban perception into reliable action in a complicated real-scale city. We propose UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data. UrbanGround supports closed-loop interaction from a first-person view and provides an interactive map for navigation. Agents can directly enter the 3D city and explore from a first-person view. Our analysis follows the growth of the spatial problem through three research questions. We first test whether an agent can ground a local scene well enough to answer spatial questions after active observation. Then we ask whether that grounding supports navigation as destinations become farther away and less explicit. Finally, we examine whether the resulting behavior survives changes in route availability and pedestrian motion. Contemporary MLLM agents usually show useful atomic abilities in visual recognition and short-range spatial reasoning, while orientation and pedestrian-aware movement remain unreliable. Their central failure emerges over extended exploration, where local abilities do not compose into sustained goal-directed behavior and errors accumulate without effective correction. We hope UrbanGround will support broader study of how far current MLLM agents can explore reliably in complex, open-ended urban environments.
Paper Context
Classified from the full extracted paper text (112,704 characters). The Paper Guide brief above is the user-facing synthesis; raw context is kept out of the page.
Full-paper context sent 112,704 of 112,704 extracted characters to classification.