Papers
arxiv:2608.27456

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City

Published on Aug 27
ยท Submitted by
Tianjie
on Aug 28
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,

Abstract

UrbanGround evaluates whether multimodal language model agents can sustain reliable navigation and spatial reasoning in a realistic 3D city replica, revealing that local perceptual skills fail to compose into extended goal-directed behavior.

Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current MLLM agents can turn local urban perception into reliable action in a complicated real-scale city. We propose UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data. UrbanGround supports closed-loop interaction from a first-person view and provides an interactive map for navigation. Agents can directly enter the 3D city and explore from a first-person view. Our analysis follows the growth of the spatial problem through three research questions. We first test whether an agent can ground a local scene well enough to answer spatial questions after active observation. Then we ask whether that grounding supports navigation as destinations become farther away and less explicit. Finally, we examine whether the resulting behavior survives changes in route availability and pedestrian motion. Contemporary MLLM agents usually show useful atomic abilities in visual recognition and short-range spatial reasoning, while orientation and pedestrian-aware movement remain unreliable. Their central failure emerges over extended exploration, where local abilities do not compose into sustained goal-directed behavior and errors accumulate without effective correction. We hope UrbanGround will support broader study of how far current MLLM agents can explore reliably in complex, open-ended urban environments.

Community

Paper submitter

๐Ÿ™๏ธ UrbanGround turns a real-scale city into an interactive sandbox for multimodal agents.

What can you do with UrbanGround?

  • Explore a georegistered 3D replica of Hong Kong from a first-person view
  • Connect MLLM agents for closed-loop perception and physical control
  • Evaluate visual grounding, navigation, exploration, planning, and adaptation
  • Test agents under different times of day, weather, road closures, and moving pedestrians
  • Build new embodied urban tasks on top of the sandbox

Key takeaways

  • Current MLLMs are already strong at local urban perception, but this does not reliably translate into long-horizon spatial agency.
  • As navigation becomes longer and more complex, small spatial errors accumulate and agents often fail to recover.
  • Dynamic changes such as road closures and moving pedestrians remain particularly challenging.

๐ŸŒ† The sandbox is available as web and native builds, together with the evaluation code and tasks. We hope UrbanGround can serve as a playground for studying what it takes to turn strong local perception into reliable city-scale agency.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.27456
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2608.27456 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2608.27456 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2608.27456 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.