Pingyue Zhang*, Zihan Huang*, Yue Wang*, Jieyu Zhang*, Letian Xue, Zihan Wang, Qineng Wang, Keshigeyan Chandrasegaran, Ruohan Zhang, Yejin Choi, Ranjay Krishna, Jiajun Wu, Li Fei-Fei, Manling Li
(* equal contribution)
[2026/01] Theory of Space is accepted by ICLR 2026
We introduce Theory of Space (ToS), a benchmark evaluating whether foundation models can actively construct spatial beliefs from partial observations. Unlike passive reasoning, ToS requires agents to explore, revise, and exploit a globally consistent spatial memory. Current multimodal models struggle with this active, self-directed construction of spatial belief, often relying on passive reasoning from static views.
Theory of Space is the ability to build a mental map from partial views. We define it as three coupled abilities:
- Construct: Actively explore and integrate partial observations into a globally consistent belief.
- Revision: Revise the belief when new evidence conflicts with earlier assumptions.
- Exploit: Use the current belief to answer spatial queries and guide the next action.
To construct a spatial belief, an agent must actively explore and integrate partial observations. We use procedurally generated multi-room layouts on N×M grids with paired environments:
- Text World: Symbolic observations with direction/distance bins (pure reasoning).
- Vision World: Egocentric RGB images from ThreeDWorld (perception + reasoning).
The agent uses an action space of Goto (move to visible object), Rotate (90°/180°/270°), Observe (get text/visual view), and Query (get coordinates).
The agent must use its current belief to answer spatial queries and guide the next action. We evaluate how the learned map is used at two levels:
- Route-level tasks test egocentric, path-based reasoning.
- Survey-level tasks test allocentric, map-like reasoning.
Survey-level probes ask whether a model can infer unseen views and handle geometric transformations beyond memorized paths.
| Methods | Avg. step | Route | Survey | Avg. | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| direction | persp.take | perc.dec | act2view | view2act | alloc.map | ment.rot | loc2view | view2loc | |||
| Vision-based World | |||||||||||
| Proprietary Models | |||||||||||
| GPT-5.2 | 17.2 | 40.0 | 36.7 | 56.2 | 43.8 | 40.3 | 43.4 | 59.7 | 56.9 | 37.8 | 46.0 |
| Gemini-3 Pro | 13.6 | 56.3 | 36.7 | 68.2 | 47.2 | 54.0 | 63.5 | 73.0 | 65.4 | 52.2 | 57.3 |
| Claude-4.5 Sonnet | 19.6 | 23.7 | 23.3 | 18.7 | 33.3 | 10.7 | 37.4 | 34.7 | 33.7 | 50.9 | 29.6 |
| Open-source Models | |||||||||||
| GLM-4.6V | 15.0 | 15.8 | 18.5 | 3.3 | 14.0 | 0.7 | 18.9 | 8.0 | 18.5 | 31.8 | 14.4 |
| Qwen3-VL | 16.3 | 16.8 | 23.3 | 13.4 | 24.8 | 5.7 | 25.8 | 16.3 | 21.5 | 43.7 | 21.3 |
| Text-based World | |||||||||||
| Proprietary Models | |||||||||||
| GPT-5.2 | 11.4 | 68.8 | 70.5 | 80.3 | 71.0 | 53.7 | 77.9 | 81.0 | 79.1 | 66.0 | 72.0 |
| Gemini-3 Pro | 13.5 | 78.0 | 79.2 | 90.6 | 75.3 | 76.3 | 81.0 | 94.0 | 83.3 | 76.2 | 81.5 |
| Claude-4.5 Sonnet | 18.7 | 65.3 | 65.3 | 79.0 | 62.7 | 51.7 | 68.8 | 76.3 | 57.0 | 67.0 | 65.9 |
| Open-source Models | |||||||||||
| GLM-4.6V | 14.5 | 20.8 | 19.7 | 12.7 | 21.8 | 3.7 | 13.9 | 9.3 | 22.7 | 26.2 | 16.8 |
| InternVL-3.5 | 15.0 | 28.8 | 44.8 | 26.0 | 36.8 | 7.3 | 31.0 | 27.7 | 33.8 | 38.9 | 30.6 |
| Qwen3-VL | 14.1 | 32.3 | 45.7 | 48.2 | 33.3 | 11.7 | 36.4 | 34.7 | 35.7 | 49.9 | 36.8 |
| Methods | Route | Survey | Avg. | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| direction | persp.take | perc.dec | act2view | view2act | alloc.map | ment.rot | loc2view | view2loc | ||
| Vision-based World | ||||||||||
| Proprietary Models | ||||||||||
| GPT-5.2 | 47.3 | 35.0 | 63.9 | 54.5 | 49.3 | 64.8 | 83.3 | 50.3 | 65.6 | 57.1 |
| Gemini-3 Pro | 63.8 | 36.3 | 57.5 | 49.0 | 58.0 | 67.2 | 85.3 | 70.4 | 57.0 | 60.5 |
| Claude-4.5 Sonnet | 47.3 | 33.5 | 37.7 | 40.8 | 15.7 | 54.8 | 58.3 | 44.7 | 54.8 | 43.1 |
| Open-source Models | ||||||||||
| GLM-4.6V | 11.5 | 24.5 | 4.7 | 19.0 | 2.7 | 22.9 | 11.7 | 20.0 | 33.6 | 16.7 |
| Qwen3-VL | 20.8 | 28.3 | 22.7 | 16.7 | 4.7 | 33.2 | 21.7 | 27.3 | 40.8 | 24.9 |
| Text-based World | ||||||||||
| Proprietary Models | ||||||||||
| GPT-5.2 | 84.5 | 88.2 | 97.0 | 89.0 | 76.0 | 96.3 | 98.3 | 94.8 | 89.2 | 90.4 |
| Gemini-3 Pro | 82.7 | 92.7 | 97.0 | 87.5 | 75.7 | 86.2 | 91.3 | 85.7 | 80.0 | 86.5 |
| Claude-4.5 Sonnet | 73.0 | 80.7 | 90.7 | 77.7 | 59.0 | 76.9 | 74.3 | 59.2 | 70.7 | 73.6 |
| Open-source Models | ||||||||||
| GLM-4.6V | 22.3 | 39.8 | 25.0 | 25.3 | 4.7 | 21.2 | 9.0 | 27.0 | 35.7 | 23.4 |
| InternVL-3.5 | 36.7 | 67.8 | 42.7 | 41.2 | 8.7 | 37.3 | 19.3 | 38.7 | 43.8 | 37.4 |
| Qwen3-VL | 40.8 | 69.3 | 56.5 | 50.0 | 17.7 | 42.8 | 40.3 | 42.5 | 54.6 | 45.6 |
We probe the agent's internal belief state to understand why failures occur. We provide a direct window into the agent's spatial belief via explicit cognitive-map probing. The agent outputs a structured cognitive map (N×M grid) at each step.
- Correctness (final): Evaluates the predicted global map at the last turn.
- Perception: Compares the predicted local map to the ground-truth local map for the current field of view.
- Self-tracking: Compares the agent pose inferred from the predicted global map to the ground-truth agent state.
- Local ↔ Global: Compares local and global predictions within the same turn (coherence check).
- Stability: Checks if previously observed objects degrade in the map over time.
- Uncertainty: Can the agent identify which regions it hasn't seen yet?
| Methods | Correctness | Perception | Local ↔ Global | Stability | Self-tracking | Uncertainty | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Ori. | Pos. | Overall | Ori. | Pos. | Ori. | Pos. | Ori. | Pos. | Ori. | Pos. | ||
| Vision-based World | ||||||||||||
| GPT-5.2 | 20.2 | 42.0 | 32.2 | 33.5 | 72.4 | 57.9 | 58.7 | 65.4 | 56.4 | 93.3 | 64.7 | 53.7 |
| Gemini-3 Pro | 32.2 | 62.5 | 52.1 | 43.8 | 68.5 | 52.9 | 68.3 | 61.8 | 62.0 | 98.8 | 73.9 | 70.2 |
| Text-based World | ||||||||||||
| GPT-5.2 | 91.0 | 75.1 | 80.0 | 100 | 86.8 | 96.4 | 86.0 | 96.7 | 67.6 | 98.0 | 86.7 | 64.5 |
| Gemini-3 Pro | 92.5 | 75.5 | 81.4 | 99.9 | 88.2 | 91.6 | 84.8 | 90.8 | 67.7 | 99.9 | 85.2 | 79.2 |
- Perception is the bottleneck: Vision perception remains a key bottleneck, especially for object orientation.
- Unstable belief: Unstable cognitive map prediction degrades spatial belief beyond initial perception.
An agent must revise its belief when new evidence conflicts with earlier assumptions. We introduce a dynamic perturbation task to probe Belief Revision. After exploration, objects are secretly relocated, creating a "false belief" that conflicts with new observations. The agent must actively re-explore to identify changes and revise its map.
| Methods | Avg. Steps | Identification | Belief Correctness | Belief Inertia | ||||
|---|---|---|---|---|---|---|---|---|
| All | Red. | Ori. | Pos. | Ori. | Pos. | Ori. | Pos. | |
| Vision-based World | ||||||||
| GPT-5.2 | 13.06 | 6.20 | 14.3 | 68.0 | 16.7 | 42.9 | 68.9 | 34.7 |
| Gemini-3 Pro | 10.29 | 3.23 | 23.9 | 82.5 | 30.3 | 63.1 | 51.1 | 14.4 |
| Text-based World | ||||||||
| GPT-5.2 | 6.92 | 0.55 | 97.9 | 98.4 | 89.5 | 69.7 | 5.5 | 12.5 |
| Gemini-3 Pro | 7.79 | 0.18 | 98.7 | 98.8 | 91.8 | 72.9 | 7.9 | 5.7 |
- Severe Belief Inertia in Vision: Vision agents fail to overwrite obsolete orientation beliefs despite new evidence.
- High Redundancy: Vision agents take many redundant steps after changes are visible, indicating failure to recognize the revision is complete.
A clear modality gap persists: text-based settings consistently outperform vision-based settings in spatial belief construction and exploitation.
Performance drops when models must actively explore.
a) Performance and Efficiency Deficit: Active agents score lower than reasoning on rule-based program histories, and explore less efficiently than the program.
b) Incomplete Coverage: Active agent fails to achieve complete information coverage. Models explore redundantly, requiring ≥14 steps without improving belief accuracy, while rule-based proxies reach target coverage in ~9 steps.
c) Complexity-Widened Gap: The active–passive gap increases with complexity; Gemini-3 Pro scales much better. As the number of rooms increases, exploration cost rises accordingly.
| Methods | 2-room | 4-room | ||||
|---|---|---|---|---|---|---|
| pass. | act. | exp. | pass. | act. | exp. | |
| Text-based World | ||||||
| GPT-5.2 | 92.3 | 77.8 | 6.2 | 86.5 | 66.0 | 16.4 |
| Gemini-3 Pro | 86.7 | 80.6 | 6.2 | 81.2 | 77.7 | 19.7 |
| Vision-based World | ||||||
| GPT-5.2 | 59.3 | 51.5 | 10.8 | 52.6 | 40.3 | 23.2 |
| Gemini-3 Pro | 58.3 | 57.8 | 6.6 | 56.2 | 51.5 | 19.7 |
Orientation, stability, and belief drift issues:
- Orientation Gap: Vision perception is a bottleneck, especially for object orientation.
- Unstable Map: Beliefs about previously observed objects degrade over time.
- Belief Drift: New updates corrupt earlier correct perceptions, lowering final correctness.
Map correctness correlates with downstream success.
- Sufficiency Test: Conditioning on ground-truth maps yields near-perfect accuracy (~95%), confirming the JSON map format captures all necessary information for tasks.
- Alignment Test: Prompting models to explicitly generate maps before answering slightly degrades performance. This externalization gap indicates the model's latent internal belief is richer than its discretized JSON output.
While lossy, the explicit map remains a strong diagnostic proxy. Map correctness correlates significantly with downstream success:
| Methods | Text (%) | Vision (%) |
|---|---|---|
| GPT-5.2 | 41.8 | 57.0 |
| Gemini-3 Pro | 46.6 | 64.5 |
Pearson correlation (r) between spatial-belief correctness and downstream evaluation performance. All correlations are significant (p<.001).
Vision agents persist in obsolete beliefs.
- Vision-based Revision Failures: Vision agents suffer from excessive exploration redundancy and poor accuracy in identifying object shifts.
- Belief Inertia: Agents, especially vision-based ones, persist in obsolete spatial coordinates despite new observations.
- Clone the
releasebranchgit clone --single-branch --branch release https://github.com/mll-lab-nu/Theory-of-Space.git cd Theory-of-Space - Configure keys and run setup
- Open
setup.shand set any required API keys / environment variables. - Run:
source setup.sh --run-exp # setup + run experiments
To add a new model, edit scripts/SpatialGym/base_model_config.yaml and add an entry under the models section. Each model requires specific parameters based on its provider:
gpt-5.2:
provider: openai # API provider (`openai`, `anthropic`)
model_name: gpt-5.2 # Model identifier used by the API
max_completion_tokens: 32768 # Maximum tokens for response
temperature: 1.0 # Sampling temperature
max_workers: 128 # Maximum parallel API calls
max_retries: 5 # Retry rounds for batch requests
max_retries_api: 5 # Retry attempts per API request
timeout: 500 # Request timeout in seconds
reasoning_effort: medium # Reasoning effort (low, medium, high)First run the following command to serve the model:
vllm serve Qwen/Qwen3-VL-2B-Instruct \
--host 0.0.0.0 \
--port 9999 \
--dtype bfloat16 \
--served-model-name qwen3-vl-2b-instruct \
--max_model_len 128000 \Then add the following entry to scripts/SpatialGym/base_model_config.yaml:
qwen3-vl-2b-instruct:
provider: openai
organization: self-hosted
model_name: qwen3-vl-2b-instruct # same as served-model-name in vllm serve command
base_url: http://localhost:9999/v1 # same as host and port in vllm serve command
max_completion_tokens: 8192
temperature: 0.0
max_workers: 16
max_retries: 3
timeout: 600For other models, you can refer to the scripts/SpatialGym/base_model_config.yaml for more details.
Run a single full pipeline (explore + eval + cogmap):
python scripts/SpatialGym/spatial_run.py \
--phase all \
--model-name gpt-5.2 \
--num 25 \ # 25 samples (run00 - run24)
--data-dir room_data/3-room/ \
--output-root result/ \
--render-mode vision,text \ # run both vision and text world
--exp-type active,passive \ # run both active and passive exploration
--inference-mode batch # run batch (or direct) inferenceSee scripts/SpatialGym/README.md for more commands.
Results are organized as:
{output_root}/
└─ {model_name}/
└─ {room_hash}/
└─ {render_mode}/
└─ {exp_type}/
└─ {think|nothink}/
├─ config.json
├─ exploration.json
├─ evaluation.json
├─ images/
└─ [proxy_agent]/ # only when exp_type = passive
├─ config.json
├─ exploration.json
├─ evaluation.json
└─ images/
After running the experiments, you can find the visulization html under result/gpt-5.2/env_data.html.
By executing python -m http.server <port> under the root directory, you can access the visulization html at http://localhost:<port>/result/gpt-5.2/env_data.html.
The main page will have explore + eval + optional cogmap + false belief + correlation metrics.
Each sample page features comprehensive data for each turn, along with relevant sample metrics.
To generate your own scenes, check https://github.com/yw2544/ToS-vision-scenes
- ToS-vision-scenes: repo for generating 3D scenes for Theory of Space.
- VAGEN: active exploration is implemented based on VAGEN.
If you find our framework and paper useful, we appreciate it if you could cite our work:
@inproceedings{zhang2026theoryofspace,
title = {Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?},
author = {Zhang, Pingyue and Huang, Zihan and Wang, Yue and Zhang, Jieyu and Xue, Letian and Wang, Zihan and Wang, Qineng and Chandrasegaran, Keshigeyan and Zhang, Ruohan and Choi, Yejin and Krishna, Ranjay and Wu, Jiajun and Li, Fei-Fei and Li, Manling},
booktitle = {International Conference on Learning Representations (ICLR)},
year = {2026},
}









