Skip to content

Spatial-IQ shows humans at 82.1% while top multimodal models hit 17.7%

Original: Spatial-IQ exposes a 82.1% versus 17.7% gap in multimodal spatial reasoning View original →

Read in other languages: 한국어日本語
AI Aug 2, 2026 By Insights AI (Twitter) 1 min read 1 views Source
Spatial-IQ shows humans at 82.1% while top multimodal models hit 17.7%

A benchmark for where spatial reasoning breaks

Seeing an image is not the same as reasoning through its hidden 3D structure. NVIDIA AI wrote on X: "Humans count the boxes in this image, hidden ones included, with 82.1% accuracy. The best off-the-shelf multimodal model manages 17.7%." The source post is available on X.

Spatial-IQ is a diagnostic benchmark from NVIDIA Research. Instead of treating 3D object counting as one final-answer task, the project decomposes it into nine hierarchical perceptual and cognitive sub-tasks. Those include counting visible columns, understanding layers, and inferring hidden blocks that must exist to support the structure. The project page frames the question directly: do current multimodal LLMs share the same hierarchy humans use, or can they arrive at answers with broken intermediate reasoning?

The numbers make the gap hard to miss. According to the tweet, humans counted hidden boxes with 82.1% accuracy, while the strongest off-the-shelf multimodal model managed 17.7%. NVIDIA also says training on the Spatial-IQ sub-tasks lifted Qwen2.5-VL-32B object-counting accuracy from 2.9% to 62.6%. That makes the benchmark useful not only as a leaderboard, but as a training signal for targeted spatial capability.

NVIDIA AI often posts about research that connects model capability with its broader GPU and AI infrastructure stack. The next thing to watch is external adoption: whether other labs reproduce the hierarchy, whether video and robotics tasks show the same failure modes, and whether improvements on Spatial-IQ transfer to real manipulation, navigation, and scene-understanding workloads.

Share: Long

Related Articles