NVIDIA Research and Yale released Spatial-IQ, a diagnostic benchmark that decomposes 3D object counting in stacked structures into nine perceptual and cognitive sub-tasks plus two target tasks (object counting and mental rotation), modeled on Piaget & Inhelder's developmental hierarchy and the KABC Block Counting subtest. The dataset comprises roughly 80,000 procedurally generated scenes built in NVIDIA Isaac Sim 5.1 on a 4x4x4 voxel grid, split into a 3,000-sample evaluation set and a disjoint 68,000-sample training set, with paper, Hugging Face dataset and GitHub code published. Across eight text models (Gemini 3 Pro, GPT 5.4, Claude Opus 4.6, Qwen3.5-27B, Kimi K2.5, GLM 4.6 and small anchors) and three image-editing models, humans scored 82.1% on object counting versus 17.7% for the best model, and models often hit the target task without preserving the prerequisite hierarchy.
- nvidia.github.io2026-07-31