A dataset named "spatial-iq" was published on the Hugging Face Hub under the account patrickqrim. It contains a single default configuration with a train split of roughly 3,000 rows, auto-converted to Parquet by the Hub. Each row describes a synthetic block-structure scene rendered from one of four viewpoints, with fields for object type (four classes, including "cube"), sample and view indices, camera offset (3 to 12.5), field of view (1 to 3), distance (0.3 to 1), total blocks (4 to 40), number of hidden blocks (0 to 13), columns (3 to 14), layers (1 to 4), a list of visible block coordinates, and eleven numbered task labels. A nested "mcq" field supplies five-option multiple-choice items with the correct letter and distractor types such as "high_by_1" and "low_by_2", suggesting use as a spatial-reasoning benchmark for vision-language models.
- huggingface.co2026-07-31