World Labs introduced Atlas, a model that generates, reconstructs and simulates 3D scenes. It was trained from scratch on text, images, video and depth data together. Atlas takes exact camera positions as direct input rather than text descriptions. It produces video up to a minute long at 1440p resolution from a few reference images. It also rebuilds real places from as few as two or three photos. Outputs include flat images, point clouds and Gaussian splats that render on a device. It can reframe footage shot on a few phones from impossible angles. It also builds simulated environments for training robots from short phone recordings. World Labs says human raters scored it above recent video models on camera following. It also reports lower reconstruction error than specialist open-source models.
What changed
Sparse-view 3D reconstruction and precise camera control needed separate specialist models or costly capture rigs.
What it unlocks
Generating consistent 3D scenes, point clouds and Gaussian splats from a handful of ordinary photos, with exact camera paths.
- up to 1 minute of video at 1440p
- 2-3 images for faithful reconstruction
- over 100 input images supported
- 3-5 cameras for video reframing
What you need to act on it
- early access request approved
- selected partner status
Sources