World Labs introduces Atlas, a world model that natively understands 3D space
Atlas: A World Model for Spatial Intelligence
Atlas is a multimodal autoregressive diffusion transformer pretrained from scratch to handle text, images, video, and 3D. It stitches reference images into a coherent 3D scene with pixel-perfect camera control, generating up to 1 minute of 1440p video. On sparse-view 3D reconstruction (2–3 images), Atlas beats specialized reconstruction models; more input images reduce guesswork. Early access is open for request, but the post doesn't disclose pricing or a public launch date.
Why it matters: World Labs drops Atlas, a from-scratch omni world model that unifies text, images, video, and 3D into a shared spatial context. The sparse-view reconstruction claim — beating specialized models with only 2-3 images — is concrete and testable. Not a 95 because it's a blog post ...