
World Labs has announced Atlas, its next generation world model, which outputs reconstructed scenes as 3D gaussian splats. RTFM, the model World Labs announced last October, generated each new frame directly from earlier frames and carried no explicit geometric representation. Atlas natively operates on both 2D image frames and 3D depth maps, which lets it write worlds out as point clouds or as splats.
Point clouds estimate a scene's geometry, and Atlas fills the remaining gaps to turn the point cloud into a complete splat scene that renders on device at high resolution and framerates. This is the same representation used in Marble, which reached version 1.1 in April. From a single input image, Atlas produces a full 3D world by jointly generating new views and estimating their geometry.
Atlas is an omni model pretrained from scratch to operate natively on text, images, video and 3D. The architecture is a multimodal autoregressive diffusion transformer, specifically a rectified flow model. Every input is folded into a shared spatial context in which each image is grounded at a 3D position, and generation is conditioned on that context.
Camera controlled generation takes camera geometry as a native input type and returns up to one minute of video at 1440p from one or more input images. World Labs claims faithful reconstruction from as few as two or three images, and says the spatial context can hold over a hundred. The Stanford Main Quad example was built from two to twenty-five ground level photographs, from which Atlas generated aerial camera paths.
The are also showcasing some dynamic work. Space time simulation covers two workflows. Atlas reframes video from three to five camera views shot on ordinary cell phones mounted on tripods and clamps, and for Real-to-Sim robotics it generates the RGB and depth a simulated robot's body mounted cameras would observe, with two large environments reconstructed from 24 frames of cell phone video each. Image generation, including 360 panoramas from text, is presented in the post as a secondary capability.
Every published figure is World Labs' own measurement, and the company states it reproduced the baseline results itself. On camera-controlled generation, third party human raters chose Atlas over MiniMax H3 75 percent of the time and over Seedance 2.5 94 percent of the time.
World Labs published no paper, no arXiv entry, no model card and no code alongside the post. World Labs says Atlas will power future versions of Marble and other products. Atlas is entering early access with select partners, but you can request early access here. World Labs counts Ben Mildenhall among its cofounders, a name anyone interested in gaussian splatting, NeRF, or any type of radiance field must know. Mildenhall has been posting some sample outputs from Atlas through his X account.
I highly suggest spending some time looking at their project page. More to come.
New to Gaussian Splatting? Start here






