Recently, Li Fei-Fei, a pioneer in the field of artificial intelligence, and the core team of World Labs visited a16z, where they officially revealed their latest multi-modal world model, Atlas. Unlike large language models that predict the next token or video models that predict the next frame, Atlas focuses on "new view prediction," demonstrating significant potential to revolutionize traditional 3D reconstruction and content generation.
New View Prediction: The New Core Primitive of the World Model
For a long time, the core of large language models has been predicting the next token, while video generation models focus on predicting the next frame. In contrast, the essence of Atlas lies in "new view prediction." Users only need to input several images of a scene or related textual descriptions, and the model can accurately output images at any spatial and temporal position.
This mechanism achieves a deep integration of "generation" and "reconstruction" within the same model, capable of natively processing multimodal data such as text, images, videos, 3D data, and camera poses. The development team believes that the importance of new view prediction can be compared to the next token prediction of large language models, and it may be one of the key fundamental primitives toward achieving artificial general intelligence (AGI).
Breaking Traditional Boundaries: Reconstructing 3D Spaces with Minimal Data Costs
In traditional 3D reconstruction, hundreds or even thousands of photos are usually required to build a complete scene. Atlas completely changes this industry pain point by drastically reducing the amount of data needed to just a few or dozens of images. For example, using just three iPhone shots, it can directly generate a bullet-time video similar to movie special effects.
This breakthrough comes from Atlas's iterative improvement over its predecessor, the Marble model. Earlier versions of Marble faced technical bottlenecks due to the use of Gaussian splats as an output representation. Atlas, however, innovatively establishes "new view prediction" as a fundamental primitive, effectively solving the problem of blind spots in traditional 3D reconstruction and the need for generative techniques to fill these gaps. It also features a highly scalable spatial context window.
Wide Industry Application Prospects
As this technology continues to mature, its application scenarios are rapidly expanding into multiple cutting-edge fields:
Firstly, in the creative design industry, Atlas can provide highly 3D-consistent generated content, significantly lowering the barriers and time required for production. Secondly, in the construction and architecture industry, its high-precision spatial reconstruction and prediction capabilities can deeply support various stages of engineering design. Lastly, in the robotics field, it can greatly improve the efficiency of transforming real environments into simulation testing, effectively alleviating the long-standing issue of insufficient data for robot training. In the future, it will seamlessly integrate dynamic data, further bridging the gap in action planning capabilities.
Currently, Atlas already possesses preliminary dynamic processing capabilities. The development team revealed that future improvements will focus on dimensions such as dynamic effects, content editing, and human-machine interaction, with the ultimate goal of building a complete system that allows ordinary users to interact intuitively with a 3D virtual world.
