In the field of computer vision, a crucial role is played by 3D reconstruction technology, which recreates the real world in 3D from images. Based on the 3D Vision Foundation Model (VFM), NAVER LABS Europe developed an AI tool called “DUSt3R”, which has garnered significant attention by instantly converting 2D images into 3D with just one round of processing. Watch the video here >
This year, at the International Conference on Computer Vision and Pattern Recognition (CVPR 2024), an upgraded version of DUSt3R, called “MASt3R,” was introduced, demonstrating significantly enhanced performance.
(The name “MASt3R” stands for “Matching And Stereo 3D Reconstruction.”)
The Key Features of MASt3R
The most notable differences between DUSt3R and MASt3R can be found in their scale and detail. While DUSt3R was praised for its ability to quickly construct a 3D model from just one or two images, MASt3R can process thousands of large-scale image datasets, allowing for fast and accurate 3D modeling of complex environments such as cities or building interiors. Moreover, the precision has been significantly improved as well.
The improved map-free localization performance is especially impressive. For instance, when a robot is placed in an unfamiliar environment, utilizing MASt3R enables it to understand its new space better and move or perform tasks efficiently. Additionally, when developing AR/VR applications, the enhanced localization accuracy and depth information can create a more immersive experience.
The Amazing Potential of the 3D Vision Foundation Model
DUSt3R was already revolutionary, but this swift upgrade to MASt3R brings even greater implications as it showcases the remarkable expansion potential of the VFM. The foundation of both DUSt3R and MASt3R is NAVER LABS Europe’s CROCO—a model designed to understand the 3D world by learning images captured from different viewpoints of the same scene. It has the advantage of being finely tuned to allow infinite expansion into a variety of AI tools.
We plan to continue developing more accurate and efficient 3D vision tools through VFM research. Our aim is to bring revolutionary innovations to how we perceive and interact with the real world across various fields such as robotics, digital twins, smart cities, and XR.