How a 3D foundation model can help robots find their bearings - MBZUAI MBZUAI

How a 3D foundation model can help robots find their bearings

Thursday, July 30, 2026

When used in a new environment, visual-inertial navigation systems (VINS), like those used in robots, drones, and VR headsets, need to carry out a process known as initialization. The process determines how the system is moving, which direction gravity is pointing in relation to the hardware, and the scale of the environment. It usually takes a few seconds, and most users don’t notice it’s happening, but getting initialization estimates wrong can send the system off course. 

“Every time you drive a car, you have to start the engine first,” says Yuantai Zhang, a doctoral student in Robotics at MBZUAI. Initialization is a similar kind of starting point. 

Zhang and researchers from MBZUAI and other institutions have developed a new initialization approach that uses an AI technology known as a feed-forward 3D foundation model. The researchers call it an efficient feature-free initialization for monocular visual-inertial systems, and it was found to initialize faster than other methods and work in settings that make initialization difficult, such as dimly lit scenes or ones with repetitive patterns.  

The researchers recently presented a study on the topic at the Robotics: Science and Systems (RSS) conference in Sydney. Jiaqi Yang, Huajian Zeng, Changhao Chen, Haoang Li, Liang Li, Dezhen Song, and Xingxing Zuo are co-authors of the study. 

How VINS work 

VINS typically combine information from a camera and an inertial measurement unit (IMU), a small chip that senses acceleration and rotation. The two sensors provide different kinds of information and compensate for each other’s weaknesses. The camera captures detail about the scene but on its own can’t determine absolute distance or scale. The IMU measures acceleration in real units and it works even in poor lighting conditions or in scenes that lack distinct features. “The inertial sensor isn’t influenced by the environment and gives you information that vision can’t provide,” Zhang says. 

The process of initialization is important because VINS can only fuse data from the camera and IMU once the system knows its initial state. And since subsequent tracking builds each estimate on top of the last, errors at the beginning of the process compound over time. “If the initial state isn’t accurate, subsequent state estimation will drift,” Zhang says. 

Initialization methods have traditionally relied on visual features, such as corners and textures, that algorithms detect in one frame and match in the next. This matching is what stitches the separate frames into a common picture, and it requires an algorithm to estimate the 3D position of every matched feature. Relying on features poses a problem for drones and robots that need to initialize in places with few visual features, such as industrial corridors or rooms with blank walls. 

The researchers’ new method removes the need for visual features. It instead relies on a feed-forward 3D model, which can predict a 3D point cloud of the scene by analyzing a few images captured by the unit’s camera. Because the model delivers its predictions already stitched together, with every point from the images expressed in a single frame, there is nothing to match and no individual feature positions to estimate. The feed-forward 3D model comes with an important limitation, however. While it can predict the shape of objects in a scene, it can’t determine their actual size. 

The IMU can fill this gap, because its measurements are based in physical units. Initialization becomes a matter of aligning the geometry predicted by the feed-forward 3D model with the motion data recorded by the IMU. And since resizing the unified point cloud stretches every point by the same amount, the visual side of the problem is reduced to one value, which corresponds to the scale of the scene. 

“The equations are simpler, and because it’s feature-free, it’s inherently immune to feature-related noise,” Zhang says. 

Faster and more reliable initialization 

Zhang and his co-authors conducted experiments on two public benchmarks and on a dataset that they created themselves. Their method achieved the highest success rate, nearly 95% compared to 75% to 83% for other approaches on the TUM-VI dataset. It also only needed about one second of sensor data, while competing approaches needed approximately three to five seconds. 

On the researchers’ own dataset, their system achieved a 90% success rate compared to 30% to 50% for the other systems. The most significant gap between the approaches appeared in difficult indoor scenes where their method initialized reliably while the others failed. 

The researchers acknowledge that their approach has limits, however. The point cloud generated by the feed-forward 3D model is a prediction, not a true measurement. Reliable predictions depend on the model having some visual cues to work from. “If it’s a completely white wall, our system will fail, because the model needs some visual clues to create reliable predictions,” Zhang says. But this is a limitation for all the systems they tested. Open-sky outdoor scenes also caused all the systems to fail. 

Perhaps the most important aspect of the team’s work is that it is agnostic to which feed-forward 3D model it uses. When the researchers swapped the default model for newer ones, accuracy improved without any other changes to the overall systems. “With a stronger model, we can have even better performance,” Zhang says. 

Related

thumbnail
Friday, June 19, 2026

At robotics’ “take-off moment”, Dezhen Song is helping shape the future at MBZUAI

From building MBZUAI's robotics department to developing the next generation of AI talent, Professor Dezhen Song is.....

  1. vice provost ,
  2. OSPA ,
  3. faculty ,
  4. robotics ,
  5. innovation ,
Read More
thumbnail
Friday, June 05, 2026

How neural networks are making soft robots easier to control

MBZUAI's Cesare Stefanini has helped develop a new AI-driven control strategy that helps soft robotic arms move.....

  1. control ,
  2. soft robotics ,
  3. ICRA ,
  4. neural networks ,
  5. conference ,
  6. robotics ,
Read More
thumbnail
Tuesday, May 05, 2026

Commencement 2026: Hard work and soft robotics

From online courses to advanced research, Ibrahim Alsarraj’s rapid rise mirrors the growth of MBZUAI’s robotics department,.....

  1. research ,
  2. graduate ,
  3. robotics ,
  4. alumni ,
  5. Ph.D. ,
  6. M.Sc. ,
  7. commencement ,
  8. Commencement 2026 ,
Read More