NAVER LABS Europe, based in the scientific city of Grenoble at the foot of the French Alps, is at the forefront of robotics and artificial intelligence. AI researchers from 26 countries around the world are working to create practical answers to the long-standing question: “Can robots understand the world like humans do?”
Among the various research topics, foundation models are attracting the most attention, and NAVER LABS Europe is playing a leading role in applying these foundation models to the field of robotics.
The following is a reconstructed interview with Martin Humenberger (Director of Science) and Florent Perronnin (Technical Advisor), who are leading these efforts.
“A Model for Robots to Understand the World”
Florent: Over the past three to four years, NAVER LABS Europe has been focusing on research into foundation models for robotics. In the past, every time a new problem arose, a dedicated AI system had to be built from scratch, which required significant time and effort. But foundation models, which are large-scale, general-purpose AI models trained on massive datasets and adaptable to a wide range of downstream tasks, are overcoming these limitations.
Martin: Our goal is to use AI to make NAVER’s service robots more versatile and trustworthy companions. To achieve this, we are focusing on three major research directions.
First is Vision. We develop AI that enables robots to understand the environments they operate in. Second is Action. We study how robots can act effectively in the physical world. And third is Interaction. We enable robots to interact efficiently not only with humans but also with other robots.
Florent: Centered on these three pillars, we’ve already achieved a variety of results. In visual perception, we’ve made major progress in enabling robots not just to see the world, but to understand the structure and depth of the spaces around them. For example, robots can now recognize where objects are located and how far away they are, allowing them to perceive their surroundings in three dimensions. In terms of decision-making, we are one of the first research labs in the world to apply foundation models to optimization problems such as path planning and task allocation for robots. In the area of interaction, we’ve also made significant advances in using large foundation models effectively in real-world environments.
Martin: For a long time, enabling robots to find their own paths under complex constraints and optimizing the execution of multiple assigned tasks were considered very difficult challenges. But foundation models have shown the potential to handle such complex problems flexibly within a single model.
Florent: Ultimately, the key is versatility. Robots must be able to operate regardless of location, environment, or conditions to respond autonomously to unexpected situations and provide a wide range of services. In that sense, it could be said that the foundation model is the crucial key to making robots true all-around assistants, capable of adapting to many different situations and tasks through a single model.
Martin: Since my Ph.D. days, I’ve been studying 3D computer vision and robotic autonomy, always driven by the question, “Can robots understand the world like humans do?” It’s truly exciting to live in a time when that possibility is becoming more and more real.
NAVER LABS Europe’s research focuses on combining foundation models, a new paradigm in AI, with robotics in order to give robots the ability to understand the physical world and carry out tasks autonomously.