Novel View Synthesis (NVS) is an AI technology that generates 3D scenes from new camera perspectives using multiple photos or video data.
Think about the photos stored on your smartphone. Each photo captures a scene from the specific angle of the camera. Traditionally, it was impossible to view the same scene from a different, unrecorded angle. But what if we could analyze just a few photos and create entirely new perspectives that were never actually captured? That’s exactly what NVS makes possible. As the name suggests, it “synthesizes (Synthesis) a new (Novel) view (View).”
At NAVER LABS, NVS has been a long-standing research focus. The technology was recently applied in a real production: the Disney+ drama Polaris. In one of the scenes filmed on a quiet suburban road, NVS was used to naturally synthesize the textures and atmosphere of an urban environment. The result is a vivid and realistic scene that looks as if it were captured in an actual city.
Large-scale action sequences aim for deep immersion, but filming them in real urban settings often involves challenges such as safety, traffic control, and staff management. To overcome these limitations, filmmakers usually shoot in controlled environments and composite city backgrounds during post-production. To make these composites appear seamless, each shot must be rendered from the precise viewpoint required by the scene. NVS makes this process possible by generating photorealistic images from any desired angle based on real spatial data. By reducing physical and logistical constraints, NVS allows creators to expand their storytelling possibilities and create scenes that would otherwise be difficult or impossible to film.
City Background Implementation Using NVS (Source: Disney+ Polaris)
By using Street View images from NAVER Map captured with the self-developed digital twin device P1, NAVER LABS precisely reconstructed the three-dimensional layout of an actual city. This data was then used to train models that learn realistic spatial representations.
During this process, moving elements such as people and vehicles were removed, while buildings and backgrounds were refined to achieve a lifelike level of detail. As a result, the system can now generate realistic and natural cityscapes from new, previously unrecorded viewpoints. This capability was applied to the drama’s background scenes, creating visuals that feel remarkably like real life.
Previously, NAVER LABS recreated a realistic skyline of Seoul using 3D digital twin technology in Sweet Home Season 2 on Netflix. This time, the team applied Street View images to generate scenes from entirely new perspectives. Although the two technologies differ in their approach, they share the same foundation: understanding and recreating real-world spaces in a digital form. Both also represent NAVER LABS’ ongoing pursuit of spatial intelligence technologies that bridge the physical and digital worlds.
When it comes to applying spatial intelligence technology to video production, the possibilities are endless. During pre-production, it allows creators to visualize and plan scenes from multiple perspectives without visiting the actual locations. During production, it enables them to generate highly realistic visuals from digital data alone, recreating scenes that once required on-site filming. This technology can also reconstruct the past appearances of locations that have since changed or visualize hard-to-access areas in vivid detail. In post-production, it can be used to extend backgrounds and enhance the overall completeness of scenes. NAVER LABS views spatial intelligence as not only a way to reduce costs but a way to expand creative freedom for filmmakers and content creators, opening up new possibilities for visual storytelling.
Learn more about NAVER LABS’ Spatial AI technology: