ImagiNav: Scalable Embodied Navigation via Generative Visual Prediction and Inverse Dynamics

Jie Chen1,2, Yuxin Cai2,3, Yizhuo Wang1, Ruofei Bai2,3, Yuhong Cao1, Jun Li2, Wei-Yun Yau2, Guillaume Sartoretti1
1Department of Mechanical Engineering, National University of Singapore, Singapore. 2Institute for Infocomm Research (I²R), Agency for Science, Technology and Research (A*STAR), Singapore. 3Nanyang Technological University (NTU), Singapore.

ImagiNav formulates navigation as generative visual prediction and inverse dynamics.

Abstract

Enabling robots to navigate open-world environments via natural language is critical for general-purpose autonomy. Yet, Vision-Language Navigation has relied on end-to-end policies trained on expensive, embodiment-specific robot data. While recent foundation models trained on vast simulation data show promise, the challenge of scaling and generalizing persists due to the limited scene diversity and visual fidelity in simulation.

To address this gap, we propose ImagiNav, a novel hierarchical paradigm that formulates navigation in visual space. Instead of predicting actions or metric waypoints, ImagiNav synthesizes a future egocentric video conditioned on language instructions, effectively serving as a high-level visual plan. This plan is subsequently interpreted by an inverse dynamics module, extracting metric trajectories for low-level execution. By decoupling visual planning from robot actuation, this paradigm enables the direct utilization of diverse, in-the-wild human navigation videos. To support this, we develop a scalable auto-labeling data collection pipeline that enhances motion annotation accuracy. Ultimately, ImagiNav demonstrates strong zero-shot transfer to robot navigation without requiring robot demonstrations, paving the way for generalist robots that learn directly from open-world data.

Overview

ImagiNav Framework

ImagiNav synthesizes future video

ImagiNav synthesizes a plausible future egocentric video from the current observation and a natural-language instruction generated by a semantic reasoner such as Gemini, producing a high-level visual plan that is then decoded into actions by an inverse dynamics module. To mitigate spatial ambiguity (e.g., confusing ``left'' with ``right''), a common failure mode in video models, we implement an Action-Conditioned Mixture-of-Experts (AC-MoE) strategy that routes the generation task to specialized experts, yielding precise motion dynamics.

Scalable Data Collection Pipeline

Geometry-first data pipeline

We develop a geometry-first data collection pipeline that enables training from in-the-wild human egocentric navigation videos without requiring precise state estimation. By extracting motion primitives through geometric inverse dynamics prior to semantic annotation, we mitigate spatial hallucination errors common in purely VLM labeling. This paradigm eliminates the need for robot teleoperation, precise localization systems, and metric calibration during data collection, making the approach inherently scalable.

Evaluation Highlights

Simulation Results

ImagiNav‑Real is trained solely on out‑of‑domain human videos with virtually no exposure to residential indoor scenes while generalizing zero‑shot to a Unitree H1 humanoid embodiment on VLN tasks on the high‑fidelity VLN‑PE benchmark.

VLN-PE results table

Imagination vs. Execution

ImagiNav generates high-fidelity visual trajectories that are interpreted by an Inverse Dynamics Model to pose trajectories for the controller. Inheriting the rich priors learned through pretraining, ImagiNav generates physically plausible videos navigating simulated house scenes despite being finetuned on nearly zero domestic environments.

Video Generation Results

Finetuned on a diverse corpus of real-world videos using our AC‑MoE strategy, ImagiNav‑Real significantly outperforms a variant trained only on simulated data and one that forgoes AC‑MoE.

AC-MoE ablation

Qualitative Comparison: Model finetuned with Real-world data demonstrates superior understanding of static geometry (Scene A), and dynamic behavior (Scene B).

Controllability and geometric grounding of our real-finetuned generation model: Given an initial observation, the model synthesizes diverse physically plausible future trajectories aligned with distinct natural language instructions.

Acknowledgements

ImagiNav is built upon the contributions of the open-source community. We specifically thank:

InternNav (baselines, VLN-PE benchmark and simulation), LTX-Video (pretrained generative video model for finetuning), VGGT (pretrained Inverse Dynamics Model), and CD-FVD (evaluation of Fréchet Video Distance).

BibTeX

@inproceedings{chen2026,
  author    = {Chen, Jie and Cai, Yuxin and Wang, Yizhuo and Bai, Ruofei and Cao, Yuhong and li, jun and Yau, Wei-Yun and Sartoretti, Guillaume Adrien},
  title     = {{ImagiNav: Scalable Embodied Navigation Via Generative Visual Prediction and Inverse Dynamics}},
  booktitle = {2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)},
  year      = {2026},
}