Enabling robots to navigate open-world environments via natural language is critical for general-purpose autonomy. Yet, Vision-Language Navigation has relied on end-to-end policies trained on expensive, embodiment-specific robot data. While recent foundation models trained on vast simulation data show promise, the challenge of scaling and generalizing persists due to the limited scene diversity and visual fidelity in simulation.
To address this gap, we propose ImagiNav, a novel hierarchical paradigm that formulates navigation in visual space. Instead of predicting actions or metric waypoints, ImagiNav synthesizes a future egocentric video conditioned on language instructions, effectively serving as a high-level visual plan. This plan is subsequently interpreted by an inverse dynamics module, extracting metric trajectories for low-level execution. By decoupling visual planning from robot actuation, this paradigm enables the direct utilization of diverse, in-the-wild human navigation videos. To support this, we develop a scalable auto-labeling data collection pipeline that enhances motion annotation accuracy. Ultimately, ImagiNav demonstrates strong zero-shot transfer to robot navigation without requiring robot demonstrations, paving the way for generalist robots that learn directly from open-world data.