Wenzhe Cai

dblp:261/2706 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 4 first-author · 12 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 ImagineNav++: Prompting Vision-Language Models as Embodied Navigator Through Scene Imagination
abstract
Visual navigation is a fundamental capability for autonomous home-assistance robots, enabling the execution of long-horizon tasks such as object search. While recent methods have leveraged Large Language Models (LLMs) to incorporate commonsense reasoning and improve exploration efficiency, their planning processes remain constrained by textual representations, which cannot adequately capture spatial occupancy or scene geometry-critical factors for informed navigation decisions. In this work, we explore whether Vision-Language Models (VLMs) can achieve mapless visual navigation using only onboard RGB/RGB-D streams, unlocking their potential for spatial perception and planning. We achieve this by developing the imagination-powered navigation framework ImagineNav++, which imagines the future observation images at valuable robot views and translates the complex navigation planning process into a rather simple best-view image selection problem for VLMs. Specifically, we first introduce a future-view imagination module, which distills human navigation preferences to generate semantically meaningful candidate viewpoints with high exploration potential. These imagined future views then serve as visual prompts for the VLM to identify the most informative viewpoint. To maintain spatial consistency, we develop a selective foveation memory mechanism, which hierarchically integrates keyframe observations through a sparse-to-dense framework, thereby constructing a compact yet comprehensive memory for long-term spatial reasoning. This integrated approach effectively transforms the challenging goal-oriented navigation problem into a series of tractable point-goal navigation tasks. Extensive experiments on open-vocabulary object and instance navigation benchmarks demonstrate that our ImagineNav++ achieves SOTA performance in mapless setting, even surpassing most cumbersome map-based methods, revealing the importance of scene imagination and scene memory in VLM-based spatial reasoning.
Teng Wang 0006, Xinxin Zhao, Wenzhe Cai, Changyin Sun 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 ImagineNav: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination
abstract
Visual navigation is an essential skill for home-assistance robots, providing the object-searching ability to accomplish long-horizon daily tasks. Many recent approaches use Large Language Models (LLMs) for commonsense inference to improve exploration efficiency. However, the planning process of LLMs is limited within texts and it is difficult to represent the spatial occupancy and geometry layout only by texts. Both are important for making rational navigation decisions. In this work, we seek to unleash the spatial perception and planning ability of Vision-Language Models (VLMs), and explore whether the VLM, with only on-board camera captured RGB/RGB-D stream inputs, can efficiently finish the visual navigation tasks in a mapless manner. We achieve this by developing the imagination-powered navigation framework ImagineNav, which imagines the future observation images at valuable robot views and translates the complex navigation planning process into a rather simple best-view image selection problem for VLM. To generate appropriate candidate robot views for imagination, we introduce the Where2Imagine module, which is distilled to align with human navigation habits. Finally, to reach the VLM preferred views, an off-the-shelf point-goal navigation policy is utilized. Empirical experiments on the challenging open-vocabulary object navigation benchmarks demonstrates the superiority of our proposed system.
Xinxin Zhao, Wenzhe Cai, Likun Tang, Teng Wang 0006
ICLR2
2025 InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts
abstract
The advancement of Embodied AI heavily relies on large-scale, simulatable 3D scene datasets characterized by scene diversity and realistic layouts.However, existing datasets typically suffer from limitations in data scale or diversity, sanitized layouts lacking small items, and severe object collisions.To address these shortcomings, we introduce \textbf{InternScenes}, a novel large-scale simulatable indoor scene dataset comprising approximately 40,000 diverse scenes by integrating three disparate scene sources, \ie, real-world scans, procedurally generated scenes, and designer-created scenes, including 1.96M 3D objects and covering 15 common scene types and 288 object classes.We particularly preserve massive small items in the scenes, resulting in realistic and complex layouts with an average of 41.5 objects per region.Our comprehensive data processing pipeline ensures simulatability by creating real-to-sim replicas for real-world scans, enhances interactivity by incorporating interactive objects into these scenes, and resolves object collisions by physical simulations.We demonstrate the value of InternScenes with two benchmark applications: scene layout generation and point-goal navigation. Both show the new challenges posed by the complex and realistic layouts. More importantly, InternScenes paves the way for scaling up the model training for both tasks, making the generation and navigation in such complex scenes possible. We commit to open-sourcing the data and benchmarks to benefit the whole community.
Weipeng Zhong, Peizhou Cao, Yichen Jin, Li Ray Luo, Wenzhe Cai, Jingli Lin, Zhaoyang Lyu, Xudong Xu, Bo Dai 0002, Jiangmiao Pang
NeurIPS5
2024 Bridging Zero-shot Object Navigation and Foundation Models through Pixel-Guided Navigation Skill
abstract
Zero-shot object navigation is a challenging task for home-assistance robots. This task emphasizes visual grounding, commonsense inference and locomotion abilities, where the first two are inherent in foundation models. But for the locomotion part, most works still depend on map-based planning approaches. The gap between RGB space and map space makes it difficult to directly transfer the knowledge from foundation models to navigation tasks. In this work, we propose a Pixel-guided Navigation skill (PixNav), which bridges the gap between the foundation models and the embodied navigation task. It is straightforward for recent foundation models to indicate an object by pixels, and with pixels as the goal specification, our method becomes a versatile navigation policy towards all different kinds of objects. Besides, our PixNav is a pure RGB-based policy that can reduce the cost of homeassistance robots. Experiments demonstrate the robustness of the PixNav which achieves 80+% success rate in the local path-planning task. To perform long-horizon object navigation, we design an LLM-based planner to utilize the commonsense knowledge between objects and rooms to select the best waypoint. Evaluations across both photorealistic indoor simulators and real-world environments validate the effectiveness of our proposed navigation strategy. More details are accessible via our project website https://sites.google.com/view/pixnav/.
Wenzhe Cai, Siyuan Huang 0004, Guangran Cheng, Yuxing Long, Peng Gao 0007, Changyin Sun 0001, Hao Dong 0003
ICRA1
2024 Discuss Before Moving: Visual Language Navigation via Multi-expert Discussions
abstract
Visual language navigation (VLN) is an embodied task demanding a wide range of skills encompassing understanding, perception, and planning. For such a multifaceted challenge, previous VLN methods totally rely on one model’s own thinking to make predictions within one round. However, existing models, even the most advanced large language model GPT4, still struggle with dealing with multiple tasks by single-round self-thinking. In this work, drawing inspiration from the expert consultation meeting, we introduce a novel zero-shot VLN framework. Within this framework, large models possessing distinct abilities are served as domain experts. Our proposed navigation agent, namely DiscussNav, can actively discuss with these experts to collect essential information before moving at every step. These discussions cover critical navigation subtasks like instruction understanding, environment perception, and completion estimation. Through comprehensive experiments, we demonstrate that discussions with domain experts can effectively facilitate navigation by perceiving instruction-relevant information, correcting inadvertent errors, and sifting through in-consistent movement decisions. The performances on the representative VLN task R2R show that our method surpasses the leading zero-shot VLN model by a large margin on all metrics. Additionally, real-robot experiments display the obvious advantages of our method over single-round self-thinking. Our project web can be seen at the https://sites.google.com/view/discussnav.
Yuxing Long, Xiaoqi Li 0020, Wenzhe Cai, Hao Dong 0003
ICRA3
2024 MO-DDN: A Coarse-to-Fine Attribute-based Exploration Agent for Multi-Object Demand-driven Navigation
abstract
The process of satisfying daily demands is a fundamental aspect of humans' daily lives. With the advancement of embodied AI, robots are increasingly capable of satisfying human demands. Demand-driven navigation (DDN) is a task in which an agent must locate an object to satisfy a specified demand instruction, such as "I am thirsty." The previous study typically assumes that each demand instruction requires only one object to be fulfilled and does not consider individual preferences. However, the realistic human demand may involve multiple objects. In this paper, we introduce the Multi-object Demand-driven Navigation (MO-DDN) benchmark, which addresses these nuanced aspects, including multi-object search and personal preferences, thus making the MO-DDN task more reflective of real-life scenarios compared to DDN. Building upon previous work, we employ the concept of ``attribute'' to tackle this new task. However, instead of solely relying on attribute features in an end-to-end manner like DDN, we propose a modular method that involves constructing a coarse-to-fine attribute-based exploration agent (C2FAgent). Our experimental results illustrate that this coarse-to-fine exploration strategy capitalizes on the advantages of attributes at various decision-making levels, resulting in superior performance compared to baseline methods. Code and video can be found at https://sites.google.com/view/moddn.
Peiqi Liu, Wenzhe Cai, Mingdong Wu, Zhengyu Qian, Hao Dong 0003
NeurIPS3
2024 DGMem: learning visual navigation policy without any labels by dynamic graph memory
Wenzhe Cai, Teng Wang 0006, Guangran Cheng, Lele Xu, Changyin Sun 0001
Appl. Intell.1
2023 Robust Navigation with Cross-Modal Fusion and Knowledge Transfer
abstract
Recently, learning-based approaches show promising results in navigation tasks. However, the poor generalization capability and the simulation-reality gap prevent a wide range of applications. We consider the problem of improving the generalization of mobile robots and achieving sim-to-real transfer for navigation skills. To that end, we propose a cross-modal fusion method and a knowledge transfer framework for better generalization. This is realized by a teacher-student distillation architecture. The teacher learns a discriminative representation and the near-perfect policy in an ideal environment. By imitating the behavior and representation of the teacher, the student is able to align the features from noisy multi-modal input and reduce the influence of variations on navigation policy. We evaluate our method in simulated and real-world environments. Experiments show that our method outperforms the baselines by a large margin and achieves robust navigation performance with varying working conditions.
Wenzhe Cai, Guangran Cheng, Lingyue Kong, Lu Dong 0002, Changyin Sun 0001
ICRA1
2023 Transmission Design and Component Allocation for STAR-RIS Assisted NOMA Systems with Direct Link
abstract
This paper investigates a simultaneously transmitting and reflecting reconfigurable intelligent surface (STAR-RIS) assisted downlink non-orthogonal multiple access (NOMA) system, consisting of a base station (BS), a STAR-RIS, a transmission user and a reflection user, where STAR-RIS uses mode switching protocols. Under the proposed transmission model, the base station serves the two users via the STAR-RIS and a direct link exists from the base station to the reflection user. By using the Beaulieu series, the statistical property of the Nakagami-m fading cascaded channel is characterized. Based on this, closed-from expressions of outage probability and diversity order for the reflection user and transmission user are obtained, respectively. In addition, to ensure the maximum gain in system throughput, a component allocation scheme is proposed, and the optimal allocation interval between the transmission and reflection components is obtained. Numerical results verify the theoretical analysis and demonstrate the superiority of the proposed scheme in terms of outage probability and system throughput compared to existing schemes.
Yuan Ren 0003, Wenzhe Cai, Suihu Yang
VTC Fall2
2023 Towards better generalization in quadrotor landing using deep reinforcement learning
Teng Wang 0006, Zichen He, Wenzhe Cai, Changyin Sun 0001
Appl. Intell.4
2023 UAV target following in complex occluded environments with adaptive multi-modal fusion
Lele Xu, Teng Wang 0006, Wenzhe Cai, Changyin Sun 0001
Appl. Intell.3
2023 Multi-objective deep reinforcement learning for crowd-aware robot navigation with dynamic human preference
Guangran Cheng, Yuanda Wang, Lu Dong 0002, Wenzhe Cai, Changyin Sun 0001
Neural Comput. Appl.4
2023 Learning a World Model With Multitimescale Memory Augmentation
abstract
Model-based reinforcement learning (RL) is regarded as a promising approach to tackle the challenges that hinder model-free RL. The success of model-based RL hinges critically on the quality of the predicted dynamic models. However, for many real-world tasks involving high-dimensional state spaces, current dynamics prediction models show poor performance in long-term prediction. To that end, we propose a novel two-branch neural network architecture with multi-timescale memory augmentation to handle long-term and short-term memory differently. Specifically, we follow previous works to introduce a recurrent neural network architecture to encode history observation sequences into latent space, characterizing the long-term memory of agents. Different from previous works, we view the most recent observations as the short-term memory of agents and employ them to directly reconstruct the next frame to avoid compounding error. This is achieved by introducing a self-supervised optical flow prediction structure to model the action-conditional feature transformation at pixel level. The reconstructed observation is finally augmented by the long-term memory to ensure semantic consistency. Experimental results show that our approach is able to generate visually-realistic long-term predictions in DeepMind maze navigation games, and outperforms the prevalent state-of-the-art methods in prediction accuracy by a large margin. Furthermore, we also evaluate the usefulness of our world model by using the predicted frames to drive an imagination-augmented exploration strategy to improve the model-free RL controller.
Wenzhe Cai, Teng Wang 0006, Changyin Sun 0001
IEEE Trans. Neural Networks Learn. Syst.1
2020 Attention-based face alignment: A solution to speed/accuracy trade-off
Teng Wang 0006, Xinjie Tong, Wenzhe Cai
Neurocomputing3