Gengze Zhou

dblp:348/6915 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0003-0279-9277ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Vision and language · 48% Robot navigation and mapping · 31% Language models and text generation · 12%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
vision-and-language navigation
4.152026
VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation Agents · ACL (1) 2026
Ground-Level Viewpoint Vision-and-Language Navigation in Continuous Environments · ICRA 2025
NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models · ECCV (7) 2024
Robotics › Robot navigation and mapping
embodied navigation
1.012026
VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation Agents · ACL (1) 2026
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal large language model evaluation
1.012026
VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation Agents · ACL (1) 2026
Machine learning › Trustworthy machine learning
robustness
1.012026
VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation Agents · ACL (1) 2026
Robotics › Robot navigation and mapping › visual navigation
language-guided navigation
0.912025
SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts · ICCV 2025
Robotics › Robot navigation and mapping
visual navigation
0.912025
SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts · ICCV 2025
Robotics › Robot navigation and mapping › mobile robot navigation › navigation planning
waypoint prediction
0.912025
Ground-Level Viewpoint Vision-and-Language Navigation in Continuous Environments · ICRA 2025
Computer vision › Vision and language › multimodal reasoning
embodied reasoning
0.812024
NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models · AAAI 2024
Natural language and speech › Language models and text generation
large language model
0.812024
NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models · AAAI 2024
Computer vision › Vision and language › vision-language model
multimodal large language model
0.812024
NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models · ECCV (7) 2024
Natural language and speech › Language models and text generation › LLM agents
web navigation
0.812024
WebVLN: Vision-and-Language Navigation on Websites · AAAI 2024
Robotics › Robot navigation and mapping › learning-based navigation
zero-shot navigation
0.812024
NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models · AAAI 2024
Robotics › Legged, aerial and field robots › legged robots
quadruped robot
0.312025
Ground-Level Viewpoint Vision-and-Language Navigation in Continuous Environments · ICRA 2025
Natural language and speech › Language models and text generation
instruction following
0.212024
WebVLN: Vision-and-Language Navigation on Websites · AAAI 2024

Methods — techniques the papers use, named apart from their topics

chain-of-thought reasoning · 1.8self-reflection · 1.0in-context learning · 1.0weighted historical observations · 0.9state-adaptive routing · 0.9mixture of experts · 0.9connectivity graph transfer · 0.9zero-shot sequential action prediction · 0.8question-based instruction tuning · 0.8multimodal transformer · 0.8
YearPublicationVenuePosition
2026 VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation Agents
abstract
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across a wide range of vision-language tasks. However, their performance as embodied agents, which requires multi-round interaction with spatial reasoning and sequential action prediction, needs further exploration. Our work investigates this potential in the context of Vision-and-Language Navigation (VLN) by introducing a unified and extensible simulation-free evaluation framework to probe MLLMs as zero-shot agents, named VLN-MME. Simplifying the evaluation with a highly modular and accessible design streamlines experiments, enabling structured comparisons and component-level ablations across diverse MLLM architectures, agent designs, and navigation tasks. Crucially, enabled by VLN-MME, we observe that enhancing prevalent agents with Chain-of-Thought (CoT) reasoning and self-reflection leads to an unexpected performance decrease. This suggests MLLMs exhibit poor context awareness in embodied navigation tasks; although they can follow instructions and structure their output, their 3D spatial reasoning fidelity is low. Furthermore, we demonstrate that agent performance could be largely improved with simple failure cases in context learning. VLN-MME lays the groundwork for systematic evaluation of general-purpose MLLMs in embodied navigation settings and reveals limitations in their sequential decision-making capabilities. We believe these findings offer crucial guidance for MLLM post-training as embodied agents.
Xunyi Zhao, Gengze Zhou, Qi Wu 0001
ACL (1)2
2026 Toward Reasoning-Centric Video Object Segmentation via Multi-Modal Large Language Models
abstract
Referring Video Object Segmentation (RVOS) aims to segment the target objects specified in human instructions. Previous approaches typically rely on explicit human instructions that contain target categories or salient appearance descriptions. These approaches tend to fail when the instructions require temporal video understanding and complex relational reasoning. In this work, we present RViSeg, a reasoning-centric video object segmentation model that leverages the reasoning capability of Multi-modal Large Language Models (MLLM) to handle complex queries. The primary challenge lies in enabling MLLM to perform efficient pixel-level video perception. To tackle this challenge, we introduce a novel Spatial Token Merge (STM) module that consolidates lengthy video tokens into compact region-level clusters, while preserving essential spatial details. This structured representation enables MLLM to infer user intention by interleaving spatial and temporal visual information. Furthermore, we propose a Query-based Target Retrieval (QTR) module that utilizes learnable tokens as the target identity for mask prediction. By propagating these instance-specific tokens both intra-clip and inter-clip, our RViSeg effectively encodes object motion, ensuring spatio-temporal consistency in segmentation results. To facilitate training and evaluation, we construct InstructVideo, a single- and multiple-object reasoning video segmentation benchmark. Comprehensive experiments demonstrate the effectiveness of the proposed components.
Yanyan Shao, Shuting He, Gengze Zhou, Qi Ye 0001, Xiufang Shi, Jiming Chen 0001, Qi Wu 0001
IEEE Trans. Image Process.3
2025 SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
abstract
The academic field of learning instruction-guided visual navigation can be generally categorized into high-level category-specific search and low-level language-guided navigation, depending on the granularity of language instruction, in which the former emphasizes the exploration process, while the latter concentrates on following detailed textual commands. Despite the differing focuses of these tasks, the underlying requirements of interpreting instructions, comprehending the surroundings, and inferring action decisions remain consistent. This paper consolidates diverse navigation tasks into a unified and generic framework -- we investigate the core difficulties of sharing general knowledge and exploiting task-specific capabilities in learning navigation and propose a novel State-Adaptive Mixture of Experts (SAME) model that effectively enables an agent to infer decisions based on different-granularity language and dynamic observations. Powered by SAME, we present a versatile agent capable of addressing seven navigation tasks simultaneously that outperforms or achieves highly comparable performance to task-specific agents.
Gengze Zhou, Yicong Hong, Zun Wang 0001, Chongyang Zhao 0003, Mohit Bansal, Qi Wu 0001
ICCV1
2025 Ground-Level Viewpoint Vision-and-Language Navigation in Continuous Environments
abstract
Vision-and-Language Navigation (VLN) empowers agents to associate time-sequenced visual observations with corresponding instructions to make sequential decisions. However, dealing with visually diverse scenes or transitioning from simulated environments to real-world deployment is still challenging. In this paper, we address the mismatch between human-centric instructions and quadruped robots with a lowheight field of view, proposing a Ground-level Viewpoint Navigation (GVNav) approach to mitigate this issue. This work represents the first attempt to highlight the generalization gap in VLN across varying heights of visual observation in realistic robot deployments. Our approach leverages weighted historical observations as enriched spatiotemporal contexts for instruction following, effectively managing feature collisions within cells by assigning appropriate weights to identical features across different viewpoints. This enables low-height robots to overcome challenges such as visual obstructions and perceptual mismatches. Additionally, we transfer the connectivity graph from the HM3D and Gibson datasets as an extra resource to enhance spatial priors and a more comprehensive representation of real-world scenarios, leading to improved performance and generalizability of the waypoint predictor in real-world environments. Extensive experiments demonstrate that our Groundlevel Viewpoint Navigation (GVnav) approach significantly improves performance in both simulated environments and real-world deployments with quadruped robots.
Gengze Zhou, Haodong Hong, Yanyan Shao, Wenqi Lyu, Yanyuan Qiao, Qi Wu 0001
ICRA2
2024 WebVLN: Vision-and-Language Navigation on Websites
abstract
Vision-and-Language Navigation (VLN) task aims to enable AI agents to accurately understand and follow natural language instructions to navigate through real-world environments, ultimately reaching specific target locations. We recognise a promising opportunity to extend VLN to a comparable navigation task that holds substantial significance in our daily lives, albeit within the virtual realm: navigating websites on the Internet. This paper proposes a new task named Vision-and-Language Navigation on Websites (WebVLN), where we use question-based instructions to train an agent, emulating how users naturally browse websites. Unlike the existing VLN task that only pays attention to vision and instruction (language), the WebVLN agent further considers underlying web-specific content like HTML, which could not be seen on the rendered web pages yet contain rich visual and textual information. Toward this goal, we contribute a dataset, WebVLN-v1, and introduce a novel approach called Website-aware VLN Network (WebVLN-Net), which is built upon the foundation of state-of-the-art VLN techniques. Experimental results show that WebVLN-Net outperforms current VLN and web-related navigation methods. We believe that the introduction of the newWebVLN task and its dataset will establish a new dimension within the VLN domain and contribute to the broader vision-and-language research community. Code is available at: https://github.com/WebVLN/WebVLN.
Qi Chen 0014, Dileepa Pitawela, Chongyang Zhao 0003, Gengze Zhou, Hsiang-Ting Chen, Qi Wu 0001
AAAI4
2024 NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models
abstract
Trained with an unprecedented scale of data, large language models (LLMs) like ChatGPT and GPT-4 exhibit the emergence of significant reasoning abilities from model scaling. Such a trend underscored the potential of training LLMs with unlimited language data, advancing the development of a universal embodied agent. In this work, we introduce the NavGPT, a purely LLM-based instruction-following navigation agent, to reveal the reasoning capability of GPT models in complex embodied scenes by performing zero-shot sequential action prediction for vision-and-language navigation (VLN). At each step, NavGPT takes the textual descriptions of visual observations, navigation history, and future explorable directions as inputs to reason the agent's current status, and makes the decision to approach the target. Through comprehensive experiments, we demonstrate NavGPT can explicitly perform high-level planning for navigation, including decomposing instruction into sub-goals, integrating commonsense knowledge relevant to navigation task resolution, identifying landmarks from observed scenes, tracking navigation progress, and adapting to exceptions with plan adjustment. Furthermore, we show that LLMs is capable of generating high-quality navigational instructions from observations and actions along a path, as well as drawing accurate top-down metric trajectory given the agent's navigation history. Despite the performance of using NavGPT to zero-shot R2R tasks still falling short of trained models, we suggest adapting multi-modality inputs for LLMs to use as visual navigation agents and applying the explicit reasoning of LLMs to benefit learning-based models. Code is available at: https://github.com/GengzeZhou/NavGPT.
Gengze Zhou, Yicong Hong, Qi Wu 0001
AAAI1
2024 NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
Gengze Zhou, Yicong Hong, Zun Wang 0001, Xin Wang 0061, Qi Wu 0001
ECCV (7)1