Keji He

dblp:319/4518 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0001-5136-8444ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Vision and language · 68% Robot navigation and mapping · 15% Deep learning architectures and training · 8%
Network and information security
1 paper
Security and privacy of machine learning · 100%
Computer graphics and multimedia
1 paper
Image and video processing · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
vision-and-language navigation
3.852026
Fine-Grained Alignment Supervision Matters in Vision-and-Language Navigation · IEEE Trans. Pattern Anal. Mach. Intell. 2026
ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Everyday Object Meets Vision-and-Language Navigation Agent via Backdoor · NeurIPS 2024
Computer vision › Vision and language
cross-modal alignment
1.522026
Fine-Grained Alignment Supervision Matters in Vision-and-Language Navigation · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Landmark-RxR: Solving Vision-and-Language Navigation with Fine-Grained Alignment Supervision · NeurIPS 2021
Robotics › Robot navigation and mapping › robot mapping
topological mapping
0.912025
ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Security and privacy of machine learning › adversarial attack
backdoor attack
0.812024
Everyday Object Meets Vision-and-Language Navigation Agent via Backdoor · NeurIPS 2024
Machine learning › Deep learning architectures and training
data augmentation
0.712023
Frequency-Enhanced Data Augmentation for Vision-and-Language Navigation · NeurIPS 2023
Image and video processing › frequency domain analysis
frequency-domain image processing
0.712023
Frequency-Enhanced Data Augmentation for Vision-and-Language Navigation · NeurIPS 2023
Machine learning › Reinforcement learning
reward design
0.512021
Landmark-RxR: Solving Vision-and-Language Navigation with Fine-Grained Alignment Supervision · NeurIPS 2021
Performance modeling and evaluation
benchmarking
0.512021
Landmark-RxR: Solving Vision-and-Language Navigation with Fine-Grained Alignment Supervision · NeurIPS 2021
Robotics › Robot navigation and mapping
obstacle avoidance
0.312025
ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Security and privacy of machine learning
adversarial attack
0.212024
Everyday Object Meets Vision-and-Language Navigation Agent via Backdoor · NeurIPS 2024
Machine learning › Representation and self-supervised learning › representation learning
visual representation learning
0.212023
Frequency-Enhanced Data Augmentation for Vision-and-Language Navigation · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

data augmentation · 2.3object-aware triggers · 1.5backdoor implantation · 1.5fourier transform · 1.3reward shaping · 1.0fine-grained alignment supervision · 1.0trial-and-error heuristic · 0.9transformer · 0.9cross-modal planning · 0.9reinitialization · 0.5re-initialization · 0.5
YearPublicationVenuePosition
2026 Fine-Grained Alignment Supervision Matters in Vision-and-Language Navigation
abstract
The Vision-and-Language Navigation (VLN) task involves an agent navigating within 3D indoor environments based on provided instructions. Achieving cross-modal alignment presents one of the most critical challenges in VLN, as the predicted trajectory needs to precisely align with the given instruction. This paper focuses on addressing cross-modal alignment in VLN from a fine-grained perspective. Firstly, to address the issue of weak cross-modal alignment supervision arising from coarse-grained data, we introduce a human-annotated fine-grained VLN dataset called Landmark-RxR. This dataset aims to offer precise, fine-grained supervision for VLN. Secondly, in order to comprehensively demonstrate the potential and advantage of the fine-grained data from Landmark-RxR, we explore the core components of the training process that depend on the characteristics of the training data. These components include data augmentation, training paradigm, reward shaping, and navigation loss design. Leveraging our fine-grained data, we carefully design methods for handling them and introduce a novel evaluation mechanism. The experimental results demonstrate that the fine-grained data can effectively improve the agent's cross-modal alignment ability.
Keji He, Yan Huang 0008, Ya Jing, Qi Wu 0001, Liang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments
abstract
Vision-language navigation is a task that requires an agent to follow instructions to navigate in environments. It becomes increasingly crucial in the field of embodied AI, with potential applications in autonomous navigation, search and rescue, and human-robot interaction. In this paper, we propose to address a more practical yet challenging counterpart setting - vision-language navigation in continuous environments (VLN-CE). To develop a robust VLN-CE agent, we propose a new navigation framework, ETPNav, which focuses on two critical skills: 1) the capability to abstract environments and generate long-range navigation plans, and 2) the ability of obstacle-avoiding control in continuous environments. ETPNav performs online topological mapping of environments by self-organizing predicted waypoints along a traversed path, without prior environmental experience. It privileges the agent to break down the navigation procedure into high-level planning and low-level control. Concurrently, ETPNav utilizes a transformer-based cross-modal planner to generate navigation plans based on topological maps and instructions. The plan is then performed through an obstacle-avoiding controller that leverages a trial-and-error heuristic to prevent navigation from getting stuck in obstacles. Experimental results demonstrate the effectiveness of the proposed method. ETPNav yields more than 10% and 20% improvements over prior state-of-the-art on R2R-CE and RxR-CE datasets, respectively.
Dong An 0002, Hanqing Wang 0001, Wenguan Wang, Zun Wang 0001, Yan Huang 0008, Keji He, Liang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Everyday Object Meets Vision-and-Language Navigation Agent via Backdoor
abstract
Vision-and-Language Navigation (VLN) requires an agent to dynamically explore environments following natural language. The VLN agent, closely integrated into daily lives, poses a substantial threat to the security of privacy and property upon the occurrence of malicious behavior. However, this serious issue has long been overlooked. In this paper, we pioneer the exploration of an object-aware backdoored VLN, achieved by implanting object-aware backdoors during the training phase. Tailored to the unique VLN nature of cross-modality and continuous decision-making, we propose a novel backdoored VLN paradigm: IPR Backdoor. This enables the agent to act in abnormal behavior once encountering the object triggers during language-guided navigation in unseen environments, thereby executing an attack on the target scene. Experiments demonstrate the effectiveness of our method in both physical and digital spaces across different VLN agents, as well as its robustness to various visual and textual variations. Additionally, our method also well ensures navigation performance in normal scenarios with remarkable stealthiness.
Keji He, Jiawang Bai, Yan Huang 0008, Qi Wu 0001, Shutao Xia, Liang Wang 0001
NeurIPS1
2024 Memory-Adaptive Vision-and-Language Navigation
Keji He, Ya Jing, Yan Huang 0008, Zhihe Lu, Dong An 0002, Liang Wang 0001
Pattern Recognit.1
2023 Frequency-Enhanced Data Augmentation for Vision-and-Language Navigation
abstract
Vision-and-Language Navigation (VLN) is a challenging task that requires an agent to navigate through complex environments based on natural language instructions. In contrast to conventional approaches, which primarily focus on the spatial domain exploration, we propose a paradigm shift toward the Fourier domain. This alternative perspective aims to enhance visual-textual matching, ultimately improving the agent's ability to understand and execute navigation tasks based on the given instructions. In this study, we first explore the significance of high-frequency information in VLN and provide evidence that it is instrumental in bolstering visual-textual matching processes. Building upon this insight, we further propose a sophisticated and versatile Frequency-enhanced Data Augmentation (FDA) technique to improve the VLN model's capability of capturing critical high-frequency information. Specifically, this approach requires the agent to navigate in environments where only a subset of high-frequency visual information corresponds with the provided textual instructions, ultimately fostering the agent's ability to selectively discern and capture pertinent high-frequency features according to the given instructions. Promising results on R2R, RxR, CVDN and REVERIE demonstrate that our FDA can be readily integrated with existing VLN approaches, improving performance without adding extra parameters, and keeping models simple and efficient. The code is available at https://github.com/hekj/FDA.
Keji He, Chenyang Si, Zhihe Lu, Yan Huang 0008, Liang Wang 0001, Xinchao Wang
NeurIPS1
2021 Landmark-RxR: Solving Vision-and-Language Navigation with Fine-Grained Alignment Supervision
abstract
In Vision-and-Language Navigation (VLN) task, an agent is asked to navigate inside 3D indoor environments following given instructions. Cross-modal alignment is one of the most critical challenges in VLN because the predicted trajectory needs to match the given instruction accurately. In this paper, we address the cross-modal alignment challenge from the perspective of fine-grain. Firstly, to alleviate weak cross-modal alignment supervision from coarse-grained data, we introduce a human-annotated fine-grained VLN dataset, namely Landmark-RxR. Secondly, to further enhance local cross-modal alignment under fine-grained supervision, we investigate the focal-oriented rewards with soft and hard forms, by focusing on the critical points sampled from fine-grained Landmark-RxR. Moreover, to fully evaluate the navigation process, we also propose a re-initialization mechanism that makes metrics insensitive to difficult points, which can cause the agent to deviate from the correct trajectories. Experimental results show that our agent has superior navigation performance on Landmark-RxR, en-RxR and R2R. Our dataset and code are available at https://github.com/hekj/Landmark-RxR.
Keji He, Yan Huang 0008, Qi Wu 0001, Dong An 0002, Shuanglin Sima, Liang Wang 0001
NeurIPS1