Lei Wang 0202

dblp:181/2817-202 · DBLP profile ↗
← Back
17ranked-venue papers
0as first author
17since 2021 · last 2026
0009-0001-6655-188XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Systems, architecture and hardware · 6 · 6 since 2021
YearPublicationVenuePosition
2026 DAWDet: A dynamic content-aware multi-branch framework with adaptive wavelet boosting for small object detection
Shaolei Liu, Dongchen Zhu, Lei Wang 0202, Jiamao Li
Pattern Recognit.4
2026 Learning to chase: Adaptive audio-visual navigation for moving sounds in complex environments
Yuanzheng He, Yuyi Liu, Chenfan Zhang, Dongchen Zhu, Lei Wang 0202
Pattern Recognit. Lett.5
2026 OMFlow: Optimizing optical flow via occlusion motion estimation
Wenjun Shi, Dongchen Zhu, Lei Wang 0202, Jiamao Li
Pattern Recognit. Lett.4
2025 $\mathbf{F}^{2} \mathbf{R}^{2}$: Frequency Filtering-Based Rectification Robustness Method for Stereo Matching
abstract
Most stereo matching networks assume that the stereo images are perfectly rectified, ignoring the perturbation of extrinsic parameters due to collisions, mechanical vibrations, and thermal expansion. This leads to poor rectification robustness in real-world stereo systems. That is, even minor rectification errors can lead to failure, making stereo systems unreliable for long-term autonomous operation in complex environments. In this paper, we are the first to propose a frequency filtering-based rectification robustness ($\mathbf{F}^{2} \mathbf{R}^{2}$) method for stereo matching, which aims to enhance the robustness of existing stereo networks to rectification errors. Specifically, we propose a sensitive frequency filter (SFF) to remove components susceptible to rectification errors within the frequency domain. SFF achieves the filtering through the learning-based adaptive filtering mask (AFM) guided by the spatial-frequency mapping modulation mask (SFM). Moreover, we build the matching feature reconstruction module (MFRM) to recover the features lost during filtering to benefit cost aggregation. Comprehensive experiments on simulated datasets and self-collected data validate that our method can significantly enhance the rectification robustness of stereo matching networks.
Haolong Zhou, Dongchen Zhu, Lei Wang 0202, Jiamao Li
ICRA4
2025 AKP: Actionable Knowledge-Augmented Agent Planning with Large Language Models
abstract
Large Language Models (LLMs) have demonstrated remarkable planning capabilities in complex language reasoning tasks. However, existing multi-step reasoning technologies easily introduces potential unreliable and inaccurate error accumulations over long-horizon action steps, and thereby making it exceedingly difficult to accurately explore the exponentially large search space. This deficiency primarily stems from the short-sighted greedy decoding of the next admissible action, and the lack of build-in actionable knowledge, which exacerbates the generation of hallucinatory or conflicting action sequences. To address these limitations, we propose Actionable Knowledge-augmented agent Planning (AKP), a novel LLM-based framework designed to enhance task planning of language agents by incorporating external actional knowledge. Specifically, we introduce a new knowledgeable policy called GVR by exploring future longer-term action paths, which leverages multiple domain foundation models (Verifier, Rewarder) trained on actional knowledge to jointly guide and calibrate more reasonable action generation. Additionally, AKP integrates the learned GVR into Monte Carlo Tree Search (MCTS) for deliberate planning, predicting various potential actions and iteratively refining alternative plans through lookahead and backtracking. Experimental results demonstrate that AKP significantly outperforms existing baselines, showcasing superior planning performance in tackling complex goal tasks across diverse embodied environments.
Dongchen Zhu, Lei Wang 0202, Jiamao Li
INDIN3
2025 WKAgent: World Knowledge-Guided Agent for Task Planning with Large Language Models
abstract
Recent advancements in using large language models (LLMs) as agent policies have demonstrated impressive planning capabilities for multi-step reasoning and planning tasks. Despite their achievements, LLM-based agents are prone to trial-and-error and planning hallucinations when generating action plans in embodied environments. This limitation stems from their lack of intrinsic world knowledge, resulting in a poor understanding of the physical world and ineffective decision-making. Imitating humans’ mental world model which provides commonsense prior knowledge when planning task, in this paper, we propose World Knowledge-Guided Agent Planning (WKAgent), a novel LLM planning approach that repurposes the LLM both a world knowledge model and an action reasoning model, and incorporates a search algorithm, such as Monte Carlo Tree Search (MCTS) to enhance the planning capabilities of language agents. Specifically, WKAgent empowers LLMs to self-synthesize dynamic world knowledge from expert trajectories to guide deliberate planning akin to human brains, which involves constantly tracking the goal instruction, summarizing world state changes, and anticipating the next course of actions during planning. Furthermore, WKAgent employs a lightweight knowledgeable policy model learned from contrastive correct-incorrect world knowledge-action pairs, aiming to constrain the next actions and calibrate more reasonable action paths. Experimental results on three complex tasks demonstrate that WKAgent can achieve superior task planning performance compared to various baselines. Further analysis indicates the explicit world knowledge from our WKAgent can improve the essential understanding capabilities of the real physical world and effectively alleviate the blind trial-and-error. Moreover, the knowledgeable learning policy model can steer the reasoning process towards generating more feasible action plans and mitigate the planning hallucinations.
Dongchen Zhu, Lei Wang 0202, Jiamao Li
INDIN3
2024 Rotated Orthographic Projection for Self-supervised 3D Human Pose Estimation
Yixuan Pan, Wenjun Shi, Dongchen Zhu, Lei Wang 0202, Jiamao Li
ECCV (69)5
2024 IPHGaze: Image Pyramid Gaze Estimation with Head Pose Guidance
Hekuangyi Che, Dongchen Zhu, Wenjun Shi, Lei Wang 0202, Jiamao Li
ICPR (28)6
2024 MemoFlow: Modifying Explicit Motion of Inconsistency in Optical Flow
Wenjun Shi, Dongchen Zhu, Lei Wang 0202, Jiamao Li
ICPR (30)4
2024 BCNet: Binocular Cooperative Network for Gaze Estimation
Dongchen Zhu, Minjing Lin, Hekuangyi Che, Wenjun Shi, Lei Wang 0202, Jiamao Li
ICPR (28)7
2024 CVFormer: Learning Circum-View Representation and Consistency for Vision-Based Occupancy Prediction via Transformers
abstract
With the increasing demands for perception accuracy in autonomous driving, there is a growing focus on fine-grained 3D semantic occupancy prediction. Effectively representing detailed three-dimensional scenes has become a significant challenge in the development of this task. In this paper, we present a novel transformer-based framework named CVFormer, which leverages two-dimensional circum-views from the ego to excavate three-dimensional features of the surrounding environment. Circum-views provide a novel solution for effectively addressing the representation of dense and fine-grained scenes. Specifically, a multi-attention module CTMA is designed for fusing temporal features from circum-views to fully exploit the spatiotemporal correlations between frames and capture more comprehensive clues. Furthermore, a novel 2D projection constraint is established by observing objects from different perspective directions, and multiple 3D constraints based on object invariance and semantic consistency are also conducted for supervising the network, which enhances its performance of understanding the scene. Experimental results on nuScenes dataset demonstrate that the proposed CVFormer obviously outperforms existing methods for occupancy prediction.
Zhengqi Bai, Wenjun Shi, Dongchen Zhu, Hanlong Kang, Gang Ye, Lei Wang 0202, Jiamao Li
ICRA8
2024 BEE-Net: Bridging Semantic and Instance with Gated Encoding and Edge Constraint for Efficient Panoptic Segmentation
abstract
Panoptic segmentation is a challenging perception task, which can help robots to comprehensively perceive the surrounding environment. In the task, we notice that semantic, instance, and panoptic have rich relations, however, which are rarely explored. In this work, we propose a novel panoptic, instance, and semantic bridged network to delve into the reciprocal relation. To make semantic and instance benefit from each other, we design a novel Gated Encoding (GE) module, incorporating complementary cues between semantic and instance heads through the gated mechanism. In addition, a novel edge-aware consistency constraint among edges of each task is presented, which exhaustedly exploits geometric constraints, to boost the segmentation quality of challenging edges. Experimental results on the Cityscapes and MS-COCO datasets demonstrate that our approach achieves state-of-the-art performance in an efficient CNN-based paradigm, attaining a balance between accuracy and efficiency.
Dongchen Zhu, Wenjun Shi, Gang Ye, Lei Wang 0202, Jiamao Li
ICRA8
2024 ESD-Pose: Enhanced Semantic Discrimination for Generalizable 6D Pose Estimation
Xingyuan Deng, Kangru Wang, Lei Wang 0202, Dongchen Zhu, Jiamao Li
PRCV (6)3
2024 Discriminative-Guided Diffusion-Based Self-supervised Monocular Depth Estimation
Dongchen Zhu, Lei Wang 0202, Jiamao Li
PRCV (6)4
2024 Continual Multiview Spectral Clustering via Multilevel Knowledge
abstract
Multiview clustering aims to integrate multiple features from different views to benefit the clustering task, which has attracted much attention in recent years. Most previous research has focused on exploring multiview clustering with a fixed set of tasks. However, it is still challenging to efficiently integrate with new clustering tasks, as it requires repeated access to previous data. To address the above challenges, this letter proposes a novel continual multiview spectral clustering model. The proposed model can efficiently achieve clustering in new tasks by transferring the accumulated knowledge from past tasks, while continuously refining the knowledge to ultimately improve the performance of all clustering tasks. Specifically, the knowledge sharing among different clustering tasks is considered at multiple levels, preserving both the heterogeneous distribution of different views and the relationships between multiple views. Meanwhile, our method is modelled by deep nonlinear structures, which allows to capture more hidden knowledge. In addition, an efficient alternating optimization algorithm is proposed to refine the knowledge online. The superior experimental results on several benchmark datasets show the effectiveness and efficiency of our method compared with other state-of-the-art models.
Kangru Wang, Lei Wang 0202, Jiamao Li
IEEE Signal Process. Lett.2
2023 Fast Extrinsic Calibration for Multiple Inertial Measurement Units in Visual-Inertial System
abstract
In this paper, we propose a fast extrinsic calibration method for fusing multiple inertial measurement units (MIMU) to improve visual-inertial odometry (VIO) localization accuracy. Currently, data fusion algorithms for MIMU highly depend on the number of inertial sensors. Based on the assumption that extrinsic parameters between inertial sensors are perfectly calibrated, the fusion algorithm provides better localization accuracy with more IMUs, while neglecting the effect of extrinsic calibration error. Our method builds two non-linear least-squares problems to estimate the MIMU relative position and orientation separately, independent of external sensors and inertial noises online estimation. Then we give the general form of the virtual IMU (VIMU) method and propose its propagation on manifold. We perform our method on datasets, our self-made sensor board, and board with different IMUs, validating the superiority of our method over competing methods concerning speed, accuracy, and robustness. In the simulation experiment, we show that only fusing two IMUs with our calibration method to predict motion can rival nine IMUs. Real-world experiments demonstrate better localization accuracy of the VIO integrated with our calibration method and VIMU propagation on manifold.
Youwei Yu, Fengjie Fu, Dongchen Zhu, Lei Wang 0202, Jiamao Li
ICRA6
2022 SRNet: Structural Relation-aware Network for Head Pose Estimation
abstract
Estimating head pose from a single RGB image has recently attracted considerable research attention. Prior arts employ a CNN backbone to process face images and then directly output Euler angles. We argue that they may ignore essential features that are highly correlated to head pose due to the non-global perspective, and the ambiguity and discontinuity issues of Euler angles representation could interfere with the performance of challenging samples. In this paper, we formulate the head pose estimation problem into quaternion representation space and propose a novel framework named Structural Relation-aware Network (SRNet). Different from previous methods, our SRNet explicitly explores the correlation among different regions of the face for mining global facial structure information. Furthermore, in order to boost robustness and generalization of the model, a hard example mining (HEM) strategy is designed to mitigate the data imbalance issue by adjusting the contributions of examples in different states to loss. Extensive experiments demonstrate that our method outperforms the current state-of-the-art alternatives on the public benchmark datasets: AFLW2000 and BIWI.
Zhaoxiang Zeng, Dongchen Zhu, Wenjun Shi, Lei Wang 0202, Jiamao Li
ICPR5