Zhen Zhou 0003

dblp:23/759-3 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
16since 2021 · last 2026
0009-0001-2112-3742ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Uncertainty-Adaptive Volume for Unsupervised Homography Estimation
abstract
Estimating homography from an image pair is crucial for image alignment, and unsupervised methods that optimize feature reprojection error between target and warped source images have gained attention for their promising performance. In real-world scenes with multiple planes, such as moving objects, outlier rejection strategies are essential to mitigate the influence of non-dominant planes. Existing methods address this by learning a mask based on reprojection error, where high errors indicate non-dominant planes misaligned by homography. However, this error-fitting mask often overextends to the dominant plane, limiting the use of valid image regions for accurate estimation. This paper proposes a novel unsupervised method to compactly exclude non-dominant planes by introducing an uncertainty-adaptive cost volume for homography estimation. We first model uncertainty by assuming image features follow a Gaussian distribution derived from a prior Normal Inverse-Gamma distribution. The network-learned distribution parameters disentangle aleatoric uncertainty, distinguishing data-dependent errors within the total reprojection error. This uncertainty reflects inherent observation noise in image data, effectively indicating non-dominant planes. We then integrate this aleatoric uncertainty into the concatenation volume across image feature maps, creating an adaptive volume that filters out unreliable matching costs associated with non-dominant planes. This adaptive volume simplifies learning homography from the rich, redundant content in the concatenation volume, enabling more efficient and accurate estimation. Experiments demonstrate that our method outperforms existing approaches, achieving state-of-the-art performance both qualitatively and quantitatively.
Jianqiao Luo, Yaonan Wang 0001, Mingtao Feng, Zhen Zhou 0003, Xuebing Liu, Yang Mo
IEEE Trans. Circuits Syst. Video Technol.4
2026 Self-Expert Imitation With Purifying Latent Feature for Generalization in Visual Reinforcement Learning
abstract
The generalization ability of visual reinforcement learning, which allows the policy trained in the source domain to guide agents in similar unknown target environments, is one of the cores applied to visual navigation and autonomous driving. Recently, methods such as data augmentation techniques, self-supervised learning methods, and the generative adversarial network were employed to enhance the generalization capability of policy neural networks in visual reinforcement learning. However, current state-of-the-art methods, after utilizing domain-general latent features to train the RL policy, result in the loss of certain state-specific features, leading to diminished policy performance following generalization. To tackle these challenges, we designed a technical framework called self-expert imitation with purifying latent features, which enables the trained policy to effectively guide agents in scenarios similar to the training environment, without compromising the performance of the policy-guided agent in task completion. Additionally, a novel method was developed for separating domain-general and domain-specific latent vectors based on a variational autoencoder, enabling the domain-general component to exhibit strong and stable zero-shot generalization performance in unseen visually similar domains. Extensive experiments on the CarRacing game demonstrated that our approach achieves strong and stable generalization performance in unseen environments, without compromising the performance of the policy in guiding agents to complete tasks.
Lin Chen 0034, Yang Mo, Yaonan Wang 0001, Zhiqiang Miao, Kai Zeng 0010, Mingtao Feng, Zhen Zhou 0003, Sifei Wang, Danwei Wang
IEEE Trans. Intell. Transp. Syst.7
2025 Multi-range Adaptive Perception Transformer for Iterative Homography Estimation
abstract
Homography estimation is fundamental to various vision tasks. Iteration-based methods have recently achieved significant success in this field. However, errors introduced during iterations can lead to increased image deformation. Existing methods often focus on capturing local correspondences in the later stages of iteration while downplaying global ones, which may cause errors to persist and propagate into subsequent iterations, ultimately leading to error accumulation. To alleviate this issue, we propose Multi-range Adaptive Perception Transformer for Iterative Homography Estimation (MAPTHomo), which integrates Multi-range Attention (MRA) and Adaptive Perception Module (APM). Specifically, MRA captures both global and local correspondences, enabling the model to adapt to varying levels of deformation. The APM dynamically adjusts attention focus based on the current context. The combination of MRA and APM enhances the error-correction capability of the iterative process, effectively mitigating error accumulation. Extensive experiments demonstrate that MAPTHomo outperforms previous methods and exhibits strong generalization ability.
Tianming Li, Qing Zhu 0003, Zhen Zhou 0003, Jianqiao Luo, Yaonan Wang 0001
ICASSP3
2025 Decentralized Multi-robot Navigation Policy with Enhanced Security Using Graph GRU Policy Network
abstract
Formulating a multi-robot obstacle avoidance policy is essential for enabling safe and efficient navigation in multi-robot environments, forming a critical component of the effective operation of multi-robot systems. Recently, reinforcement learning has been applied to improve the performance of decentralized, policy-driven robots in task execution. However, ensuring the safety of these agents during movement remains a significant challenge due to the inherent risks associated with the reinforcement learning process, such as frequent collisions. To address this issue and enhance the safety of policy-guided multi-robot navigation, we propose a novel policy based on imitation learning. This framework introduces a novel policy neural network that integrates a graph attention mechanism with the GRU network structure. The key innovation lies in utilizing the interactions between neighboring robots to enhance the safety of their movements. In a multi-robot simulation environment, robot behaviors are directed by the proposed policy. A comparative analysis was conducted between our approach and RL-RVO, one of the advanced methods in the field. The results demonstrate that our approach outperforms RL-RVO, achieving a higher success rate and significantly improving safety performance.
Lin Chen 0034, Yuxuan Ao, Zhen Zhou 0003, Yaonan Wang 0001, Danwei Wang
IROS3
2025 DRLHomo: Disentangled Representation Learning for Cross-Modal Homography Estimation
Tianming Li, Zhen Zhou 0003, Qing Zhu 0003, Jianqiao Luo, Yaonan Wang 0001
PRCV (9)2
2025 RAMPGrasp: Retentive Attention-Based Multiscale Perception Grasp Detection Network
abstract
In robotic grasp detection, challenges such as uncertainty in object type, size, and placement within the scene diminish grasping accuracy. However, the inability to effectively locate the graspable area and incomplete feature extraction for grasp detection are two key factors that hinder grasp detection accuracy and are not considered in current methods. This paper presents a novel retentive attention-based multiscale perception grasp detection network (RAMPGrasp) to address this constraint. First, we introduce retentive attention in the feature extraction module, which significantly improves the efficiency of attention score computation for long sequences in visual tasks. Second, we propose a multiscale spatial pyramid attention module, which can effectively adjust the importance of multiscale feature sequences and feature channels, while enhancing the correlation of multiscale features. Third, we design the prediction module as a coarse-to-fine framework, improving feature representation for grasp detection by considering the distribution trend of grasp poses. As a result, RAMPGrasp achieves state-of-the-art grasp detection accuracy, with 98.4% and 95.6% on the Cornell and Jacquard datasets, respectively.
Jianan Huang 0002, Xuebing Liu, Qing Zhu 0003, Yaonan Wang 0001, Mingtao Feng, Zhen Zhou 0003, Lin Chen 0034, Danwei Wang
IEEE Trans. Circuits Syst. Video Technol.7
2025 Multiscale Spherical Feature Decoupling Network for Multimodal Image Registration
abstract
Multimodal image registration plays a crucial role in advancing Earth science. However, significant appearance variations and geometric deformations between multimodal images pose considerable challenges to this task. In this paper, we propose a novel multiscale spherical feature decoupling network (MSFDNet) for multimodal image registration by combining a multiscale iterative strategy with a multimodal decoupling strategy. MSFDNet adopts a multiscale architecture, with each scale incorporates a spherical feature decoupling (SFD) module with a carefully crafted three-stage decoupling strategy to bridge the modality gap. Specifically, we first introduce asymmetric shared and unique feature encoders to extract modality-shared and modality-unique features. Next, we design a spherical constraint learning (SCL) module to project the extracted features into spherical space, leveraging its regularized distance properties to enhance feature separability during the decoupling process. Finally, we propose a dual-path reconstruction mechanism that combines self-reconstruction with homography-guided cross-reconstruction to reconstruct the original multimodal images from the decoupled features, thereby simultaneously enhancing the learning of both feature decoupling and registration network. Based on the decoupled modality-shared features, we predict the registration function in a multiscale iterative manner, effectively bridging the geometric gap. Extensive experiments on multiple multimodal datasets validate the effectiveness of MSFDNet and demonstrate its state-of-the-art performance.
Tianming Li, Zhen Zhou 0003, Qing Zhu 0003, Jianqiao Luo, Yaonan Wang 0001
IEEE Trans. Geosci. Remote. Sens.2
2025 Uncertainty Guided Deep Lucas-Kanade Homography for Multimodal Image Alignment
abstract
Homography estimation for multimodal images poses a considerable challenge in computer vision because of content disparities and the diverse feature points captured by different sensors. Existing methods typically extract feature maps using neural networks and apply the Lucas-Kanade (LK) algorithm, which is based on the brightness constancy assumption, to solve the homography matrix. However, applying this assumption across all pixel features in multimodal images can lead to inaccuracies, as these images often contain noise, such as homogeneous regions or considerable appearance variations, which can corrupt the network’s training. To address this problem, we propose an uncertainty-guided deep LK (UG-DLK) framework that integrates uncertainty predictions to enhance the network’s iterative learning process. Specifically, we employ a probabilistic approach where the network predicts the distribution of the feature map rather than fixed values. By designing an uncertainty neighborhood estimator, we unfold the cost volume along the channels into 2-D slices, allowing the model to focus on neighborhood information at specific locations, effectively reducing the interference from spatial neighborhoods in the estimation of feature uncertainty. Through uncertainty modeling, the network can accurately identify scenes and objects that comply with the brightness constancy constraint, leading to more robust learning outcomes. Additionally, we introduce a novel loss function that incorporates feature uncertainty, leading to a smoother optimization landscape near the true homography parameters and reducing convergence oscillations. Our method, which is evaluated on benchmark datasets such as Google Maps, Google Earth, MSCOCO, and DPDN, demonstrates state-of-the-art performance, confirming the robustness and adaptability of our model across various scenarios.
Zhen Zhou 0003, Jianqiao Luo, Qing Zhu 0003, Yaonan Wang 0001, Hang Zhong, Mingtao Feng, Lin Chen 0034
IEEE Trans. Geosci. Remote. Sens.1
2024 Domain Adaptation in Visual Reinforcement Learning via Self-Expert Imitation with Purifying Latent Feature
abstract
Generalizing visual reinforcement learning is fundamental to robot visual navigation, involving the acquisition of a policy from interactions with source environments to facilitate adaptation to analogous, yet unfamiliar target environments. Recent advancements capitalize on data augmentation techniques, self-supervised learning methods, and the generative adversarial network framework to train policy neural networks with enhanced generalizability. However, current methods, upon extracting domain-general latent features, further utilize these features to train the reinforcement learning policy, resulting in a decline in the performance of the learned policy guiding the agent to accomplish tasks. To tackle these challenges, a framework of self-expert imitation with purifying latent features was devised, empowering the policy to achieve robust and stable zero-shot generalization performance in visually similar domains previously unseen, without diminishing the performance of guiding the agent to accomplish tasks. The extraction method of domain-general latent features is proposed to enhance their quality based on the variational autoencoder. Extensive experiments have shown that our policy, compared with state-of-the-art counterparts, does not diminish the performance of the policy guiding the agent to accomplish tasks after generalization.
Lin Chen 0034, Jianan Huang 0002, Zhen Zhou 0003, Yaonan Wang 0001, Yang Mo, Zhiqiang Miao, Kai Zeng 0010, Mingtao Feng, Danwei Wang
IROS3
2024 Decentralized Multi-Robot Navigation Coupled with Spatial-Temporal RetNet Based on Deep Reinforcement Learning
abstract
Navigating robots through dynamic multi-robot environments, avoiding collisions with both other robots and obstacles, has emerged as a central challenge in robotics. The existing approaches fall short in allowing the policy network to effectively capture spatial-temporal reciprocal collision avoidance in multi-robot environments, comprising both static and dynamic obstacles, resulting in inadequate safety and efficiency in directing robot movement. In this study, we introduce a novel policy neural network called Spatial-Temporal RetNet (STR), designed to encode reciprocal collision avoidance states between robots in spatial and temporal dimensions. The goal is to improve the safety and efficacy of the policy neural network in directing robots to complete assigned tasks. The spatial state encoder module is built upon a parallel RetNet structure, which strengthens the neural network's capacity in extracting reciprocal collision avoidance states between robots in spatial dimensions. This module addresses the limitations of position encoding in transformer-based multi-robot navigation policy neural networks. We design a temporal state encoder utilizing a recurrent RetNet structure. This innovation bolsters the multi-robot navigation policy neural network's capability to capture features in the temporal dimension of multi-robot movements. It addresses the limitations of transformer-based multi-robot navigation policy neural networks, particularly in recurrently inferring information across time dimensions. Simulation experiments were conducted to showcase the superior safety and effectiveness of our proposed method compared to previous state-of-the-art approaches in guiding robots to accomplish tasks.
Lin Chen 0034, Yaonan Wang 0001, Zhiqiang Miao, Mingtao Feng, Yuanzhe Wang, Yang Mo, Zhen Zhou 0003, Hesheng Wang 0001, Danwei Wang
IROS7
2024 Interpretable Unsupervised Homography Estimation
Zhen Zhou 0003, Qing Zhu 0003, Yaonan Wang 0001, Yang Mo, Lin Chen 0034, Jianan Huang 0002, Tianjian Jiang
PRCV (2)1
2024 Toward Safe Distributed Multi-Robot Navigation Coupled With Variational Bayesian Model
abstract
Designing a safe and effective collision avoidance policy for multiple robots is essential in decentralized scenarios, where each robot is responsible for generating its own paths, to ensure their safe operation. Recently, the utilization of reinforcement learning to develop decentralized policies that enable multiple robots to move cooperatively and accomplish tasks has yielded positive outcomes. However, the presence of exploration unsafe actions during the reinforcement learning training process results in inadequate safety. We seek to enhance the safety of distributed multi-robot navigation policies and propose a new imitation learning framework based on the variational Bayesian model, which enables robots to learn safe actions by anticipating the subsequent state they are expected to reach. In addition, a new policy neural network structure for multi-robot navigation is proposed by introducing the transformer structure, which encodes the significance of nearby robots in relation to their forthcoming conditions. Experiments demonstrated that our policy can more safely guide robots to navigate in multi-robot environments under conditions of limited information, outperforming the state-of-the-art RL-RVO method in terms of success rate.Note to Practitioners—The motivation of this paper is to address the problem of collision avoidance in a multi-robot environment under limited information, which can also be applied to autonomous driving, crowd simulation, and other related fields. Positive outcomes have been observed in the utilization of reinforcement learning to create decentralized policies that enable multiple robots to move cooperatively and complete tasks. However, inadequate safety remains a challenging task due to the possibility of exploring hazardous actions during training. This article aims to enhance the safety of distributed policies guiding robots to accomplish navigation tasks in dynamic multi-robot environments. To begin with, we introduce a novel framework for imitation learning that is based on the variational Bayesian model. This framework facilitates the learning of safe actions by the policy to improve its performance and guide the robot in navigating and avoiding obstacles more securely. A loss function is proposed that enables the anticipation of the future state expected to be reached by the robot. By incorporating the transformer structure, a new neural network structure is designed for multi-robot navigation that encodes the significance of nearby robots concerning their upcoming conditions. This network structure employs a BiGRUs to facilitate the assimilation of observations from multiple agents by the policy. Compared to existing works such as GA3C-CADRL, SARL, and RL-RVO, our proposed method achieves a higher success rate. In our future research, we will investigate methods to enhance the policy’s performance in guiding robots to complete tasks by focusing on improving travel time and average speed, while also strictly ensuring safe navigation. Furthermore, we plan to extend this approach by addressing navigation challenges in more densely populated multi-robot environments.
Lin Chen 0034, Yaonan Wang 0001, Zhiqiang Miao, Mingtao Feng, Zhen Zhou 0003, Hesheng Wang 0001, Danwei Wang
IEEE Trans Autom. Sci. Eng.5
2024 Unsupervised Homography Estimation With Pixel-Level SVDD
abstract
Homography estimation is a common image alignment method. Unsupervised learning, which uses unlabeled training and exhibits excellent performance, has attracted much attention in this field. When there are multiple planes in the scene, using features over the entire image for matching will lead to compromised results. However, existing methods for learning focused principal plane masks through deep neural networks lack explicit guidance. In this paper, we propose a novel unsupervised method to explicitly model anomaly descriptor removal and mask generation. Specifically, reliable feature descriptors are selected from a novel perspective, and regard the features that are not responsible for alignment as outliers. The pixel-level support vector data description (PL-SVDD) module is designed. This module learns the feature representation of image pixels and fits a hypersphere to exclude the feature redundancy information that is not responsible for alignment from the hypersphere, thereby optimizing the feature descriptor. Based on the optimized image features, a correlation learning (CL) module is designed. This module displays a generated mask through mathematical modeling to select reliable areas for homography estimation. Specifically, the feature descriptor of one unaligned images is modeled as a multivariate Gaussian distribution by Gaussian density estimation (GDE). Then, The Mahalanobis distance is combined with the multivariate Gaussian distribution of the model and the feature descriptor of another image to generate the mask. Experiments show that our method achieves good performance compared with previous methods.
Zhen Zhou 0003, Qing Zhu 0003, Mingtao Feng, Yaonan Wang 0001, Jianqiao Luo, Zhiqiang Miao, Lin Chen 0034, Yang Mo
IEEE Trans. Circuits Syst. Video Technol.1
2024 A Depth Adaptive Feature Extraction and Dense Prediction Network for 6-D Pose Estimation in Robotic Grasping
abstract
Estimating the 6-D pose of an object is a vital and challenging task for robot vision systems in industrial robotic grasping. With the wide use of 3-D cameras, the additional acquired depth image provides geometric information of the scene to increase the pose estimation performance but leads to a challenge, fully leveraging the two-modal data, the color image and the depth image. Previous works usually adopt two individual strategies to handle the data, which suffer from limited accuracy and efficiency since the two complementary data are not fully explored. Thus, we propose a depth adaptive feature extraction and dense prediction network that decouples the scale-dependent and the scale-invariant information from the depth image. The former guides the network to perceive the 3-D structure of the scene, and the latter, together with color image, provides the scene textures for feature extraction. The proposed network not only fuses multimodal textures but also retains their 3-D structure. In addition, a dense prediction strategy is adopted to regress the object pose; this approach can mitigate the instability caused by outliers. We conduct various evaluations on a real-world industrial dataset to illustrate the advantages of the proposed approach; and a practical robotic grasping platform is presented to demonstrate its application performance.
Xuebing Liu, Xiaofang Yuan, Qing Zhu 0003, Yaonan Wang 0001, Mingtao Feng, Zhen Zhou 0003
IEEE Trans. Ind. Informatics7
2023 Improved YOLOv7 Based on Transformer for Object Detection in UAV-Captured Images
abstract
As the drone captures image targets at different flying altitudes, their scales may vary significantly, which can pose challenges for the object detection model to accurately detect them. Additionally, tiny objects in the image contain minimal information, making them difficult to distinguish from the background. To overcome these two challenges, we proposed a network architecture that aims to improve the accuracy of tiny object detection in drone images. Specially, we designed a tiny object detector(TOD) that can effectively extract features of tiny objects and distinguish between tiny object features and image background. Furthermore, this TOD module contains a Convolutional Visual Attention Network (CVAN) to better focus on the regions of tiny objects. Experimental results demonstrate that the proposed method achieves [email protected] accuracy of 53.9% on the VisDrone2021-test-dev dataset and improves by 2.8 % compared to YOLOv7.
Yuefan Luo, Qing Zhu 0003, Zhen Zhou 0003, Lin Chen 0034, Tianjian Jiang, Yijiang Li, Danwei Wang, Yaonan Wang 0001
SMC3
2023 Transformer-Based Imitative Reinforcement Learning for Multirobot Path Planning
abstract
Multirobot path planning leads multiple robots from start positions to designated goal positions by generating efficient and collision-free paths. Multirobot systems realize coordination solutions and decentralized path planning, which is essential for large-scale systems. The state-of-the-art decentralized methods utilize imitation learning and reinforcement learning methods to teach fully decentralized policies, dramatically improving their performance. However, these methods cannot enable robots to perform tasks efficiently in relatively dense environments without communication between robots. We introduce the transformer structure into policy neural networks for the first time, dramatically enhancing the ability of policy neural networks to extract features that facilitate collaboration between robots. It mainly focuses on improving the performance of policies in relatively dense multirobot environments under conditions where robots do not communicate with each other. Furthermore, a novel imitation reinforcement learning framework is proposed by combining contrastive learning and double deep Q-network to solve the problem of difficulty training policy neural networks after introducing the transformer structure. We present results in the simulation environment and compare the resulting policy against advanced multirobot path-planning methods in terms of success rate. Simulation results show that our policy achieves state-of-the-art performance when there is no communication between robots. Finally, we experimented with a real-world case using a total of three robots in our robotic laboratory.
Lin Chen 0034, Yaonan Wang 0001, Zhiqiang Miao, Yang Mo, Mingtao Feng, Zhen Zhou 0003, Hesheng Wang 0001
IEEE Trans. Ind. Informatics6